ongridio/ongrid
Ongrid is an AI-powered ops agent that autonomously investigates alerts, performs root-cause analysis, and orchestrates fixes for IT infrastructure issues, interacting via chat platforms.
Awesome AI for Infra › Incident Management & Response
RunbookHermes leverages the core capabilities of the Hermes Agent while adding a specialized AIOps/SRE domain layer to address real-world production incidents. Its primary function is to process alerts, gather comprehensive evidence from metrics, logs, traces, and deployment history, analyze root causes, propose remediation actions, and facilitate approved execution. A key design principle is "Evidence first," ensuring that decisions are based on concrete data rather than speculative guesses. It also emphasizes that "Memory is weak prior," meaning historical data serves as context but not a substitute for current evidence. For critical and potentially destructive operations, RunbookHermes implements a robust safety chain, requiring approval, checkpoints, dry-runs, and recovery verification to prevent unintended consequences. Crucially, the system is designed to learn from every incident. After each resolution, RunbookHermes refines its knowledge base by capturing incident summaries, fault patterns, service governance rules, team habits, and generating reusable runbook skills, RAG documents, and evaluation benchmarks. This self-evolving memory ensures that the system continuously improves its understanding of the IT environment and becomes more effective with each use, thereby embodying the principle of "Every incident improves the next incident." The project includes a web console for AIOps overview, real-time monitoring, incident command and detail views, approval workflows, and a self-evolving memory console.
https://github.com/Tommy-yw/RunbookHermes
Ongrid is an AI-powered ops agent that autonomously investigates alerts, performs root-cause analysis, and orchestrates fixes for IT infrastructure issues, interacting via chat platforms.
Aurora is an open-source, AI-powered incident management platform that uses AI agents to autonomously investigate incidents, perform root cause analysis, and suggest remediations across multi-cloud...
OpenSRE is an open-source AI SRE agent that automates incident investigation, root cause analysis, and learns from past incidents using episodic memory and a knowledge graph.
AI SRE for Kubernetes providing zero-instrumentation eBPF observability, an AI copilot for incident remediation via guardrailed, self-verifying actions, and air-gapped capable operations.
This project implements a multi-agent ChatOps system that uses AI/ML to autonomously triage, investigate, and propose fixes for infrastructure alerts, including self-improving prompts and a causal ...