Tommy-yw/RunbookHermes
RunbookHermes is an AIOps agent built on the Hermes Agent framework, designed for evidence-driven incident response, approval-gated remediation, and self-evolving runbook learning.
Awesome AI for Infra › Incident Management & Response
Ongrid provides an AI-driven operations agent designed to understand and manage IT infrastructure, pinpointing root causes of incidents and facilitating their resolution directly through chat interfaces like Slack or Telegram. It leverages a coordinator-specialist agent architecture, where a main coordinator dispatches tasks to specialized sub-agents focused on SRE, network, or database issues. The tool automatically investigates alerts, performing root cause analysis by correlating metrics, logs, traces, and topological blast-radius information, often tracing issues to specific lines of code. Ongrid features a secure architecture with zero inbound ports and browser-based SSH for audited remote execution. It comes with built-in observability using Prometheus, Loki, Tempo, and Grafana, with the agent intelligently crafting queries. The platform supports various large language models (LLMs) and integrates with existing observability stacks and communication channels, making it a comprehensive solution for AI-powered incident management and automated IT operations.
https://github.com/ongridio/ongrid
RunbookHermes is an AIOps agent built on the Hermes Agent framework, designed for evidence-driven incident response, approval-gated remediation, and self-evolving runbook learning.
Aurora is an open-source, AI-powered incident management platform that uses AI agents to autonomously investigate incidents, perform root cause analysis, and suggest remediations across multi-cloud...
OpenSRE is an open-source AI SRE agent that automates incident investigation, root cause analysis, and learns from past incidents using episodic memory and a knowledge graph.
AI SRE for Kubernetes providing zero-instrumentation eBPF observability, an AI copilot for incident remediation via guardrailed, self-verifying actions, and air-gapped capable operations.
This project implements a multi-agent ChatOps system that uses AI/ML to autonomously triage, investigate, and propose fixes for infrastructure alerts, including self-improving prompts and a causal ...