Agentic AIOps β real detection, diagnosis, and automated remediation with a human in the loop.
A production-credible, end-to-end AIOps platform built from scratch: seven Python microservices, a React operator console, a real Kubernetes remediation path, a self-improving ML pipeline, and a Deloitte-style financial platform as its live target. Every number on the dashboard is real; every fix is reversible; every decision is audited.
πΉ Demo video coming soon β live incident detection β diagnosis β HITL approval β real pod remediation, in under 3 minutes.
Modern cloud estates emit 500β1,200 alerts per day. The large majority is noise. SREs burn their on-call hours triaging it instead of fixing things, which keeps MTTR stuck in the hours range and drives burnout.
IntelliOps attacks that from four angles simultaneously:
| Problem | What IntelliOps does |
|---|---|
| Alert storm β unreadable noise | Collapses raw alerts into a handful of meaningful Situations using online ML (z-score, robust seasonal, or a trained IsolationForest) |
| "Which service, which cause?" | Enriches each Situation with deploy history, topology, and metric evidence; ranks hypotheses; explains the top one in plain language |
| "Can we automate the fix?" | Executes safe, reversible remediation β scale, restart, rollback, config patch β after a pre-flight rehearsal on an isolated cluster clone |
| "Did it work, and who approved it?" | Re-checks the firing metric after every fix; rolls back if unhealthy; writes every decision to an immutable audit trail |
And the system learns: every outcome (success, rollback, escalation) feeds back into the correlation model, so accuracy compounds rather than freezing at deployment-day quality.
This isn't a mock dashboard. The full stack runs in a single docker compose up:
- Seven Python microservices (FastAPI) talking over Redis Streams (or Kafka)
- Real Prometheus scraping a real breakable target β every metric on the console is scraped, not generated
- A live SSE pipeline that streams each incident through Detect β Correlate β Diagnose β Approve β Execute β Verify in real time
- Real Kubernetes remediation on a local kind cluster β approving a fix restarts/scales a real pod; a real health check verifies it
- Postgres-backed durable state β the audit trail, playbook registry, pending approvals, and the detector's learned baseline all survive a restart
- Meridian β a four-service Deloitte-style financial/audit platform (gateway, validation, aggregation, reporting + its own portal UI) running alongside IntelliOps as a genuine production-shaped target, emitting 11 USE+RED gauges across 8 typed fault scenarios
telemetry βββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
sources β governance-service (CoE) β
(Prometheus, β RBAC Β· audit log Β· playbook registry β
Loki, OTel) ββββββββ² sync approval gate ββββββββββ² audit (async)ββββ
β β β
βΌ βββββββ΄βββββββ ββββββββββββ ββββ΄ββββββββ ββββββββββββ
ββββββββββββ βcorrelation β β rca β β action β β feedback β
βingestion βββrawβββββββββΆβ detect + ββββΆβ enrich + ββββΆβ approve, ββββΆβ label + β
βnormalize β β cluster β β β rank + β β execute, β β retrain β
β + dedup β β Situation β β runbook β β rollback β β (metrics)β
ββββββββββββ βββββββ²βββββββ ββββββββββββ ββββββββββββ ββββββ¬ββββββ
β β
βββββββββββββββββ retrain (closed loop) ββββββββ
β²
read-service (CQRS)
React console (SSE)
Six services communicate over an event bus. The only synchronous step is action β governance β the human-in-the-loop gate is enforced in the call graph itself, not by convention. Full walkthrough in flow.md; the reasoning behind every decision is in architectural.md.
- Augment, don't replace β IntelliOps sits alongside your existing observability and ticketing stack; it never becomes the system of record
- Human-in-the-loop by construction β automated action is structurally gated on a governance decision; the gate fails closed
- Reversible-only automation β the system only automates what it can undo; it verifies health after acting and rolls back when the metric doesn't recover
- The loop closes β every remediation outcome is training data; the correlation model improves as it operates
- Open-source-first β every named tool sits behind an interface and is swappable; no vendor lock-in
| Concern | Default | Opt-in / swap path |
|---|---|---|
| Services | Python 3.11 Β· FastAPI Β· Pydantic v2 | β |
| Event bus | Redis Streams | Kafka (BUS_BACKEND=kafka) |
| Telemetry | Prometheus Β· Loki Β· OpenTelemetry | any, via TelemetrySource |
| Correlation / ML | River (online z-score) | Robust seasonal z-score Β· scikit-learn IsolationForest (CORRELATOR_KIND) |
| Runbook selection | Keyword rules | Embedding similarity β sentence-transformers all-MiniLM-L6-v2 (RUNBOOK_SELECTOR_MODE=embedding, now on by default) |
| Remediation | Dry-run (safe default) | Kubernetes API real pod ops (REMEDIATOR_MODE=k8s) |
| Pre-flight | β | Sandbox rehearsal on isolated namespace clone (SANDBOX_MODE=k8s) |
| Persistence | Postgres (SQLAlchemy Core + Alembic) | β |
| Frontend | React 19 Β· TypeScript Β· Tailwind | β |
| Auth | Off (dev default) | Bearer token (AUTH_MODE=token) |
| Deploy | Docker Compose | Helm chart (deploy/k8s/) |
# 1. Start the full stack (Redis, Postgres, 7 services, Prometheus, demo app, Meridian)
docker compose -f deploy/docker-compose.yml up --build
# 2. Start the console in live mode
cd frontend && cp .env.example .env.local
# edit .env.local β VITE_DATA_MODE=live
npm install && npm run dev
# open http://localhost:5173Logs, metrics and the audit trail in one place: Grafana at http://localhost:3000 (on kind: http://localhost:30300). Every container's logs (via Loki), the Prometheus metrics, and the Postgres audit trail on one dashboard, viewable without logging in. See docs/OBSERVABILITY.md.
Drive a real incident end to end:
# Break the demo app and watch the pipeline animate in the console
./scripts/chaos.sh
# Approve the fix in the UI (Incidents β open Situation β Approve)
# Reset between runs β clears the detector baseline and read model
./scripts/reset.shDetection takes 15β30 seconds β that's a real Prometheus scrape cycle plus River needing a few samples to flag the anomaly, not a mock timer.
For a step-by-step demo walkthrough (including real kind-cluster remediation), see docs/DEMO.md.
# deploy/.env (gitignored)
VITE_OPERATOR_NAME=your-name # shows in approvals + audit trail
INTELLIOPS_LLM_ENDPOINT=https://api.groq.com/openai/v1 # optional: LLM explanations
INTELLIOPS_LLM_API_KEY=gsk_...
INTELLIOPS_LLM_MODEL=llama-3.3-70b-versatileEvery incident in a demo needs a real target. Meridian (services/meridian/) is a four-service Deloitte-style financial/audit reporting platform β gateway, validation, aggregation, reporting β with its own client portal and ops panel. It runs alongside IntelliOps in the same compose file and gives the pipeline something genuinely production-shaped to watch.
Each service emits a USE+RED metric set (11 gauges: CPU, memory, disk, queue depth, DB-pool, request rate, error rate, p50/p99 latency) and accepts 8 typed fault scenarios, each moving a realistic cluster of those metrics rather than a single gauge.
Three scenarios were verified live end-to-end against real Docker, each producing the expected, genuinely different diagnosis:
| Fault | Service | Correct diagnosis |
|---|---|---|
| CPU saturation | meridian-aggregation | scale-service |
| Error-rate spike | meridian-validation | restart-pod |
| Recent deploy + saturation | meridian-gateway | rollback-deploy (outranks saturation) |
Meridian is wired to IntelliOps through additive-only changes β a Prometheus scrape job per service and a broadened ingestion selector. No IntelliOps service code changed. See docs/MERIDIAN.md.
- HITL gate fails closed β a timeout, rejection, or unreachable governance service means no action, never a default execute
- Destructive-action denylist β a typed vocabulary gate prevents catastrophic actions from being expressed at all; the
Literalaction type is permanently closed - Pre-flight sandbox β fixes are rehearsed on an isolated namespace clone; a failed rehearsal blocks auto-remediation and surfaces the verdict in the UI
- Immutable audit trail β every RBAC decision, approval, execution, and outcome is recorded in Postgres, threaded by
correlation_id - AI is bounded β the AI can draft a runbook for a gap, but it cannot join the registry without a human approval; it can rank vetted playbooks by embedding similarity, but it never chooses an unvetted fix
- Compliance-aligned β NIST AI RMF (Govern/Map/Measure/Manage), EU AI Act risk-tiered documentation, DORA 4-hour notification window; deployable on-prem for sovereign-cloud needs
uv run pytest # 433+ tests, no Docker needed for the core suite
uv run ruff check . # linting
npm run build # frontend type-check + contrast gate (112 checks)CI runs on every PR: lint β test β frontend build β compose smoke test. The correlator improvement is covered by a reproducible benchmark committed to the repo.
| Area | Status |
|---|---|
| Seven-service closed loop (detect β diagnose β approve β execute β verify β retrain) | β |
| Real Kubernetes remediation + rollback on a kind cluster | β |
| Postgres persistence β audit trail, playbook registry, durable approvals, baseline snapshot | β |
| Pluggable detectors β River (online z-score), robust seasonal, trained IsolationForest | β |
| Reliability-weighted + LLM-explained RCA with a committed CI benchmark | β |
| Semantic runbook selection via embedding similarity (now on by default) | β |
| AI-authored runbook drafting β propose-only, human-approval gate, type-safe actions | β |
| Pre-flight sandbox rehearsal on an isolated namespace clone | β |
Destructive-action denylist (7 typed Deployment-scoped verbs; Literal closed) |
β |
| Real-time React console β SSE pipeline view, audit explorer, light/dark theme | β |
| Agent Activity trace β live step-by-step view of the AI runbook author as it works | β |
| Meridian sample financial platform β 4 services, 8 fault scenarios, verified live | β |
| Edge auth (bearer token, timing-safe), structured JSON logging, readiness probes | β |
| Kafka bus binding + whole-stack Helm deploy | β |
| 433+ tests Β· CI pipeline Β· ruff-clean | β |
| Document | What it covers |
|---|---|
| architectural.md | Design principles, layer mapping, 30 ADRs with context and trade-offs |
| flow.md | One-incident journey, bus topics, data contracts, per-function service reference |
| docs/DEMO.md | Guided two-act demo β live loop then real kind-cluster remediation |
| docs/MERIDIAN.md | The sample financial platform β services, fault scenarios, verified results, honest limits |
| docs/BENCHMARKS.md | Correlator benchmark β methodology, results table, where robust/trained win and cost |
| docs/PERSISTENCE.md | Postgres backend β schema, STORE_BACKEND switch, migrations, durable runtime state |
| docs/UI.md | Operator console β five views, mock vs. live mode, SSE architecture, theme system |
| docs/OPERATIONS.md | Deploy guide, full env-switch table, auth model |
| docs/OBSERVABILITY.md | Structured JSON logging, liveness vs. readiness probes |
| deploy/k8s/README.md | Real-remediation demo on a kind cluster |
intelliops/
βββ README.md β you are here
βββ architectural.md β why: 30 ADRs, layer model, compliance mapping
βββ flow.md β how: data flow, bus contracts, per-function reference
βββ common/ β shared library: contracts, interfaces, bus client, config, auth
βββ services/
β βββ ingestion/ β normalize, dedup, ingest from Prometheus/Loki
β βββ correlation/ β detect anomalies, cluster into Situations
β βββ rca/ β enrich, rank hypotheses, select + embed runbooks, explain
β βββ action/ β HITL gate, sandbox, execute, verify, rollback
β βββ governance/ β RBAC, audit log, playbook registry, AI runbook author
β βββ feedback/ β label outcomes, retrain, graduate playbooks, compute KPIs
β βββ read/ β CQRS read model, SSE push, Prometheus proxy
β βββ meridian/ β 4-service sample financial platform + portal UI
βββ frontend/ β React 19 operator console (TypeScript, Tailwind)
βββ playbooks/ β YAML playbook definitions
βββ alembic/ β Postgres schema migrations
βββ deploy/
β βββ docker-compose.yml β full dev stack
β βββ k8s/ β Helm chart + kind demo
βββ tests/ β 433+ tests (pytest)
βββ docs/ β DEMO, MERIDIAN, BENCHMARKS, PERSISTENCE, UI, OPERATIONS
βββ pyproject.toml
SRE/DevOps engineers spend most of their on-call hours on manual log correlation and alert triage β not on resolution. Downtime costs enterprises roughly $15,000/minute (Splunk & Cisco with Oxford Economics, 2026), yet most AIOps setups treat correlation and remediation as disconnected, static systems that can't improve over time.
IntelliOps is built around two ideas:
- The loop closes. Remediation outcomes feed back into the correlation model so accuracy compounds instead of freezing at deployment-day quality.
- A governed Center of Excellence, not a point tool. RBAC, audit, rollback, and a shared playbook registry are a single control plane β countering the well-documented point-solution anti-pattern in AIOps adoption.
Target outcomes (from the proposal, grounded in cited industry data):
| KPI | Target |
|---|---|
| MTTR reduction | 40β60% |
| Alert volume reduction | 80β95% |
| Low-risk incidents auto-remediated | 30β60% (phased) |
| SRE on-call burden | ~30β40% reduction |