AI· Case study 07
AI on-call triage agent
A production AI on-call agent that turns monitoring alerts into triaged, deduplicated tickets and Slack updates, so engineers start from a diagnosis instead of a raw alert.
Auto-triage
alerts classified, deduplicated, and ticketed
Architecture at a glance
- Monitoring webhook
- FastAPI
- LLM structured output
- Ticket + Slack
01 The problem
Alert noise and repeated failures ate on-call time. Every alert needed someone to open logs, classify the issue, and file or update a ticket.
02 What I built
- 1A FastAPI service with authenticated webhook endpoints for monitoring alerts.
- 2An LLM pipeline that reads logs and health status and returns structured-output classifications: issue category, likely cause, and severity.
- 3Ticket automation that creates or updates issues only for persistent failures (for example, reconcile errors lasting more than 3 days) and avoids duplicates.
- 4Slack notifications with the diagnosis and next steps.
- 5Containerized and deployed on Kubernetes with Helm.
03 Impact
Less manual triage, fewer duplicate tickets, and faster time to a first diagnosis for platform and customer-project failures.