Skip to content

AI· Case study 07

AI on-call triage agent

A production AI on-call agent that turns monitoring alerts into triaged, deduplicated tickets and Slack updates, so engineers start from a diagnosis instead of a raw alert.

Auto-triage

alerts classified, deduplicated, and ticketed

Architecture at a glance

  1. Monitoring webhook
  2. FastAPI
  3. LLM structured output
  4. Ticket + Slack

01 The problem

Alert noise and repeated failures ate on-call time. Every alert needed someone to open logs, classify the issue, and file or update a ticket.

02 What I built

  1. 1A FastAPI service with authenticated webhook endpoints for monitoring alerts.
  2. 2An LLM pipeline that reads logs and health status and returns structured-output classifications: issue category, likely cause, and severity.
  3. 3Ticket automation that creates or updates issues only for persistent failures (for example, reconcile errors lasting more than 3 days) and avoids duplicates.
  4. 4Slack notifications with the diagnosis and next steps.
  5. 5Containerized and deployed on Kubernetes with Helm.

03 Impact

Less manual triage, fewer duplicate tickets, and faster time to a first diagnosis for platform and customer-project failures.

08 · Contact

Let’s build something that holds up in production.

If something here resonates, whether it’s a data problem, an idea, or a conversation worth having, I’d love to hear from you. The fastest way to reach me is a message on LinkedIn.

Good reasons to reach out

  • Data platform help

    ClickHouse or Druid performance, Kubernetes operations, migrations, and cost reviews.

  • Collaboration

    Open-source work, writing, talks, or comparing notes on a hard data problem.

  • Just to talk shop

    Real-time analytics, AdTech data, MCP and agent tooling, or life as an FDE.

esc
↑ ↓ navigate↵ select