Observability· Case study 08
Fleet observability, data-freshness alerting & usage billing
The observability layer for a 30+ cluster, 2+ PB ClickHouse and Druid fleet: templated monitoring dashboards, data-freshness alerting, SLOs, and a usage-billing pipeline that shows customers exactly what they consume.
30+
clusters monitored from one values file
Architecture at a glance
- values.yaml
- Jinja templates
- Rill dashboards
- Alerts + billing
01 The problem
With dozens of clusters and customer pipelines, problems surfaced when customers noticed stale dashboards, and resource usage wasn’t visible enough to manage cost.
02 What I built
- 1Jinja-templated monitoring projects that generate connectors, models, metrics views, and dashboards for every ClickHouse cluster from one values file, covering query latency, query_log with projection usage, parts, merges, disk, and memory.
- 2A template-driven alert generator that produces lag and gap alerts for every customer pipeline from a single source-of-truth config, regenerated in CI.
- 3SLOs for availability and data freshness, burn-rate alerts, runbooks, and postmortems.
- 4An internal usage-billing pipeline that breaks resource usage down per customer.
- 5Customer-facing usage and product-telemetry dashboards.
03 Impact
Issues get caught before customers see them, adding a new cluster or pipeline to monitoring is a config change, and cost and usage are visible to both the business and customers.