Platform· Case study 01
Self-hosted ClickHouse platform on Kubernetes
A self-hosted, multi-tenant ClickHouse platform on Google Kubernetes Engine that runs 14 production ClickHouse clusters for enterprise real-time analytics, as part of a 2+ PB OLAP fleet.
14
production ClickHouse clusters
Architecture at a glance
- GitOps repo
- Helm + Terraform
- ClickHouse operator
- Sharded ClickHouse
01 The problem
Enterprise customers needed dedicated, low-latency ClickHouse clusters with predictable cost and strong isolation, without the price and limits of a managed service.
02 What I built
- 1Sharded, replicated ClickHouse clusters on GKE, run by the ClickHouse Kubernetes operator with ClickHouse Keeper coordination and persistent volumes.
- 2Distributed / _local table patterns so queries fan out across shards while writes and DDL stay replica-safe.
- 3Helm charts and GitOps repos where every cluster, user, profile, quota, and setting is declared in code and rolled out through CI/CD.
- 4Terraform for node pools, networking, service accounts, and storage, with separate staging and production environments.
- 5Per-cluster query memory limits, concurrency caps, and user profiles so one heavy dashboard can’t take down a shared cluster.
- 6Runbooks for rolling ClickHouse binary upgrades, operator upgrades, replica recovery, and scaling out shards.
03 Impact
Dedicated ClickHouse clusters for customers, delivered through a repeatable GitOps workflow instead of hand-built infrastructure, with TB-scale scans in under 30 seconds and ~30–40% lower storage and CPU after tuning.
Related case studies
- AIAI-assisted ClickHouse optimization with MCPMulti-TB single columns found in one audit pass
- MigrationWarehouse & lakehouse to ClickHouse, with reverse ETL and integrity checksEvery hop reconciled for drift, hourly
- PerformanceClickHouse performance & cost re-architecture with projections0 dashboard out-of-memory failures after rollout