Streaming· Case study 04
Real-time AdTech bidstream pipelines
Streaming and batch pipelines that ingest high-volume programmatic-advertising data (bid requests, wins, impressions, and auctions) into Apache Druid and ClickHouse for real-time analytics.
GBs/hr
of AdTech events into real-time dashboards
Architecture at a glance
- Scala intake
- Kafka
- Beam on Dataflow
- Druid / ClickHouse
01 The problem
AdTech platforms produce huge, bursty event streams across several cloud regions, and customers need them queryable within minutes, correct down to the deal ID.
02 What I built
- 1HTTP intake services in Scala that accept event data and publish to Apache Kafka.
- 2Streaming and batch pipelines on Apache Beam (Scio) running on Google Cloud Dataflow that parse OpenRTB-style payloads (impressions, PMP deals, creatives), then enrich and aggregate them.
- 3Batch ingestion from multi-region AWS S3 and GCS buckets, orchestrated by a central Apache Airflow service on Kubernetes.
- 4Store monitoring that tracks landed files and partitions, so late or missing data is detected automatically.
- 5Idempotent backfills and replay for failed windows, plus source-to-field lineage maps used to answer customer data questions quickly.
03 Impact
GBs per hour of AdTech events land in real-time dashboards, with a clear path from any dashboard number back to the raw field that produced it. The inherited legacy pipelines were stabilized and their deployment modernized along the way.