From Kafka to Iceberg: High-Throughput Ingestion at Starburst | Trino Meetup
Real-time analytics is critical for modern businesses — but bridging the gap between fast-moving #kafka streams and query-ready #lakehouse tables remains a complex challenge. At Starburst, we encountered this firsthand while ingesting our internal telemetry data into #iceberg tables. Existing solutions fell short, plagued by issues such as lack of exactly-once guarantees, limited scalability, head-of-line blocking, small file proliferation, and high operational overhead and cost. This prompted us to rethink streaming ingestion from the ground up. Our goal was a fully managed, highly available system that makes data actionable within minutes, optimized for both performance and usability. We designed and built a custom Kafka-to-Iceberg ingestion service that directly writes data in Iceberg format with strong guarantees, minimal latency, and continuous data maintenance for optimal query performance. Along the way, we developed novel techniques — such as Iceberg-aware commit coordination and adaptive #kafka consumer assignment — to overcome typical #ingestion bottlenecks and deliver best-in-class price/performance. Speaker: Lakshmikant (Pachu) Shrinivas, Staff Software Engineer at Starburst, hosted by Lester Martin, Developer Advocate at Starburst. ------ This recording was made during the 2026-02-18 live meetup described at https://www.meetup.com/trino-americas/events/312487299/