Apache Iceberg at Sub-Second Speed: Architecting the Serving Layer
Is Apache Iceberg too slow to serve? 👉 See PhoenixAI: https://bit.ly/4r3Cwv1 Questions about your workload? 👉 https://bit.ly/3UAulu7 Apache Iceberg gives teams one open copy of their analytical data, but latency-sensitive workloads often get pushed into separate serving systems. Is Iceberg itself actually the bottleneck? PhoenixAI PM Sida Shen breaks down where time goes in an Iceberg query and how to achieve sub-second analytical serving without moving the data. Eric Sun, former Head of Data Platform + Datastores at Coinbase, brings the production perspective on ingestion, data modeling, joins, concurrency, and real-time updates. In this video: 🌟 What actually determines Iceberg query performance 🌟 The four stages of an Iceberg query: metadata, planning, scanning, and execution 🌟 Table layout, statistics, caching, and cost-based optimization 🌟 Why real-time updates remain difficult on Iceberg 🌟 When and how to add a real-time serving layer 🌟 Production lessons on ingestion, joins, and concurrency 🌟 Benchmarks and real-world examples from Demandbase and Conductor ---------------------------------------------------------------------------------------- Timestamps 00:00 Intro and Agenda 02:29 The Problem: Latency-Sensitive Workloads Still Move to Separate Systems 04:13 Where the Time Goes: The Table You Write, the Engine You Choose, the Format Itself 06:42 The Table: Partition, Sort, Compact, and Keep Your Statistics Fresh 10:46 The Engine: Why Data Lake Engines Are Slow at Serving 15:59 The Four Stages in the Life of an Apache Iceberg Query 17:03 Resolve Metadata — The First Iceberg-Specific Cost: Avro Manifests, ~1s per Core for 8 MB, Single-Threaded 18:41 Planning: Why a Cost-Based Optimizer Is Non-Negotiable with Joins; Runtime Filters as a Cost Decision 21:01 Scanning: Caching, Pruning, Late Materialization 22:42 Execution: Vectorized SIMD Operators, Pipeline Execution, Runtime Filters in Scan 25:06 The Format: Why Real-Time Updates Are the Hard Part on Iceberg 26:04 Copy-on-Write vs. Merge-on-Read — Neither Is a Real-Time Answer 27:04 The Pattern That Works: A Real-Time Layer Beside Iceberg 27:49 PhoenixAI: Metadata, Planning, Scan, Execution, Primary Key Table 31:05 Iceberg-Specific Optimization: Parallel Manifest Parsing, Cached Avro 32:17 Benchmarks: Trino, Databricks Delta, PhoenixAI on Iceberg 33:17 Accelerate Inside the Engine: Query Iceberg in Place; Materialize into Native Format; Preprocess the Hottest Few 35:34 The Real-Time Layer: Primary Key Table — Indexing Updates at Write Time, Not at Read 36:38 Eric Sun Takes Over — Nothing Is Free: End-to-End Latency Has Three Layers 38:11 Ingestion: Why 5-Minute Commit Intervals Are the Norm 39:22 Transformation: The Step Teams Skip, and What MPP Buys You 41:30 Do All Tables Belong in the Real-Time Layer? About 10% — One Mutable Table Sets the Whole Query's SLA 44:31 Dual Ingestion: The Tax for Keeping Iceberg as the Source of Truth 45:56 What "a Healthy Iceberg Table" Actually Means 47:15 Joins: The Industry Said You Didn't Need Them 49:32 The Concurrency Trap: 20 ms in Testing, 500 Concurrent in Prod 50:10 Materialized Views as Lightweight ETL, and a Recap of the Three Layers 53:04 Case Study: Demandbase — 49 ClickHouse Clusters to 1 54:13 Case Study: Conductor — 240M Rows in Hundreds of Milliseconds 56:19 Live Q&A: Tiering Hot Data Into the Real-Time Layer ---------------------------------------------------------------------------------------- Learn more at https://www.phoenixdata.ai/ Connect with us: LinkedIn: https://www.linkedin.com/company/phoenixai-data/ Twitter: https://x.com/phoenixdataai CelerData Website: https://www.phoenixdata.ai/ StarRocks GitHub: https://github.com/StarRocks/StarRocks StarRocks Website: https://www.starrocks.io/ Slack: https://starrocks.io/redirecting-to-slack #ApacheIceberg #DataEngineering #DataLakeAnalytics #DataAnalytics #OLAP #DataAnalyst #DataEngineer #DataInfrastructure #UserFacingAnalytics #Database #AnalyticalDatabase #DataLake #DataLakeHouse #Trino #Presto #DataWarehouse #DataScience