Posts
All the articles I've posted.
- 8 MIN READ•Oct 14, 2025
The State of Apache Iceberg v4 - October 2025 Edition
What's Coming in Apache Iceberg v4: A Deep Dive into the Future of Open Table Formats
Data LakehouseData EngineeringApache Iceberg - 21 MIN READ•Sep 24, 2025
The Ultimate Guide to Open Table Formats - Iceberg, Delta Lake, Hudi, Paimon, and DuckLake
Understanding Iceberg, Delta Lake, Hudi, Paimon, and DuckLake
Data LakehouseData EngineeringAgentic AI - 42 MIN READ•Sep 23, 2025
The 2025 & 2026 Ultimate Guide to the Data Lakehouse and the Data Lakehouse Ecosystem
What is the Data Lakehouse and the Data Lakehouse Ecosystem? This comprehensive guide covers everything you need to know about the Data Lakehouse architecture, open table formats like Apache Iceberg, Delta Lake, Apache Hudi, and Apache Paimon, and the modern data ecosystem that supports them.
Data LakehouseData EngineeringAgentic AI - 3 MIN READ•Sep 17, 2025
Composable Analytics with Agents - Leveraging Virtual Datasets and the Semantic Layer
The power of Dremio's Semantic Layer for Agentic AI
Data LakehouseData EngineeringAgentic AI - 4 MIN READ•Sep 16, 2025
The Endgame – Building an Autonomous Optimization Pipeline for Apache Iceberg
Learn how to automate compaction, snapshot expiration, and layout optimization in Apache Iceberg using metadata-driven triggers and orchestration tools for a self-healing lakehouse.
Apache IcebergAutomationCompaction - 3 MIN READ•Sep 9, 2025
Managing Large-Scale Optimizations – Parallelism, Checkpointing, and Fail Recovery
Learn how to scale Apache Iceberg table optimizations across large datasets using parallelism, checkpointing, and fail recovery to ensure reliability and performance.
Apache IcebergOptimizationCompaction - 9 MIN READ•Sep 5, 2025
Unlocking the Power of Agentic AI with Apache Iceberg and Dremio
Unlocking the Power of Agentic AI with Apache Iceberg and Dremio
Data LakehouseData EngineeringAgentic AI - 4 MIN READ•Sep 2, 2025
Hidden Pitfalls – Compaction and Partition Evolution in Apache Iceberg
Partition evolution in Apache Iceberg is a powerful feature, but if not managed carefully, it can introduce fragmentation and impact compaction performance. Learn how to handle it effectively.
Apache IcebergPartition EvolutionTable Optimization - 4 MIN READ•Aug 26, 2025
Using Iceberg Metadata Tables to Determine When Compaction Is Needed
Discover how to use Apache Iceberg's metadata tables to proactively detect small files, bloated manifests, and table fragmentation - so you can trigger compaction only when it's needed.
Apache IcebergMetadata TablesTable Optimization - 4 MIN READ•Aug 19, 2025
Designing the Ideal Cadence for Compaction and Snapshot Expiration
Learn how to design an effective schedule for compaction and snapshot expiration in Apache Iceberg to balance cost, performance, and data freshness.
Apache IcebergTable OptimizationCompaction - 4 MIN READ•Aug 12, 2025
Avoiding Metadata Bloat with Snapshot Expiration and Rewriting Manifests
Learn how to prevent and clean up metadata bloat in Apache Iceberg by expiring snapshots and rewriting manifests for better performance and manageability.
Apache IcebergMetadata OptimizationSnapshot Expiration - 4 MIN READ•Aug 5, 2025
Smarter Data Layout – Sorting and Clustering Iceberg Tables
Improve query performance in Apache Iceberg by organizing your data layout with sorting and Z-order clustering. Learn how to reduce scan cost and improve filter effectiveness.
Apache IcebergClusteringSorting