7:50 · Music video
The Iceberg Open Lakehouse
A beat-synced explainer about the open table format and the interoperable lakehouse around it.
Open videoOpenDataLakehouse
A vendor-neutral reference.
A data architecture that stores data in open file and table formatson commodity object storage, so that any compliant engine can read and write it without vendor lock-in.
The Iceberg Open Lakehouse
An eight-minute visual tour of Apache Iceberg, open storage, catalogs, engines, and the architecture that keeps your data portable.
7:50 · Music video
A beat-synced explainer about the open table format and the interoperable lakehouse around it.
Open videoAn open lakehouse is assembled from layers, each with one job and several implementations. Any one of them can be replaced without disturbing the others, which is the entire point of the architecture.
Four pages that cover the whole subject, in the order they make sense.
Each layer has several viable options. These comparisons lead with the table.
201 in-depth entries, each mapped to the layer it belongs to, plus a one-line glossary for when that is all you need.
Short answers, each with a link to the page that covers it in full.
A data architecture that keeps data in open file and table formats on object storage you control, so any engine that implements those formats can read and write it. You attach engines to the data instead of loading the data into an engine.
Read the full definitionA data lake gives you cheap storage without guarantees, and a warehouse gives you guarantees without engine choice. A lakehouse adds transactions and schema enforcement to files on object storage, so you keep the low-cost storage and the choice of engines while gaining the guarantees.
See the side-by-side comparisonThe metadata layer that turns a directory of data files into a table. It records which files make up the table right now, what the schema is, and how concurrent writers commit changes, which is what gives you transactions, schema evolution, and time travel. Apache Iceberg, Delta Lake, Apache Hudi, and Apache Paimon are table formats.
Read the table format entryIt tracks which tables exist, where each table's current metadata lives, and who is allowed to read or write them. Once more than one engine touches the same tables, the catalog is the single authority they all ask.
Compare lakehouse catalogsYes, and that is the main reason to build one. Spark, Trino, Flink, DuckDB, Dremio, and others can read and write the same tables as long as they implement the table format and connect through a shared catalog. There is still one copy of the data.
Read the five principlesThe architecture is, when the formats are publicly specified, the catalog speaks an open API, and leaving costs no more than arriving. Individual products can still lock you in, so this site gives a test for each of those properties rather than taking a product's word for it.
See the tests for opennessRead the definition, then the reference architecture to see how the layers fit. If you learn best by doing, run the local example: it writes rows to an Iceberg table with PyIceberg and reads a snapshot back on your laptop.
Run the local exampleWritten by vendors, listed here because the technical content is useful. Nothing on this site is a recommendation to buy anything.
More reading, including independent sources, is collected in theblog roll, video roll, andbook roll.