OpenDataLakehouse

What is an open lakehouse?

A vendor-neutral reference.

A data architecture that stores data in open file and table formatson commodity object storage, so that any compliant engine can read and write it without vendor lock-in.

Reference entries
201
Architecture layers
9
Comparisons
3
SparkTrinoFlinkDuckDBSAP/DremioCatalogTable formatData fileData fileData fileData fileObject storage you control

The stack, one layer at a time

An open lakehouse is assembled from layers, each with one job and several implementations. Any one of them can be replaced without disturbing the others, which is the entire point of the architecture.

  1. Agents & AI18 entriesHow do agents use any of this safely?Agentic analyticsRAGText-to-SQLVector search
  2. Semantics12 entriesWhat does this data mean?Semantic layerMetricsOntologyModeling
  3. Movement15 entriesHow does data get in, and stay current?AirflowdbtDagsterCDCStreaming
  4. Interchange3 entriesHow does data move between systems?Apache ArrowArrow FlightFlight SQL
  5. Compute43 entriesWhat actually runs the query?SparkTrinoFlinkDuckDBDremio
  6. Catalog17 entriesHow do engines find and govern those tables?Apache PolarisIceberg RESTNessieGlueUnity
  7. Table format50 entriesWhat makes those files behave like a table?Apache IcebergDelta LakeApache HudiApache Paimon
  8. Storage format26 entriesHow is the data physically written down?Apache ParquetORCAvroObject storage
  9. Foundations17 entriesWhat kind of system are we building?LakehouseData lakeData warehouseMedallion

Start here

Four pages that cover the whole subject, in the order they make sense.

The choices, side by side

Each layer has several viable options. These comparisons lead with the table.

All comparisons

Look something up

201 in-depth entries, each mapped to the layer it belongs to, plus a one-line glossary for when that is all you need.

Browse the library

Questions newcomers ask

Short answers, each with a link to the page that covers it in full.

Full FAQ

What is an open lakehouse?

A data architecture that keeps data in open file and table formats on object storage you control, so any engine that implements those formats can read and write it. You attach engines to the data instead of loading the data into an engine.

Read the full definition

How is it different from a data warehouse or a data lake?

A data lake gives you cheap storage without guarantees, and a warehouse gives you guarantees without engine choice. A lakehouse adds transactions and schema enforcement to files on object storage, so you keep the low-cost storage and the choice of engines while gaining the guarantees.

See the side-by-side comparison

What is a table format?

The metadata layer that turns a directory of data files into a table. It records which files make up the table right now, what the schema is, and how concurrent writers commit changes, which is what gives you transactions, schema evolution, and time travel. Apache Iceberg, Delta Lake, Apache Hudi, and Apache Paimon are table formats.

Read the table format entry

What does a catalog do?

It tracks which tables exist, where each table's current metadata lives, and who is allowed to read or write them. Once more than one engine touches the same tables, the catalog is the single authority they all ask.

Compare lakehouse catalogs

Can multiple engines share the same tables?

Yes, and that is the main reason to build one. Spark, Trino, Flink, DuckDB, Dremio, and others can read and write the same tables as long as they implement the table format and connect through a shared catalog. There is still one copy of the data.

Read the five principles

Is an open lakehouse vendor neutral?

The architecture is, when the formats are publicly specified, the catalog speaks an open API, and leaving costs no more than arriving. Individual products can still lock you in, so this site gives a test for each of those properties rather than taking a product's word for it.

See the tests for openness

Where should I start?

Read the definition, then the reference architecture to see how the layers fit. If you learn best by doing, run the local example: it writes rows to an Iceberg table with PyIceberg and reads a snapshot back on your laptop.

Run the local example

Implementation reading

Written by vendors, listed here because the technical content is useful. Nothing on this site is a recommendation to buy anything.

More reading, including independent sources, is collected in theblog roll, video roll, andbook roll.