Hands-on architecture / checked September 2026

Run the layers on your laptop

This small example uses local disk as storage, Parquet as the file format, Apache Iceberg as the table format, SQLite as the catalog backend, and PyIceberg to write and read data. A named net-revenue calculation stands in for the semantic layer. It does not simulate a distributed query engine or an agent.

Prerequisites and versions

Python 3.10–3.14, a terminal, and an internet connection for the first install. The example pins PyIceberg 0.12.0 and PyArrow 25.0.1. Run it from the repository root:

python3 -m venv .venv
. .venv/bin/activate
pip install -r examples/local-iceberg/requirements.txt
python examples/local-iceberg/run.py

Expected first line: rows=3 net_revenue=210. The second line contains the Iceberg snapshot ID and varies per run. The example uses a temporary directory, so it leaves no table behind.

What each step proves

  1. Catalog: SQLite records a namespace and Iceberg table metadata pointer.
  2. Write: an Arrow batch becomes Iceberg-managed Parquet data and a new snapshot.
  3. Read: a fresh catalog lookup scans the committed snapshot.
  4. Metric: paid amounts (230) less refunds (20) yield net revenue (210).

Change a row and run again to see how the metric follows data. For a real platform, replace local storage and SQLite with durable services, put access policy at the catalog and query layers, and publish the metric definition to every consumer.

Troubleshooting

If installation cannot find a wheel, check the Python version and architecture. If the script reports a catalog or file error, run it from a writable directory and use a fresh virtual environment. The full source is in the repository; the library behavior follows the Apache PyIceberg documentation.