Apache Iceberg Open-source Lakehouse with Postgres Managed Catalog
Details
Hello friends,
Our next Dallas–Fort Worth PostgreSQL Meetup is on Tuesday, Sept 08.
Event presenters: Van Dorsey, Product Owner
Abstract: An open-source data lakehouse running on a single on-prem container —
Apache Iceberg on SeaweedFS object storage, written by Trino, read by
DuckDB — ingesting a CDC feed of NYC taxi data unattended for weeks, now
past half a billion rows. The piece that makes it a database rather than
a folder of Parquet files is the catalog, and the catalog is Postgres:
S3 gives you an atomic single-object write and no multi-object
transaction, so something has to swap one pointer transactionally. Van will
show it running live, including what broke getting there and the
measured answer to "so why not just use Postgres?"
We'll have pizza, soft drinks, and social from 6:30 PM to 7:00 PM. The session will start at 7:00 PM.
Our generous sponsor, Improving, provides the place and pizza at every meeting. Thank you to everyone who RSVPs; this helps ensure we order enough food and reduce waste.
See you there!


