What Happens When A Database Gets Popular? (with Hannes Mühleisen)

Developer Voices

Developer Voices

24,155 views

Building something people actually want is supposed to be the happy ending. But it arrives with a bill attached: feature requests you didn't ask for, pull requests you'd rather not maintain forever, users demanding the one thing you swore you'd never build, and — if you're unlucky — a company where the sales team quietly starts deciding what engineering works on. DuckDB has spent the last two and a half years working through that list. So how do you stay a database engineering team when success keeps trying to turn you into something else?

Hannes Mühleisen, co-creator of DuckDB, is back to talk through the answers they've landed on. Their fix for unwanted pull requests was an extension mechanism, which then forced them to make every part of the engine pluggable — including the parser, which meant ripping out 20,000 lines of Postgres' yacc grammar and rewriting SQL parsing on top of PEG. Their fix for the users demanding client-server was Quack, a protocol designed by people who'd already published a paper on why every existing database wire protocol is wrong. And their answer to Apache Iceberg, after three years of implementing it, was DuckLake: throw out the Avro-and-JSON metadata files and keep the metadata in a database, on the grounds that the Iceberg REST catalog has a Postgres in it anyway.

Which brings us to the news Hannes breaks in this episode: DuckDB Labs is being acquired by AWS, while the DuckDB Foundation, the project and its licence stay where they are. There's the question of why a profitable, self-funded, 30-person company in Amsterdam would take that deal, what commitments you write into the contracts when you're worried today's promises might outlive today's management, and what it's actually like to have a boss again after five years without one. If you're curious how an open source project keeps its technical soul once the enterprise arrives — or you just want to know why parsing SQL is harder than parsing almost anything else — Hannes has some good answers.

---

Support Developer Voices on Patreon: Patreon: DeveloperVoices
Support Developer Voices on YouTube: @developervoices

Our previous episode with Hannes: DuckDB: How to Build 100x Faster Analytics...

DuckDB: https://duckdb.org/
DuckDB Foundation: https://duckdb.foundation/
DuckLabs (formerly DuckDB Labs): https://ducklabs.com/
DuckLake: https://ducklake.select/
Quack (DuckDB's client-server protocol): https://duckdb.org/quack/
DuckDB v2.0: Your Database Deserves a Better Parser: https://duckdb.org/2026/08/20/duckdb-...

Runtime-Extensible Parsers (CIDR 2025 paper): https://duckdb.org/pdf/CIDR2025-muehl...
Don't Hold My Data Hostage (VLDB 2017 paper): https://www.vldb.org/pvldb/vol10/p102...

cpp-peglib: https://github.com/yhirose/cpp-peglib
GNU Bison: https://www.gnu.org/software/bison/
PEP 617 – New PEG parser for CPython: https://peps.python.org/pep-0617/
PRQL: https://prql-lang.org/

Apache Iceberg: https://iceberg.apache.org/
Apache Parquet: https://parquet.apache.org/
Apache Avro: https://avro.apache.org/
Apache Thrift: https://thrift.apache.org/
Protocol Buffers: https://protobuf.dev/
Apache Arrow Flight SQL: https://arrow.apache.org/docs/format/...
Amazon S3 Tables: https://aws.amazon.com/s3/features/ta...

SQLite: https://www.sqlite.org/
PostgreSQL: https://www.postgresql.org/
pandas: https://pandas.pydata.org/
Apache Spark: https://spark.apache.org/
Snowflake: https://www.snowflake.com/
Databricks: https://www.databricks.com/

CWI (Centrum Wiskunde & Informatica): https://www.cwi.nl/en/
DuckCon #7, Amsterdam: https://duckdb.org/events/2026/06/24/...

Kris on Bluesky: https://bsky.app/profile/krisajenkins...
Kris on Mastodon: http://mastodon.social/@krisajenkins
Kris on LinkedIn: LinkedIn: krisjenkins

---

0:00 Intro
4:42 What Is DuckDB, And Why Build One From Scratch?
8:41 Making A Database Both Fast And User-Friendly
15:05 Why Single-Node Beats Distributed
20:30 Quack, And Why Database Wire Protocols Are Terrible
30:32 The Hidden Costs Of Going Client-Server
37:33 How DuckDB's Extension Mechanism Works
42:40 Why SQL Is The Final Boss Of Parsing
50:14 Rewriting DuckDB's Parser With PEG
57:02 Iceberg, And Why DuckDB Had To Support It
1:06:52 The Trouble With Avro
1:10:05 DuckLake: Put The Metadata In A Database
1:15:29 From Laptop Tool To Enterprise Data Stack
1:23:43 The Big News: DuckDB Labs Joins AWS
1:32:31 The Missing Feedback Loop Of Real Workloads
1:38:43 What It's Like Having A Boss Again
1:41:44 Triggers, Stored Procedures And DuckDB 2.0
1:45:02 Outro