What is Trino? Trino Explained for Beginners!

The Data and AI Guy

The Data and AI Guy

1,901 views

One SQL statement that joins a Postgres table to a folder of Parquet files in object storage, with no pipeline, no copy, and no warehouse holding both. The engine answering it owns none of the data. That engine is Trino.

This is a beginner explainer for Trino: a distributed SQL query engine with no storage of its own. I cover why Facebook built the original Presto in 2012 (every cross-system question carried a two-week pipeline with it), how a query is actually answered by a coordinator, workers, and connectors registered as catalogs, and why file-based sources need a Hive Metastore, Glue, or Iceberg REST catalog before Trino can see a single table. Then the part that decides whether your queries are fast or dragging whole tables over the wire: predicate pushdown, and how to use EXPLAIN to check whether your filter survived the trip into the connector. Finally the real limit, which is memory: Trino streams through RAM and fails queries that don't fit where Spark would spill to disk, so I close with when to reach for Trino vs Spark vs Snowflake and BigQuery, the Presto fork history, and where you've already used Trino without knowing it (Starburst, Athena, lakehouse SQL over Iceberg).

⏱️ Chapters
0:00 - One query across Postgres and a data lake
0:39 - What this video covers
0:55 - The definition: a SQL engine with no storage
1:32 - Why it exists: Facebook, Presto, and the pipeline tax
2:28 - Coordinator and workers
3:01 - Connectors and catalogs
3:36 - Why file sources need a metastore or catalog
4:17 - Trino holds no state you'd miss
4:40 - Predicate pushdown: the thing that makes it fast
5:23 - When pushdown fails and how EXPLAIN tells you
6:16 - The real limit: memory and failed queries
7:32 - Trino Vs. Spark: who is waiting?
8:25 - Trino Vs. Snowflake and BigQuery
8:57 - Trino Vs. Presto: the fork
9:28 - Where you've already seen Trino: Starburst, Athena, Iceberg
10:08 - Recap

🔗 Links
Trino: https://trino.io
Trino connectors: https://trino.io/docs/current/connect...
Apache Iceberg: https://iceberg.apache.org
Starburst: https://www.starburst.io
Amazon Athena: https://aws.amazon.com/athena/

#Trino #Presto #distributedSQL #queryengine #dataengineering #datalake #lakehouse #ApacheIceberg #Starburst #AmazonAthena #ApacheSpark #Postgres #Parquet #predicatepushdown #federatedquery