Query the data where it lives
Copying data into a warehouse to join it is a habit born of necessity,
not design.
Trino queries sources in place. Postgres, S3, Kafka, ClickHouse - one SQL
dialect across all of them, no pipeline in between:
SELECT p.name, c.revenue
FROM postgres.public.customers p
JOIN clickhouse.default.sales c ON p.id = c.customer_idNo ETL ran. Nothing was copied. The join just happened across two systems that
have never met.
This removes an entire category of pipeline - and the failures that came with
it, like the sync job that silently stopped three weeks ago.
The trade is that federated queries are only as fast as the slowest source, so
it is not a replacement for a warehouse. It is a replacement for copying data
you did not need to copy.
What are you federating, and what turned out to be too slow to federate?
