So what does this mean for ingestion (and indexing)? Iceberg V3? Paimon? Bespoke ingestion through the DB engine to do the indexing?
So what does this mean for ingestion (and indexing)? Iceberg V3? Paimon? Bespoke ingestion through the DB engine to do the indexing?
I assume the native formats will always be faster / more optimized but the need for Trino as a separate executor while running either of these databases seems to be close to gone.
Native format is faster (especially for colocated joins), but it's way more expensive if you have to run a bunch of separate storage nodes vs just using S3, especially your query volume isn't that high.
I liken it to the BigQuery cost model, where storage is effectively free.
Actually query on Iceberg is end-to-end faster on ClickHouse than MergeTree native format on disk. It's mostly a matter of how much compute you throw at it. Native MergeTree still wins in cases that depend on use of indexes to reduce I/O but scan speed is no longer an issue.
You do not even need to change the Iceberg spec, although you could. Just rejig things a bit to produce better statistics (think liquid clustering but on the write path). You still write the same dataset queried the same way, but automatically get better query performance. Workload owners should be able to decide the tier at which this streaming reorganization happens: they know their data staging and bandwidth constraints the best.