Near Real-Time Ingestion for Trino
starburst.io
starburst.io
> Here, the continuously running Flink job parses the JSON data source directly following the Iceberg table schema definition. This means future schema changes can be facilitated through just one source of truth.
This is the most interesting point. I'd love to take this implementation out for a spin - could solve a lot of pain I know other people deal with too.
Most of the pain (or even grunt-work) of managing data pipelines is updating, validating and managing schemas.
In the past the common approach people suggested was to have each application write the data with the same schema but in practice it's never possible unless it's a greenfield project and the services don't need to talk to other external or pre-existing systems. What ends up happening is that a translator (or validator) service comes up whose job it is to translate across the various schemas. 2x storage in Kafka, 2x compute for the consumers and probably 20x more maintenance and ways things can go wrong.
Maybe near real-time is just considering ingestion and query latency?
To be pedantic though Flink also operates on a flush interval so not "real time" by that definition.