Thoughts on OLAPs ClickHouse vs. Apache Druid vs. Starrocks in 2023/2024
The ideal solution would be a db that can make use of an S3 bucket as long term storage, while being easy to deploy and manage on Kubernetes. They will be consuming kafka topics.
I have narrowed down the solutions between Apache Druid, Starrocks and of course Clickhouse.
While Druid shows some very neat tricks on real time data storing/indexing, its deployment on kubernetes is really ugly.
Starrocks is not so far, also a bit complex to deploy and manage the "front end and back end" clusters, while Clickhouse makes the k8s deployment a bliss.
What got me in Clikchouse is the kafka topics ingestion, its a lot of work to create and keep managing sql materialized things everytime I need to get some new data. Druid otherwise makes it super cool by almost automatically understanding the json received from the kafka topics.
However I feel like Druid is getting a bit old and with not so much community or development resources or attention. ChatGPT is saving my ass I must admit. Tho, a lot of "what corps are using this tech" sites show that Twitter, Neflix and Reddit itself use Druid. I'm not sure if that is true or old data.
Clickhouse otherwise feels like the industry baby, ultra-seeded and megacorp friendly, but with its own problems like painful cross table joins. Starrocks have a nice fanbase, but falls a bit in the same place as Druid, even being some years ahead of Apache Doris, plus the fact it is mantained by the Linux Foundation.
So, what are your thoughts on these three OLAP databases today?