Bulk inserts on ClickHouse: How to avoid overstuffing your instance
runportcullis.co
runportcullis.co
It just buffers inserts in a queue and then does them all at once after a second has passed. This document does briefly mention it but it's such a nice feature that saves you from setting up a whole Kafka thing.
It is really nice in a pinch though.
However, during the last years I always find myself using https://vector.dev/ for all sort of tasks, including bulk inserts in ClickHouse.
At PostHog we were inserting something around 50M rows a day into CH with it and it was quite nice to be able to pause ingestion to a table by just detaching the table via SQL :D
I think they're still using the Kafka engine today but not sure. In our case (at least back then) we had to live with suboptimal batch sizes because we were providing near realtime analytics so Kafka was a solid fit.
- https://github.com/turbolytics/sql-flow/issues/100
- https://github.com/turbolytics/sql-flow/pull/116
The Clickhouse python API made this trivial thanks to supporting Arrow!
We have a tutorial explaining how to stream data from Kafka to Clickhouse using SQLFlow:
https://sql-flow.com/docs/tutorials/clickhouse-sink
Local benchmarks show 70,000 rows / second insert using a batch size of 50k!