Here is another example: https://datastudio.google.com/u/0/reporting/6a2c38d4-3a22-41... - startup times for serverless engines. Note: DuckDB is not present in this comparison, I'm not sure why.
Both comparisons are independent.
I know a few cases when clickhouse-local is worse than DuckDB:
- it does not use the embedded metadata to filter while processing Parquet files;
- the syntax for accessing the files in s3 is clunkier in clickhouse-local;
- finally, there is no Python module and integration with dataframes.
I know many cases when clickhouse-local is better than DuckDB. Performance is mostly better, because ClickHouse is more advanced in the query engine. DuckDB mostly keeping up, but sometimes is already ahead, in some scenarios. Query language, data types support, feature completeness, stability and testing - much better in ClickHouse.
I did not make uncharitable comments about DuckDB. We have recently met with Hannes Mühleisen and the team from DuckDB labs, and I have very good impression about the technology and the team. I see them as our friends. I'm also enthusiastic about every data processing technologies.
(co-founder at MotherDuck)
When I tried to use DuckDB on the same dataset as ClickHouse, it simply did not work due to OOM: https://github.com/duckdb/duckdb/issues/3969
I also told them about our experience of using various memory allocators, and why you should never use the GLibC's malloc.
This issue was fixed.
I did not catch you address correctness?
FWIW I have some experience with Clickhouse as I ran product at Firebolt and played a critical role in being more transparent about their foundations and giving credit where it is due. However, I do have some first hand experience with Clickhouse, which was jarring, considering my previous experience was with BigQuery.
I tried to play with "serverless ELT" (the other kind of serverless) where I would define AWS Lambda that turns incoming CSV files into Parquet for archiving and querying.
It seems that in DuckDB, the amount of memory it needs to do that is always at least a few gigs, and the memory goes up proportionally with the file size (which I think is due to data type inference?), which is both expensive and annoying because you are either overpaying or you need to go and increase a size of your Lambda when things start crashing.
I wonder if ClickHouse-local or some other tool can do that with constant memory, no matter the file size. I know Spark can, but Spark is kind of pain to work with.
(Yes, I do realize this is a bizarre use case.)
But it did not work (the server became unresponsive after consuming all memory).
> How confident are you that your chosen dataset is neutral?
I have no idea if it is "neutral", I picked it randomly.
I test ClickHouse on every interesting dataset, see here: https://github.com/ClickHouse/ClickHouse/issues?q=is%3Aissue...
The reason - I love working with data :) If I see a dataset, I load it into ClickHouse - this is the first thing I do. This is not a kind of marketing or promotion of ClickHouse - you know, if it were some directed task, it would be uninteresting for me.
> Query language, data types support, feature completeness, stability and testing
nothing about correctness.
In terms of stability, I see a couple of pretty old, and still unresolved issues about memory safety (data races, segmentation faults) in your repository, found by users.
In contrast, most of the memory safety issues in ClickHouse are found by continuous fuzzing before the release. And finding similar issues will give you a reward: https://github.com/ClickHouse/ClickHouse/issues/38986
Our testing system successfully finding issues in well known and widely used libraries - jemalloc, rocksdb, grpc, AWS, Arrow, Avro, ZooKeeper, Linux kernel... It is kind of surprising, and it makes an impression like we are the only product that does testing for real.
I also remember an example of using SQLancer from 1.5 years ago. When SQLancer appeared, we started to use it on ClickHouse, and it has found a few issues and one crash. At the same time, it has found a lot of crashes in DuckDB. But this example is very old, and DuckDB evolved a lot since then - it is a much younger technology after all.
I felt this too [0], but I guess it is mostly down to the fact that perhaps English may not be their first or even second language.