(co-founder at MotherDuck)
(co-founder at MotherDuck)
When I tried to use DuckDB on the same dataset as ClickHouse, it simply did not work due to OOM: https://github.com/duckdb/duckdb/issues/3969
I also told them about our experience of using various memory allocators, and why you should never use the GLibC's malloc.
This issue was fixed.
I did not catch you address correctness?
FWIW I have some experience with Clickhouse as I ran product at Firebolt and played a critical role in being more transparent about their foundations and giving credit where it is due. However, I do have some first hand experience with Clickhouse, which was jarring, considering my previous experience was with BigQuery.
I tried to play with "serverless ELT" (the other kind of serverless) where I would define AWS Lambda that turns incoming CSV files into Parquet for archiving and querying.
It seems that in DuckDB, the amount of memory it needs to do that is always at least a few gigs, and the memory goes up proportionally with the file size (which I think is due to data type inference?), which is both expensive and annoying because you are either overpaying or you need to go and increase a size of your Lambda when things start crashing.
I wonder if ClickHouse-local or some other tool can do that with constant memory, no matter the file size. I know Spark can, but Spark is kind of pain to work with.
(Yes, I do realize this is a bizarre use case.)
But it did not work (the server became unresponsive after consuming all memory).
> How confident are you that your chosen dataset is neutral?
I have no idea if it is "neutral", I picked it randomly.
I test ClickHouse on every interesting dataset, see here: https://github.com/ClickHouse/ClickHouse/issues?q=is%3Aissue...
The reason - I love working with data :) If I see a dataset, I load it into ClickHouse - this is the first thing I do. This is not a kind of marketing or promotion of ClickHouse - you know, if it were some directed task, it would be uninteresting for me.
> Query language, data types support, feature completeness, stability and testing
nothing about correctness.
In terms of stability, I see a couple of pretty old, and still unresolved issues about memory safety (data races, segmentation faults) in your repository, found by users.
In contrast, most of the memory safety issues in ClickHouse are found by continuous fuzzing before the release. And finding similar issues will give you a reward: https://github.com/ClickHouse/ClickHouse/issues/38986
Our testing system successfully finding issues in well known and widely used libraries - jemalloc, rocksdb, grpc, AWS, Arrow, Avro, ZooKeeper, Linux kernel... It is kind of surprising, and it makes an impression like we are the only product that does testing for real.
I also remember an example of using SQLancer from 1.5 years ago. When SQLancer appeared, we started to use it on ClickHouse, and it has found a few issues and one crash. At the same time, it has found a lot of crashes in DuckDB. But this example is very old, and DuckDB evolved a lot since then - it is a much younger technology after all.