DuckDB 0.9
duckdb.org
duckdb.org
The recent support in DuckDB for "AS OF" has really helped our performance when doing point-in-time joins across tables.
https://www.hopsworks.ai/post/python-centric-feature-service...
[0]: https://www.birdspot.co.uk/identifying-birds/types-of-ducks
It has been around since 2016, and it covers and extends the feature set of DuckDB with a huge margin. Worth noting that it never has breaking changes in its table format MergeTree.
I'm tracking the progress of DuckDB and see that it is modeled after ClickHouse, but does not approach it in terms of feature completeness, stability, or performance.
The closest to DuckDB option is to use its self-contained version, clickhouse-local: https://clickhouse.com/blog/extracting-converting-querying-l... or an embedded version, chdb: https://github.com/chdb-io/chdb
(Head of Produck and Co-founder at MotherDuck)
Next up adding a caching layer in.
Anyone know if duckdb is considering adding some sort of caching mechanism? i.e. determining which queries or intermediate queries to store in-memory?
Would be interesting to write this as an extension
While we're running vanilla duckdb, our extension does sopen heart surgery on some of the components, to make it suitable for a multi tenant serverless cloud service.
(Co-founder and head of produck at MotherDuck)
A friend has a use case involving searching structured data (like everyone). I think it's basically reports on different companies, so have fields like name, location, market cap, etc. They come as JSON documents, with very similar structures, which change slowly over time as fields are added and removed. He needs to give users an interface to these, where they can filter on various fields ("show me all small hairdressers in Minnesota" etc).
He currently dumps the JSON documents into ElasticSearch, and queries that. It works, but my gut feeling is that ElasticSearch is wildly less efficient than the theoretical optimum for this. Could DuckDB be an alternative?
The easiest thing would be if he could put whole JSON documents into a JSON column. Less easy would be breaking them up into real columns; i know that DuckDB is able to do this automatically, but it feels like a brittle thing to depend on.
I did some playing around with searching data from JSON, comparing a single JSON column to automatic 'proper' columns, and the proper columns were massively faster. Is there scope for making queries on JSON much faster, using indexes, or some setting i missed, or something?
Is DuckDB even likely to be a good solution for this? It's intended for analytics, whereas this is more of a retrieval / search / filtering use.
If you are compiling applications with clang (or zig), some (all?) extensions will not work, because C++ does not have a stable ABI. And thus an application linked with lld/zig cc will fail to load an extension linked with ld/gcc (aka the official duckdb extensions).
We found this the hard way after enabling zig-c++ as the C++ compiler in prod. Took some time to recompile (and statically link) the extensions we actually use, but now the system is much better off.
I would prefer a default that does not download anything from the internet (especially .so files), but that's quite easy to change when you know it.
Still wip to actually poke with duckdb for myself. :)
DuckDB also has potential for IoT/embedded use cases as a data buffer with its compression and efficient reads, instead of just sqlite. but for that it needs to free up disk after expiring data
I think I have just found out yet another great DuckDB-weekend project!
Dale from ClickHouse wrote a pretty extensive blog series on the Hacker News dataset, ingest approach, some queries of interest...
Could be a good bit of reading alongside the project?
https://clickhouse.com/blog/getting-data-into-clickhouse-par...
Go to https://shell.duckdb.org, and type
FROM 'https://hacker-news.firebaseio.com/v0/item/37663308.json';
Querying hacker news, from a browser tab (passing through a bunch of database and Web technology that make it possible for DuckDB to be executed within a browser tab)
[0] https://duckdb.org/docs/archive/0.8.1/api/python/relational_... [1] https://github.com/duckdb/duckdb/pull/8083
Storage is naturally distributed due to underlying IaaS.
Rather, that seems to be an implementation detail. What use cases are relevant is the key. Are we chasing big Apple-sized workloads? No. [1]
That said, scaling-up on EC2, especially with our one-instance-per-user architecture, fits vast majority of workloads we've seen in our past lives building Exabyte-scale systems.
[0] https://motherduck.com/blog/the-simple-joys-of-scaling-up/ [1] https://motherduck.com/blog/big-data-is-dead/
it is a scope, not inability. Building single node system is much easier than distributed for fast complex algos, meaning team can ship more features.
> boring open-source technologies, such as ClickHouse.
to me ClickHouse is not boring, it is fragile and failing all the time with OOMs in various places.
so, how did you scale to exabyte with one server?..
In ClickHouse, we approach this with continuous fuzzing. When we tried to integrate duckdb as one of the storage engines into ClickHouse, our CI system immediately found an uninitialized memory read inside it: https://github.com/duckdb/duckdb/issues/7433
@wenc, the op might have meant the native storage format of DuckDB, which does not use Parquet.
| Title Snippet | YYYY-MM | comments | HN URL |
| ------------------------------- | ------- | --------:| --------------------------------------------- |
| "DuckDB 0.8.0" | 2023-05 | 0 | https://news.ycombinator.com/item?id=35974719 |
| "DuckDB for Swift" | 2023-05 | 0 | https://news.ycombinator.com/item?id=35711430 |
| "Calculate the Digits of Pi..." | 2023-03 | 1 | https://news.ycombinator.com/item?id=35153824 |
| "...in-process OLAP DBMS " | 2023-02 | 103 | https://news.ycombinator.com/item?id=34741195 |
| "Modern Data Stack in a Box..." | 2022-10 | 5 | https://news.ycombinator.com/item?id=33191938 |
| "Querying Postgres Tables..." | 2022-09 | 39 | https://news.ycombinator.com/item?id=33035803 |
| "...embeddable SQL database..." | 2020-09 | 160 | https://news.ycombinator.com/item?id=24531085 |
| "DuckDB: SQLite for Analytics" | 2020-05 | 67 | https://news.ycombinator.com/item?id=23287278 |