Researchers, please replace SQLite with DuckDB now
dirk-petersen.medium.com
dirk-petersen.medium.com
Researchers, please work with that you know and works well for you.
No problems with SQLite.
I ended up putting the code in Debian because of other problmes I found but the decision was already made.
Don’t researchers know how to create database indexes? That would explain a lot of what I see in day to day pseudo-benchmarking like this.
(And yes, DuckDB is pretty cool. But trying to compare things without understanding how they should work is… sub-optimal)
For hosted foss services, write.as (https://write.as/) and bearblog (https://bearblog.dev/) are good. If self-hosting, the choices are infinite.
Hi. No thank you? I'm tired of greedy software that I have to literally pin down by shackling it to a cgroup via systemd or docker.
1 core is fine. We have a group of 20, and limited resources, and there are no fires.
Would be nice to have an option to limit the number of cores used by DuckDB though, if it doesn't have one already. (`duckdb -j4`?)
I'm just not a fan of software that immediately tries to use all possible resources for 2% gain in a performance benchmark over peers.
Is there some specific software you are referring to?
I agree that database software likely does not behave so aggressively, but the terms I saw in the duckdb post triggered me to believe they were trending in that direction.
(Disclaimer: I work at DuckDB Labs)
See, as I recall WebSQL was deprecated because to be a proper standard it needed more than one implementation, but nobody thought they could improve on Sqlite.
So the draft standard ended up saying user agents must implement the SQL dialect supported by Sqlite 3.6.19 which is.... not normal, for a standard.
But fundamentally, an argument in favour of Sqlite
A good resource to better understand how SQLite is tested is https://www.sqlite.org/testing.html. I found it to be a fascinating read, and quite impressive. I came away with a lot of confidence in SQLite's robustness.
I really don't like the tone of this article. Author, please change the tone to something more less aggressive.
The ones that irritate me are the ones with titles in the form of a direct order, such as "Stop using (technology)!".
That one makes me feel like starting to use (technology) just to demonstrate that I don't take orders from this person. :-)
I use sqlite for storing configs with data rather than dumping everything in one big YAML.
I might give duckdb a try in our next project
> not trivial to install
It's a single binary, there is nothing to install. If this isn't trivial to install I can't imagine installing anything else.
497MB is big, though, and 100 MB is as well, but I agree for desktop use in research it is more than OK.
In my field many people use MongoDB to store data - if they feel like it... I personally use SQLite for keeping logs sometimes though that is quite annoying when all your files live on NFS - btw: how does DuckDB deal with NFS if we are at that?
Because if it doesn't, you can just let your experiments log to CSV - which likely still is more than what some people do and a perfect input to DuckDB...
One typical way to work around this limitation is to create a new table, copy the data there, and drop the old table. Copying from table to table is in general very fast.
(Disclaimer: I work at DuckDB Labs)
[1] https://duckdb.org/docs/guides/performance/schema#microbench...
Does this work on Windows (which is more likely what they're using)?
Does it come with conda (which pandas does)?
I'm not sure the author is familiar with researchers, nor if this is just an ad...
It runs on macOS, Windows and Linux.
It's not yet available as a conda package but can be installed from conda-forge (or from PyPI via pip). Additionally, it also works in R (via CRAN) and has bindings to many other languages.
(Disclaimer: I work at DuckDB Labs)
Also, my naive and quick look at your docs gives me the impression this isn't actually a storage engine (which didn't come across in the article), but instead an in-memory query engine? The FAQ doesn't really clarify this (e.g. is there an actual storage format, how does this compare to using one of the more traditional dbs, how stable it is (c.f. sqlite being recommended by the library of congress), does this assume all the data can be loaded into memory, or can I stream the data in).