Choose DuckDB rather than SQLite
tracewayapp.com
tracewayapp.com
Please, folks, write with your own voice -- especially if it's for your business blog. It's good for you as an author (practice makes perfect) and it's good for your readers (whom you want to influence).
That is workload specific. Title should be "Choose DuckDB rather than SQLite for Analytics" IMHO
Can you explain this more, especially why SQLite is best at OLTP and what happens at scale?
I've just checked their website, and they state "relying on ClickHouse to power these analytics use cases". That's not OLTP. https://clickhouse.com/comparison/postgresql
Fair to say, seeing 1000x w/o any trace of proof won't help me to choose.
It has some characteristics typical of OLTP engines. But they are targeted and limited to areas DuckDB feels are important.
So yes, all these benchmarks are great, but it wasn't so fun working with DuckDB when I had to close duckdb cli, just so a query in another script could run.
Now it's debt. Oops.
I ran into this myself; I tried using SQLite to store the results of whole-internet rDNS scan and a count() over the entire DB could take 8 minutes. I used the wrong DB for the job and the narrative around SQLite/duckdb is around reckoning with perfectly reasonable limitations and tradeoffs that SQLite made.
And DuckDB is reasonably fast for even single record writes. ~1000x slower than SQLite, but that's still pretty fast if you're only doing a few hundred writes per second or batching.
And I am fatigued by the AI style in all code comments, reviews, PRs etc :(
That server is now $51.09 for those wondering See https://news.ycombinator.com/item?id=48540844
TLDRTL;DR: If everything you do is column-store territory, use a column-store.
That said, I do think duckdb has a wider range of use cases than people here might think. It can whip through fairly large datasets (I use it for ~1B row tables all the time)
DuckDB can't be. PR was sent a year ago. Blocked on the same concurrency model issue in the other sub thread.
Specifically on windows, the database can't read its own WAL file from a different thread in the same process.
Love DuckDB for being permissively open source, great tech and performance!
It's a big additional leap to for our software to try to reach into others' sites, get through any anti-bot defenses they may be running, try to scrape their content and evaluate on whether it's sufficiently human-authored to be on HN.
There's generally a wider range of LLM involvement with a long-form post than the typical, relatively brief HN comment, which then opens the way for more debate on HN about "how much" LLM influence the post has and how much should be allowed on HN. Part of what we're trying to optimize for on HN is minimizing offtopic/meta discussion, so we don't want to encourage this kind of debate.
Our heuristic about article quality is largely unchanged from before LLMs were an issue: if an article is badly written, it shouldn't be on HN, and should be flagged.
As to what should be tested: front-page items, possibly even a subset of those (top 10--15 of 30). That's going to be a limited set of items per day, though more than just 30. (I don't know how many items cycle through the front page on a daily basis, though I believe daily submissions as of 2022 were about 1,000/day (<https://web.archive.org/web/20220116193045/https://whaly.io/...>)).
Working this into the HN story-processing lifecycle might be a good call.
I'd much prefer not seeing a bunch of AI slop in submissions, by way of generated output. AI as part of the resarch process I think I could live with.
AI-generated content seems, definitionally, not to be intellectual in nature, and would seem to go against HN's prime directive. It also seems to make HN lose its collective mind, which has long been another mod consideration.
In the meantime, please feel free to flag items that are badly written/unpleasant to read, and email us if something is on the front page that shouldn't be there.