D1: Improvements to performance and scalability
blog.cloudflare.com
blog.cloudflare.com
DuckDB works great as an in-memory database (it's also the default mode).
(I use it personally, but it's not the same thing as what we're building with D1)
Basically, we already have a product for doing DuckDB-like things, and the D1 architecture isn't great for high volume OLAP workloads.
Disclaimer: I'm an engineer at Cloudflare, but not on a developer platform team. I'm not speaking for our developer platform strategy or anything, I'm just commenting on how it looks from where I sit.
I'm sure there's a lot of really cool local-first databases out there, but SQLite has the benefit of being incredibly widely battle-tested, with literally billions of installations worldwide. It has received thorough security research and fuzzing (it's part of Chrome's attack surface after all). And there's tons of resources online to help people understand how to use it. Although I'm sure there are alternatives that serve certain use cases better it's hard to imagine anything coming close for ours.
That said, the storage engine we've built is not that heavily dependent on SQLite specifically. Any database that uses a write-ahead log like SQLite does should be possible to adapt to it in the future. So maybe we'll eventually open it up to a variety of choices, or even let you bring your own as a Wasm module.
> DuckDB is indeed a free columnar database system, but it is not entirely built on top of SQLite. It exposes the same front-end and uses components of SQLite (the shell and testing infrastructure), but the execution engine/storage code is new.
Quite the hobby project for a lot of people too! Other folks doing this:
Rqlite https://hn.algolia.com/?query=Rqlite&sort=byDate , Dqlite https://hn.algolia.com/?query=dqlite&sort=byDate , Litestream https://hn.algolia.com/?query=Litestream&sort=byDate / LiteFS https://hn.algolia.com/?query=LiteFS&sort=byDate, marmot, mvsqlite https://hn.algolia.com/?query=mvsqlite&sort=byDate
https://www.philipotoole.com/9-years-of-open-source-database...
You can't just round down 1.82ms to 1ms.
D1 was in open alpha prior to today, iirc.
Though they do say:
> when we enable global read replication, you won’t have to pay extra for it, nor will replication multiply your storage consumption
With AWS you’ll have at least one read replica for failover, so $0.23/GB. And if you really want global read replicas, with AWS you might end up with something like a primary in North America and read replicas in South America, Europe and Asia. That would work out to $0.46/GB, so closes the gap a bit.
The vision of the Supercloud is that you give us your code and we'll figure out where and how to execute it: https://blog.cloudflare.com/welcome-to-the-supercloud-and-de...
And for reference, here’s the original D1 announcement with some additional info https://blog.cloudflare.com/introducing-d1/
Note there are several alternatives here, too. Workers Durable Objects[0] provide a lower-level primitive for building advanced distributed systems. But D1 is easier to use for typical use cases. For blob storage you might use R2[1]. And for large databases Workers can easily integrate with several serverless database providers.[2]
[0] https://blog.cloudflare.com/introducing-workers-durable-obje... [1] https://www.cloudflare.com/products/r2/ [2] https://blog.cloudflare.com/announcing-database-integrations...
What I'd like to enable here is a progression where you start out prototyping your app with a single D1 database, which is easy to use and reasonably fast. Then as you grow we provide tools to let you transition to many D1 databases sharded in a way that makes sense (e.g. per-user). Apps that want even more control can move to using full-on Durable Objects (which will soon support a SQLite database per-object).
That said, there are certainly many use cases out there where simple monoliths make sense, especially non-interactive data crunching. I'm not sure yet if D1 will ever be the right choice for those, but the Workers platform aims to provide many options.
I've only started to think about this and I'm thinking the hardest part will be dealing with cross-cutting concerns (in a non-auto sharded world manually creating multiple database) and trying to find a way to keep each database isolated without extra burden compared to using a hosted Postgres.
As an aside, that lan optimized house was a gaming dream. Hope your new house is as awesome.
Can you elaborate this little bit more? Im using DO today and i have a bad time sharding my data (works, but i hate it);
So i will have the option to use the standard store or/and SQLite?
If so, i dont can keep with my DO (because i have control of everything) and use SQLite for things that is bigger than what the value store supports.
In the future each DO will have a private SQLite database. The key/value store will actually be redirected to store into a special table in this database, but probably new apps will just use the database and not the KV store.
Separately from that, I would like to develop tools that make sharding Durable Objects (and D1 databases) easier. Today it's a pain to do manually. This is independent from the underlying storage model, though.
For comparison, fly.io, turso's provider, has 34 locations and well-documented reliability issues.
There is danger in centralized systems.
(Disclosure/disclaimer: I work for Cloudflare but in a different department; I'm not an expert on Privacy Pass.)
Sounds like something which is designed to create hard to detect tracking vectors.
All in the name of "privacy" and "security". Though of course, the only security that is served here is the ad network's revenue stream, who are the only ones who really care if a real device or a bot is accessing their services. For most other cases, a DDoS is a DDoS regardless of it being initiated by actual users (e.g. an HN hug of death) or by a bot network.
But once cloud flare hates you, you have to bow to them.
Try browsing the normal web over Tor sometime and see how bad it can get.
I still use Temporary Containers and lately, I've noticed a sharp decline in these Cloudflare captchas. I don't know if its because people are moving away from them, or Cloudflare just found a better way to finger print me.
People like to blame the easiest target.
This is true. Neon however offers bottomless storage and D1 is 100Mb currently going to 1Gb.
Neon doesn't have read replicas and even if they already work on it, I wouldn't expect it before 2025 if at all and never at CF's pricing (Neon still charges for egress).
[1] I would compare D1 rather to Turso or LiteFS from Fly or PlanetScale with many read-replicas
- Development since start took time (with a good reason), hope we see soon read replication and higher pace
- So, some roadmap _with_ ETAs would be great, e.g., will replication come this or next year?
- I don't care about read performance because I know that sqlite paired with CF's infrastructur will perform but I would love to have more infos on write performance; this is really sqlite's weakest point and I haven't seen any implementation which could compete anywhere with other dbs; sqlite slows down very quickly when doing a couple of small writes/sec
Then, I'd like to see the close competition here on HN, so db-providers with many read-replicas (CEO's and/or devs from Turso and Fly/LiteFS), commenting on D1 and how they compare against and what they plan to compete. This is a too exciting space and time to be laissez-faire.
There are a lot of ways to approach SQLite replication and one isn’t better than another necessarily. They simply have different trade-offs.
I think Deno is lightyears ahead.
And the team has been aware of the issue for years now
Prisma isn't standard and too slow anyway.
Among other things, we made Wrangler (Workers CLI tool) use the open source workerd by default for local development, so local dev should produce a much more precise simulation now (since it's literally running the same code).
I tried to use that recently and it was a disaster. I wrote about my experience here:
https://twitter.com/pierbover/status/1641474067013271552
I then opened these two issues:
https://github.com/cloudflare/workers-sdk/issues/2962
https://github.com/cloudflare/workers-sdk/issues/2964
I ended up moving the project over to Netlify + Edge functions. I had it all working in like 5-10 mins as it should. Took me two hours to figure out why Workers weren't working in my Pages project, and could never get Workers working properly with my Astro project.
I think you're working exclusively on the engine of Workers which is really top notch, but Cloudflare really needs to improve the outer layer which affects DX considerably.
https://blog.cloudflare.com/pages-and-workers-are-converging...
Edit:
It does!
I thought this was a wrapper for Litestream but apparently it's a parallel project by the same author who Fly hired.