HNHacker News
TopNewBestAskShowJobs

NathanFlurry

512 karma · joined September 16, 2016

Building open-source actor library, a serverless primitive for stateful workloads.

https://rivet.dev

submissionscomments
NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
Author here! I agree it's very similar to the actor model, but I kept the article's scope small, so I didn’t cover that.

In fact – Durable Objects talks a bit about its parallels with the actor model here: https://developers.cloudflare.com/durable-objects/what-are-d...

You might also appreciate this talk on building a loosely related architecture using Erlang, though it doesn't implement an actor-per-database pattern – https://www.youtube.com/watch?v=huGVdGLBJEo

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
Author here, love this take.

I've chatted with a few medium-sized companies looking at Durable Objects for this reason. DB-per-tentant removes much of the need for another dedicated team to provision & maintain infrastructure for the services. It's almost like what microservices were trying to be but fell woefully short of achieving.

It's disappointing (but understandable) that "serverless" received a bad rap. It's never going to fully replace traditional infrastructure, but it does solve a lot of problems.

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
Fair point, noted.
NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
Hopefully, it matures into a healthy open-source ecosystem that doesn’t rely on proprietary databases.

More companies than people realize are already building and scaling with DO SQLite or Turso internally. Almost every company I've talked to that chooses Postgres hits scaling issues around Series A — these companies aren’t.

NathanFlurry··on SQLite-on-the-Server Is Misunderstood: Better at Hyper-Scale Than Micro-Scale
If you care only about serverless, databases like PlanetScale, CockroachDB Cloud, and DynamoDB work well.

The biggest strength of using SQLite here is that it provides the benefits of a familiar SQL environment with the scaling benefits of Cassandra/DynamoDB.

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
The two tables are intended to be part of the same "chat" partition (ie SQLite database). You can join them with a native SQLite query. Seems I should make this more clear.

Cheers

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
Author here.

To clarify — is your concern that the only scaling options I listed are proprietary services?

If so, I completely agree. This article was inspired by a tool we're building internally, based on the same architecture. We knew this was the right approach, but we refuse to rely on proprietary databases, so we built our own in-house.

We’re planning to open-source it soon.

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
This seems to be the biggest hesitation I've heard over and over by far. There absolutely needs to be a good story here for both (a) ad-hoc cross-partition queries and (b) automatically building a datalake without having to know what ETL stands for.

However, this isn't so much different from Cassandra/DynamoDB which have a similar problem. You _can_ query cross-partition, but it's strongly discouraged and will strain any reasonably sized cluster.

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
Great Scott!
NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
Yep.

Nile (https://www.thenile.dev/) is trying to address this use case with a fully isolated PG databases. Though, I don't know how they handle scaling/sharding.

NathanFlurry··on SQLite-on-the-Server Is Misunderstood: Better at Hyper-Scale Than Micro-Scale
DuckDB crushes SQLite in heavy data workloads according to ClickBench by 915x. (Link below since it's looong.)

DuckDB also has a WASM target: https://duckdb.org/docs/stable/clients/wasm/overview.html

I don't know enough about DuckDB to understand the tradeoffs it made compared to SQLite to achieve this performance.

https://benchmark.clickhouse.com/#eyJzeXN0ZW0iOnsiQWxsb3lEQi...

NathanFlurry··on SQLite-on-the-Server Is Misunderstood: Better at Hyper-Scale Than Micro-Scale
I think the other comments have the application-level approaches covered.

However, I suspect the infrastructure will provide this natively as it matures:

- Cloudflare will probably eventually add read replicas for Durable Objects. They're already rolling it out for D1 (their other SQLite database offering). [1]

- Turso has their own story for read replicas. [2]

[1] https://blog.cloudflare.com/building-d1-a-global-database/#s... [2] https://docs.turso.tech/features/embedded-replicas/introduct...

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
Author here.

Agreed — I think adding some comparisons to other database partitioning strategies would be helpful.

My 2 cents, specifically about manually partitioning Postgres/MySQL (rather than using something like Citus or Vitess):

SQLite-on-the-server works similarly to Cassandra/DynamoDB in how it partitions data. The number of partitions is decoupled from the number of databases you're running, since data is automatically rebalanced for you. If you're curious, Dagster has a good post on data rebalancing: https://dagster.io/glossary/data-rebalancing.

With manual partitioning, compared to automatic partitioning, you end up writing a lot of extra complex logic for:

- Determining which database each piece of data lives on (as opposed to using partitioning keys which do that automatically)

- Manually rebalancing data, which is often difficult and error-prone

- Adding partitions manually as the system grows

- (Anecdotally) Higher operational costs, since matching node count to workload is tricky

Manual partitioning can work fine for companies like Notion, where teams are already invested in Postgres and its tooling. But overall, I think it introduces more long-term problems than using a more naturally partitioned system.

To be clear: OLTP databases are great — you don’t always need to reach for Cassandra, DynamoDB, or SQLite-on-the-server depending on your workload. But I do think SQLite-on-the-server offers a really compelling blend of the developer experience of Postgres with the scalability of Cassandra.

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
I (sort of) disagree. Cassandra- & DynamoDB-based systems which are also partitioned do fine without a central OLTP DB.

Wrote a bit about it here: https://news.ycombinator.com/item?id=43246212

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
> Is it even realistic to depend on transactional guarantees, with hundreds of services hammering the DB(s) more or less concurrently?

If a single request frequently touches multiple partitions, your use cases may not work well.

It's the same deal as Cassandra & DynamoDB: use cases like chat threads or social feeds fit really well because there's a clear ownership hierarchy. e.g. message belongs to a single thread partition, or a social post belongs to a feed partition.

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
> The obvious caveat here is any situation where you need global tables

A lot of people still end up storing data that's not frequently updated in a traditional OLTP database like Postgres.

However:

I think it always helps to think about these problems as "how would you do it in Cassandra/DynamoDB?"

In the case of Cassandra/DynamoDB, the relevant data (e.g. user ID, channel ID, etc) is always in the partitioning key.

For Durable Objects, you can do the same thing by building a key that's something like:

``` // for a simple keys: env.USER_DO.idFromName(userId);

// or for composite keys: env.DIRECT_MESSAGE_CHANNEL_DO.idFromName(`${userAId}:${userBId}`); // assumes user A and B are sorted ```

I've spoken with a lot of companies using _only_ this architecture for Durable Objects and it's working well.

NathanFlurry··on SQLite-on-the-server is misunderstood: Better at hyper-scale than micro-scale
This.

> In this case I think you can let them become inconsistent in the face of, e.g., write errors.

For devs using CF Durable Objects, people frequently use CF Queues or CF Workflows to ensure that everything is eventually consistent without significant overhead.

It's a similar pattern to what large cos already do at scale with keeping data up to date across multiple partitions with Cassandra/DynamoDB.

NathanFlurry··on SQLite-on-the-Server Is Misunderstood: Better at Hyper-Scale Than Micro-Scale
The folks over at StarbaseDB (https://starbasedb.com/) are working on building tools for shareded SQLite.

From the companies I've talked to, most developers using this architecture are building quick scripts to do this in-house. Both Turso and Durable Objects SQLite already a surprising amount of usage that people don't talk about much publicly yet, so I suspect some of this tooling will start to be published in the next year.

NathanFlurry··on SQLite-on-the-Server Is Misunderstood: Better at Hyper-Scale Than Micro-Scale
Author here, happy to answer questions!
NathanFlurry··on Patterns for Building Realtime Features
Durable Objects solves so many problems for realtime & persistence. The biggest problem is they're vendor locked, so we can't use them if we want to keep all of our infrastructure on AWS.

So we built out a library that lets you run Durable Object-like backends on any cloud: https://github.com/rivet-gg/actor-core

NathanFlurry··on Multiplayer Filesystem in Durable Objects
> Our multiplayer filesystem persists its data to Durable Objects with the help of DOFS - a mostly POSIX compatible filesystem layer with Emscripten's filesystem interface.

This reminds me of HAFS – a FUSE-based filesystem based on FoundationDB [1]. It does a lot of the same things under the hood to work with FDB's requirements (i.e. slicing, etc). Looking forward to seeing your time travel implementation.

Regarding Gitlip – I've spent a lot of time in the gaming industry where many workflows rely on Perforce for collaboration/file locking instead of Git & Git LFS. It's a widely disdained tool because of its rigidity and complexity. Gitlip feels like a better approach to collaboration than Perforce, but applied to wider use cases. I'll have to give it a spin some time.

[1] https://github.com/andrewchambers/hafs

NathanFlurry··on We will rewrite SQLite. And we are going all-in
I'm really excited for the idea of Limbo, though I have a few concerns:

#1:

I would like to see some sort of guarantee by Turso that they plan on sticking to an MIT or Apache license and don't plan on re-licensing to AGPL or non-OSI-approved licenses.

Some context – the founders of Turso were early employees at ScyllaDB (phenomenal Cassandra rewrite, powering Apple, Discord, etc). I'm speculating on how much influence they had over the licensing of the database, but ScyllaDB was kept under an AGPL license during their ~8 year tenure there. However, ScyllaDB still rug pulled their license and switched to enterprise-only in 2024 [1].

I also run a venture-backed startup with an open-source product (Apache 2.0), so I completely understand the pressures that cause these decisions. A rug pull with something like Limbo seems very tangible as Turso competitors like Cloudflare D1 start making cloud-based SQLite a commodity.

Regardless, we're avoiding touching libSQL or Limbo because we've already been burned by Nomad (BSL relicense), Redis (SSPL relicense), and CockroachDB (custom enterprise relicense); all of which affect our ability to keep an Apache 2.0 license. (e.g. AGPL is a viral license that affects us to change our license if we use software like Minio.)

It's worth noting that Limbo does not have a CLA, which is a good indicator that they intend to keep Limbo licensed permissively.

#2:

One big focus of Limbo (and libSQL) is the WASM-based extensions for user-defined functions. Postgres also has a powerful UDF system because users need to be able to define custom logic without an extra network hop between client <-> database. However, SQLite runs on the same machine as the code that's querying it, so it doesn't make sense to make the database heavier to allow UDF when there's already a ~0-latency way of providing extensions.

Frankly, this feels a lot like Turso investing in features to match the constraints of Turso Cloud. Turso Cloud can't run arbitrary code, so they need to implement user logic in to the database itself.

I'd much rather a Supabase-like approach: provide vanilla Postgres-as-a-service and add Auth/PostgREST/Realtime/Functions/etc as a separate bundled service instead of rewriting Postgres itself. Similarly, Neon innovated on Postgres' storage architecture, but still provides full compatibility.

StarbaseDB [2] is taking the Supabase-like approach with SQLite, and I'm really excited for this. They're leveraging Cloudflare Durable Objects + SQLite to do things like provide a native auth & rate limiting & sharding & Stripe integration in to the database without modifications to SQLite itself.

#3:

I'm not looking forward to the fracturing of the SQLite ecosystem – similar to what happened with the enshittification of the MySQL ecosystem with MariaDB or NodeJS when io.js forked off. This is not a comment against Limbo itself, just something that's inevitably going to happen.

---

To be clear – despite these concerns, I'm very excited for Limbo for:

- The new governance model

- Heavier investment in to SQLite tooling

- Promote SQLite-in-the-cloud

- Promoting local-first database architectures

Props to the engineering team for their work on this (esp the DST) and the founders/management to green-light such an ambitious project.

[1] https://www.scylladb.com/2024/12/18/why-were-moving-to-a-sou...

[2] https://starbasedb.com/

NathanFlurry··on Kronotop: Redis-compatible, transactional document store backed by FoundationDB
What fascinates me most is not the protocol itself, but the potential to add a thin scripting layer (similar to Redis' EVAL or the newer modules API) to enable complex logic with atomic transactions and high QPS.

In the past, we relied heavily on using EVAL[SHA] with 200+ loc Lua scripts in order to implement high throughput, atomic transactions for realtime systems. We also used the JSON & Redis Query Language (previous named "full-text search") to build a more maintainable & strongly consistent system than using raw key-values and manually building secondary indexes.

We’ve since migrated to a native FoundationDB and SQLite hybrid setup, but this approach would have been really helpful for early-stage prototyping with a higher performance ceiling (thanks to FDB sharding) than a single-node Redis with AOF.

Related: Redis Cluster is a world of pain when handling clustering keys and cross-node queries and orchestration. DragonflyDB is chasing after the market of companies considering sharding Redis because of performance issues by providing better single-node performance. There's probably an alternative approach that could work by using an architecture like this.

NathanFlurry··on Kronotop: Redis-compatible, transactional document store backed by FoundationDB
There's a version written in Go, though clearly just a hobbyist project: https://forums.foundationdb.org/t/introducing-the-redis-prot...
NathanFlurry··on Kronotop: Redis-compatible, transactional document store backed by FoundationDB
Redis on Flash [1] is one of the key ways Redis gets people on their enterprise plan (esp before the SSPL relicense). I've spoken with their enterprise team before; it's ungodly expensive, even compared to AWS MemoryDB (wish I had the numbers on hand).

[1] https://redis.io/resources/building-large-databases-redis-en...

NathanFlurry··on Show HN: Instantly visualize any codebase as an interactive diagram
This is fun! I’ve come across a few tools like this before and almost dismissed it due to past poor experiences. However, I just diagrammed our startup’s codebase, and it was surprisingly similar to our hand-made diagram. I tried customizing it with specific instructions to ignore legacy code & provide an understanding of our edge architecture, but the one-generation-per-day rate limit is a bit restrictive.

For comparison:

- Hand made diagram: https://github.com/rivet-gg/rivet/blob/d45bf556e903404ab2df0...

- GitDiagram (no instructions): https://gitdiagram.com/rivet-gg/rivet

NathanFlurry··on Show HN: An edge first feature flag implementation on Cloudflare
Neat! Love the ability to deploy some like as powerful as Launch Darkly effortlessly with Workers.

D1’s read replicas [1] is still pretty new, how has this worked out for a read-heavy workload like FlagShip?

[1] https://blog.cloudflare.com/building-d1-a-global-database/#s...

NathanFlurry··on Show HN: Rivet Actors – Durable Objects build with Rust, FoundationDB, Isolates
I skimmed a bit about Durable Promises, and if I understand correctly, they’re similar to workflow tools like Temporal but come with high-quality language bindings that let you write workflow steps using async/await.

We built something almost identical in Rust to let us use async/await for long-running & failure-prone workflows. This powers almost everything we do at Rivet. We have a technical writeup coming later, but here's a couple links of interest:

- Rough overview: https://github.com/rivet-gg/rivet/blob/58b073a7cae20adcf0fa3...

- Example async/await-heavy workflow: https://github.com/rivet-gg/rivet/blob/58b073a7cae20adcf0fa3...

- Example actor-like workflow: https://github.com/rivet-gg/rivet/blob/58b073a7cae20adcf0fa3...

Rivet Actors can start, stop, or crash at any time and still continue functioning, much like Durable Promises.

However, their scope differs: Rivet Actors are broader and designed for anything stateful & realtime, while Durable Promises seem focused on workflows with complex control flow.

Rivet Actors can (and likely will) support workflow-like applications in the future, since state management and rescheduling are already built-in.

For a deeper dive, Temporal has a writeup comparing actors and workflows: https://temporal.io/blog/workflows-as-actors-is-it-really-po...

NathanFlurry··on Show HN: Rivet Actors – Durable Objects build with Rust, FoundationDB, Isolates
We document as many design decisions like this as we can. Here’s a related bit on serial vs. parallel RPC/message handling [1].

> I guess this sort of comes by forcing everything to round-trip through FoundationDB for state?

If you’re using the KV API directly, this is correct.

Actors also have a `this._state` property, which is automatically written to FDB after each RPC call if modified [2]. This allows developers to rapidly prototype by writing vanilla JS code like `this._state.count += 1` without having to worry about writing state and its associated edge cases.

> What about calls to other actors? If I make a potentially state-changing RPC call to some other actor as part of handling an RPC, do those commit together, or is it possible for the other actor to commit without me?

Not at the moment. You’d need to use a 2-phase commit for that.

[1] https://rivet.gg/docs/internals/design-decisions#parallel-rp...

[2] https://rivet.gg/docs/state#state-saves

NathanFlurry··on Show HN: Rivet Actors – Durable Objects built with Rust, FoundationDB, Isolates
Haha, I’m with you there. I was considering calling that out in the post.

I dearly love Erlang & co., but its learning curve is way too steep for most developers today. Our goal is to bring the benefits of Erlang/Akka/Orleans/etc. to more developers by:

(a) supporting mainstream languages, and

(b) lowering the technical and conceptual barriers to entry.

← PreviousPage 2 of 4Next →