> What are some reasons you reckon that the current setup won't scale beyond 10GB?
It's more of an arbitrary threshold right now. A lot of testing that we do right now is chaos testing where we frequently kill nodes to ensure that the cluster recovers correctly and we try to test a range of database sizes within that threshold. Larger databases should work fine but you also run into SQLite limitations of single writer. Also, the majority of databases we see in the wild are less than 10GB.
> Leaving aside stability related bugs, what design decisions previously made were the caused painful bugs / roadblocks?
So far the design decisions have held up pretty well. Most of the PRs were either stability related or WAL related. That being said, the design is pretty simple. We convert transactions into files and then ship those files to other nodes and replay them.
We recently added LZ4 compression (which will be in the next release). There was a design issue there with how we were streaming data that we had to fix up. We relied on the internal data format of our transaction files to delineate them but that would mean we'd need to uncompress them to read that. We had to alter our streaming protocol a bit to do chunk encoding.
I think our design decisions will be tested more once we expand to doing pure serverless & WASM implementations. I'm curious how things will hold up then.
> Consequently, what things majorly surprised you in a way that perhaps has altered your approach / outlook towards this project or engineering in general?
One thing that's surprised me is that we originally wrote LiteFS to be used with Consul so it could dynamically change its primary node. We kinda threw in our "static" leasing implementation for one of our internal use cases. But it turns out that for a lot of ancillary cache use cases, the static leasing works great! Losing write availability for a couple seconds during a deploy isn't necessarily a big deal for all applications.