Around the World with SQLite3 and Rsync
fly.io
fly.io
To be fair, "explains a solution without explaining a problem" is one of my most common criticisms for technical content in general, but I am pretty sure this wouldn't have gotten to front page if it wasn't for fly's reputation. There's a lot of passive voice which makes it really hard to figure out who is doing what to what, and most importantly _why_?
Can I ask you not to say things like this? Specifically: this idea that there's a bar for technical stunt driving that we need to clear to get things up on the blog. I know I can't demand that of anybody; I'm just asking.
I can't tell you how many things we haven't written on our main blog because we're worried about whether it's as technically interesting as the "Gossip Glomers" thing we did with Aphyr.
Sam's a good writer and this is a pragmatic post about an approach he took to building something. The world is incrementally better with this post in it than without it. Holding off until we could have integrated LiteFS and written a 4 paragraph coda about sqlite3 internals would not make the world a better place.
I worry all the time about "Fly.io fatigue" on HN, for what it's worth! But the only sane response we can have to do that is: we just write what we want to write, and we let the community do what it's going to do. If 'dang gave us a header we could use to dampen some of our posts, we'd use it. :)
If I'm crazy to ask this, say so. We've been grappling with this for years now; if we knew exactly how to handle it, I wouldn't have to write weird things like this.
> But the only sane response we can have to do that is: we just write what we want to write, and we let the community do what it's going to do.
Move over paxos and sqlite3; HN needs FlyLLM and Fly vGPU to satiate its appetite now ;)
But I decided to leave it because I think I stand by the implication that reputation for specific technical blogs can be earned or used up. Fly.io has earned a lot of reputation and I (and I guess others) clicked on a title I would not usually have because I saw the domain. In this case I feel like some reputation was "used up" and I am (slightly) less likely to click on a post solely because I see a fly.io domain in future.
Maybe I am just not the target audience for the post and it is better than my assessment, but if not then I think people value HN as an honest audience that doesn't hold back on criticism (I know I do and I have received harsh but valid criticism from hn on stuff I have written).
Fly.io can probably keep making front page based on reputation if there is a balance of very high quality content and other posts, but I feel like something like this could have been 1000% better with another revision round and some editing (probably only 20% more effort) so I will keep saying stuff like this I guess?
Don't. I can honestly say that I didn't write this post targeting HN. I'll go further... this post wasn't meant for people who are unlikely to use https://github.com/fly-apps/dockerfile-rails#overview. I recently added some features to that gem whose usage may not be intuitively obvious. I wrote this post to explain some of the motivation for those features.
I don't know how to mark posts as not intended for HN (and truth be told, if there was such a feature, I'd be inclined to overuse it). I don't know where else I should have posted this content, but I'm not sure I would be inclined to move it. In any case, this post, as written, serves a purpose for me. Somebody not you and not me felt it belonged here. We can both second guess that decision. Either way, there is no reason for either of us to feel bad.
I love fly.io and appreciate the content but that doesn't sit right with me.
I just worry about us writing less stuff because people assume there's going to be some stunt driving sequence or shootout or CGI or whatever, and not everything people write here is like that.
Minute later
(I may be getting pushed into responding adversarially here just by message board dynamics. Sorry. Like I said up front: it's a weird thing to write! But what am I going to do, filter my thoughts before writing them on HN?)
As for the database that handles your business logic … 3 minutes of googling would show that rsync is a poor (or at least a brittle) substitute for multi-master replication when it comes to database engines like InnoDB for MySQL (which is what most websites on the Internet use):
https://dba.stackexchange.com/questions/91287/is-possible-to...
However, you can consider setting up multi-master replication, or if you want your app to be portable, you can set up a replication layer in the app itself. The order of operations should only matter “per activity” (eg chatroom, document etc) and not globally.
That’s what we set up at https://qbix.com/platform — except that the instances don’t all have to be under the control of one company. It’s basically a decentralized social networking platform. Federated, like Mastodon, but far more general-purpose, so you can build any applicatiom with it
So, I was very interested in this article which seemed to suggest that fly.io, somehow found a way to make this work anyhow. TL;DR — looks like the answer is: they did not.
SQL databases are not great at storing huge chunks to data (petabytes of video) but those are easily handled by the native filesystem. Just make sure the write access is safely wrapped in a DB transaction, RAII style. This works and scales well enough except for the failure mode: over stressed egress (typically SELECT queries) will starve ingest (a more or less constant load of UPDATE & INSERT queries).
At that point, you want to scale horizontally. As mentioned, replicating the SQL backend file over NFS or rsync is simply a horrible idea. Surprisingly we ultimately discovered that SQL-level replication did not solve the problem reliably either: scale improved, but the starvation failure mode (read locks blocking DB writes) persisted (be it less frequently).
The replication layer we needed was at the application level where it's much easier to express that actually we did not need the real-time coupling between media ingest and egress. Eventual consistency is good enough: if the system is about to collapse, simply fail by adding latency so you can always keep ingest healthy.
Worst case, people watching a popular sports event will see players buffering. Which is bad — don't get me wrong — but far better than a catastrophically blocked pipeline and congested video encoders.
Is there a tl;dr to what this does and why?
1) Delete the first two paragraphs.
2) Tell me what you're going to tell me, but without all the highlights and bullet points.
3) Sell me on why I want to read this article. Don't focus on the technical; focus on the human, what I as a person will get from it.
4) In later technical sections, don't give me instructions. Tell me what you are doing in general to solve some problem, then just show me the commands, and I will infer that I could also do X, do Y, do Z.
5) If you want to give a lot of technical detail and backstory, either put it in an expand, or in some kind of colored box so I have the visual context that I can probably skip this whole, continue the story, and come back to this later.
6) The recap shouldn't be larger than two paragraphs.