RDS has some significant downsides which I would personally consider no go if I were in a position to evaluate it. The one which stands out as particularly painful from my past experience is… it’s excruciatingly slow to provision or make configuration changes. Like lose whole days of work to a few iterations of trial and error slow. Combined with AWS’ sprawling and inscrutable set of authorization and configuration options, the weird idiosyncrasies between most of their offerings, and the absolutely opaque naming applied to most of those offerings… trying to use RDS effectively as a managed database service felt more to me like becoming a full time ops professional.
But tailscale isn't some random group of engs. They've probably got the chops to pull off literally anything they want to. I mean TFA casually mentions online cross-database transfers, multiple zero-downtime schema migrations, inspecting litestream's replication code for feasibility, deftly modifying sqlite WAL checkpoints... all in one breath.
It seems to me that tailscale engs want to avoid DBA work but also not use managed offerings, and so, they're comfortable paying the costs they have to (such as multiple migrations).
> ...rather than improving the product?
Well, you'd guess they want to be able to continually improve their already credible product too. When TFA points out that zero vendor lock-in and hassle-free, local end-to-end tests are non-negotiable, I think it is for this reason.
----
> if they're this talented, is their time really best spent on...
From: https://tailscale.com/blog/go-linker/
"People are often surprised and sometimes horrified when they learn that Tailscale maintains its own fork of the Go toolchain. Tailscale is a small startup. Isn't that a horrible distraction, a flagrant burning of innovation tokens?"
"Maybe. But the thing is, you write code with the engineers you have."
"We had a problem: We kept crashing on iOS, and in addition to being awful, it was preventing us from adding features."
"Another team might have decided to cut even more features on iOS to try to achieve stability, or limited in some way the size of the tailnet that iOS could interact with."
"Another team might have radically redesigned the data structures to squeeze every last drop out of them."
"Another team might have rewritten the entire thing in Rust or C."
"Another team might have decided to accept the crashes and attempted to mitigate the pain by making re-establishment of connections faster."
"Another team might have decided to just live with it and put their focus elsewhere."
"The Tailscale team has Go expertise, spanning the standard library to the toolchain to the runtime to the ecosystem. It’s an asset, and it would be foolish not to use it when the occasion arises. And the fun thing about working on low level, performance-sensitive code is that that occasion arises with surprising frequency."
"Blog posts about how people solve their problems are fun and interesting, but they must always be taken with a healthy dose of context. There may be no other startups in existence for which working on the Go linker would be a sensible choice, but it was for us."
if zero vendor lock-in and hassle-free, local end-to-end tests are non-negotiable, why are they using s3? migrating to another s3 compatible backend would be similar in effort to migrating from aurora mysql or postgres to another managed mysql or postgres service or to self-hosted mysql or postgres
You may be right. I have no experience migrating litestream but from the docs (https://litestream.io/guides/) it is literally cp'ing files from S3 to wherever and exec'ing one of these one-liners (of course, the devil is in the details):
litestream restore -o my.db s3://BUCKETNAME/PATHNAME
litestream restore -o my.db abs://STORAGEACCOUNT@CONTAINERNAME/PATH
litestream restore -o my.db gcs://BUCKET/PATH
litestream restore -o my.db s3://SPACENAME.nyc3.digitaloceanspaces.com/db
litestream restore -o my.db s3://BUCKETNAME.us-east-1.linodeobjects.com/db
litestream restore -o my.db sftp://USER:PASSWORD@HOST:PORT/PATHIt's a super confusing argument regardless, because the industry is lousy with "S3-compatible backends".
I just don't believe the tale of "such skilled engineering teams" which don't show that in their products but blogposts.