Which leads to the real takeaway which is "Tiger Style": https://tigerstyle.dev/ which I am partial to, along with Rich Hickey's "Hammock Driven Development" https://www.youtube.com/watch?v=f84n5oFoZBc
"Tiger on Hammock" will absolutely smoke the competition.
(edit: add links)
To keep things simple. My current company is running multiple instances of back-end services for absolutely no fucking reason, and I had to fix numerous race condition bugs for them. I had an interview with a startup where, after I asked why they were using distributed DynamoDB locks in a monolith app with only a single instance running, the person said "it works for us" and got defensive. Later they told me I wasn't experienced enough. I am so frustrated that there appears to be zero basic engineering rigor anywhere I can find nowadays.
> And for up to how many hundred terabytes of data can you get away with the single beefy server?
Do you even need to store many hundred terabytes of data? I have never encountered a scenario in my career (admittedly not very long so far) where there was a need to store even one terabyte of data. But in case of TigerBeetle, from skimming through the video, it appears they offload the main bulk of data to a "remote storage."
Oh I am so with this. The "why" I posed was in context of TigerBeetle's design choices to solve for high-contention OLAP.
Sorry, I framed my question loosely --- too much implicit context.
Me personally, I'm learning from the lesson of TigerBeetle and others, and just using SQLite for my multi-ten-gigabyte ambitions :D
You can have multiple petabytes on a single beefy server with excellent performance characteristics. This is already a thing.
The part of the database that doesn't scale with storage density is cache replacement algorithms, as a matter of theory. There are some problems a cache can't solve. These degrade long before you get to petabytes. If you replace cache replacement architectures with cache admission architectures then things scale just fine.
The main reason we use cache replacement architectures is that they are simple to design and understand. They do have fundamental limits though. Right tool for the job and all that.
Kubernetes is not just for scaling, it's a way to standardize all ops.
Boot it up again. You'll still have higher availability than AWS, GitHub, OpenAI, Anthropic, and many others.
> Where do you think those object storage live exactly?
On a RAID5 array with hot-swappable disks, of course.
(Edit to add: this is just a comment on Kubernetes being invoked whenever someone talks about scalability; I have massive respect for what the TigerBeetle folks are doing)
Me too. Why did you have to add this edit though? Is there anything that suggests either of us disrespects the TigerBeetle folks? I swear, I'm going crazy.