A distributed systems reading list
ferd.ca
ferd.ca
“Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems” by Martin Kleppmann as the more advanced deep dive.
Both books provide timeless conceptual advice. Kleppmann’s description of developing a database by starting from an append-only text file really stuck with me.
Really shows the experience of someone who understands this stuff inside and out (was one of the main people behind Temporal).
From a high-altitude view, that's why splitting a huge database table into smaller partitions is not an automatic performance win. If you have M partitions with N rows each, then a lookup might require O(log M) time to find a partition and O(log N) time to find a row within the partition. But O(log M + log N) = O(log MN) which is what you would get from a single big table with appropriate indexing.
Of course, in the real world constant factors and implementation details matter, so this is just a heuristic. But it seems to run contrary to a lot of novice programmers' intuition that a large DB table must automatically be a slow one.
In addition, Lamport's paper "Time, Clocks, and the Ordering of Events in a Distributed System" [1].
[0] https://www.microsoft.com/en-us/research/wp-content/uploads/...
I know this article is a classic. I studied it at school but I've always found it very hard to understand. Maybe I'm wrong but I have the feeling that relatively few engineers use these formalisms as their mental models when designing distributed systems.
Until you have personally battled with replication lag, real-life impacts of eventual consistency and distributed writes, Data-Intensive Applications feels like a dry theoretical read. If you do come across the book with the scars and lessons, it does open the world up.
Can anyone familiar on the topic suggest a list? Perhaps starting with a "101" item for those that want a general understanding / scratch a curiosity itch and perhaps proceeding to more technical items for those that want to dig deep.
2. A bunch of communicating local VMs (easier with a beefier machine like my current desktop).
3. Mininet (there are other options) to simulate a network environment, can fully control the topology very easily. Lighter weight than (2), more control for simulating different network effects than (1) alone.
Use a small k8's distro (kind, minikube, k3s) and build something that talks amongst itself and is resilient.
If you want to be "optimal, neat, or best practiced" read a book, and get stuck in tutorial hell. If you actually want to learn how to do something, literally go do it. Nobody has ever built anything of value (whether that is financial, intellectual, or emotional) by leetcoding.
The problems they're trying to solve are related to really large distributed systems that fail a lot, and their design decisions are basically a "this is how we worked around that problem."
You can also look for the LISA archives (https://www.usenix.org/publications/loginonline/thirty-five-...). System administrators were the first people that had to deal with large distributed systems at scale, and university system administrators led the charge.
You might want to hunt down the comp.sys.admin archives (I can't remember the newsgroup anymore).
Most of the ideas and issues behind distributed computing are obvious if you think about it. Many of the actual implementation and mitigation of those are not obvious, though.
And there's also the client side of distributed computing, which I don't think is discussed as much.
As an example, exponential backoff is one of the go-to techniques for clients when the servers are under load. Unfortunately that doesn't really work IRL, because instead of spreading the load you get waves of load coming back over and over. Likewise on the server side you have problems with peak load.
https://www.amazon.com/Designing-Data-Intensive-Applications...
Maybe some Internet of Things applications would provide a good avenue for some distributed systems exploration?
The cheapest way is to pick up an old ThinkStation (or other tower), load it up with 128GB (or more) of ram and install ESXI on it. That's a perfectly good baseline, and you can run about 30 4gb linux VMs on it.
Ideally you'd have a bit less than 1 core per VM, just so it's a bit slow. Lots of people assume your nodes are quick, but in real life they may not be. And really, most of the time your machines won't be doing squat.
You might want to have SSDs in there too, because ESXI doesn't have RAID capability (or at least mine didn't). I don't think you can get a cloud device that uses spinning disks anymore, and you wouldn't use it in real life anyway.
A 2tb drive is cheap these days, or just slap all those old small SSDs in there. Everyone has a bunch of those small SSDs left over, and they're perfect.
1. Write any stateful program
2. Now look at every single LOC and imagine what happens to the system if the service crashes before executing the next LOC. Then modify the system to deal with those scenarios.Yeah, people miss this. If your app interacts with another app - bam distributed.
"Singularity" systems are an abstraction afforded to us by the grace of the hardware we run them on. If you start pushing their performance hard enough, however, you inevitably get distributed behavior.
This is also a good potential career reason to try to make software which is as performnt as possible - you'll get all the tasty edge cases and complexity war stories to talk about.
By designing twitter in a 45 min interview.
Some people are just up front about it - I've read a lot, and practiced the best I can, but am looking for some real world experience to marry that too.
It's written in Go, so it'll help if you are familiar with Go. But the code is not difficult to understand even if you don't.
I have my current thing set up to create VMs from a downloaded cloud-VM image with minimal updates (add my ssh pubkey and install Python), and then use Ansible for everything further.
Along with "Mastering Bitcoin: Programming the Open Blockchain Book" by Andreas Antonopoulos
Sometimes I only finish half the paper, but damned if I haven't learned a lot.
Disclaimer: I can could never go through and systematically work through a giant list like this. If you know yourself and you can, this may be more effective.