Patterns of Distributed Systems (2022)
martinfowler.com
martinfowler.com
Sentences like this will make me never regret to moving my infrastructure to bare-metal. My clocks are synchronized down to several nano-seconds, with leap-second skew and all kinds of shiny things. It literally took a day to set up and a blessing from an ISP in the same datacenter to use their clock sources (GPS + PTP). All the other servers are synchronized to that one via Chrony.
Seems like poor engineering practice.
That's why Google built True Time, which provides physical time guarantee of [min_real_timestamp, max_real_timestamp] for each timestamp instant. You can easily know the ordering of 2 events by comparing the bounds of their timestamps as long as the bounds do not overlap. In order to achieve that, Google try to keep the bound as small as possible, using the most accurate clocks they can find: atomic and GPS clocks.
logical causality does not represent poor engineering practice :)
if you stick to a single source of truth - only one machine's time is used as a source of truth - then the problem disappears.
for example instead of using java/your-language's time() function (which could be out of sync across different app nodes) just use database's internal CURRENT_TIMESTAMP() when writing to db.
another alternative is compare timestamps with up to 1 minute/hour precision, if you carry over time from one machine to another. That way you have a a buffer of time for different machines to synchronize clocks over NTP
but that's a sort of trivial base case -- the interesting bit is when you can't make that kind of simplifying assumption
There are ways around this but they are restrictive or come at the cost of increased latency. Sometimes those are acceptable trade offs and sometimes they are not.
if you use a single source of truth for clocks (simplest example is use RDBMS's current_timestamp() instead of your programming language's time() function), and the problem disappears
Now two operations come, one adding $300, other one withdrawing $400. What the result would be, depending on thd order of operations?
Not everyone has the luxury of being able to procure and install hardware and/or run an antenna to someplace with gpc reception.
For a small web app, fine, but if you're running enterprise level software processing billions of DB transactions per day, clocks just don't cut it.
Race conditions are mitigated, not by clocks, but by other logics. The clock was just something done after frustrations in reading distributed logs and seeing them out of order. Logs are basically never out of order any more and there is sanity.
ntp can fail, chrony can fail, system clocks can always drift undetectably
you can treat the system clock as an optimistic guess at the time, but it's never a reliable way to order anything across different machines
node clocks are unreliable by definition, it's a fundamental invariant of distributed systems
Node clocks can be plenty reliable, but like any other hardware, sometimes they get defects.
the A->B link is under DDoS or whatever and delivers packets with 10s latency
the A->C link is faulty and has 50% packet loss
the A->{D,E,F} links are perfectly healthy
node B has one view of A's skew which is pretty bad, node C has a different view which is also pretty bad for different reasons, and nodes D E and F have a totally different view which is basically perfect
you literally cannot "detect skew" in a way that's reliable and actionable
issues are not a function of the node, they're a function of everything between the node and the observer, and are different for each observer
even if clocks were perfectly accurate, there is no such thing as a single consistent time across a distributed system. two events arriving at two nodes at precisely the same moment require some amount of time to be communicated to other nodes in the system, that time is a function of the speed of light, the "light cone" defines a physical limit to the propagation of information
The clock sync is just to keep human-readable logs in order for debugging. It’s ok if it is sometimes out of order, though in practice, it never is.
Interesting excerpt from [0]:
> Finally, it seems that even the weak IQ-job performance correlations usually reported in the United States and Europe are not universal. For example, Byington and Felps (2010) found that IQ correlations with job performance are “substantially weaker” in other parts of the world, including China and the Middle East, where performances in school and work are more attributed to motivation and effort than cognitive ability.
[0]: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4557354/ [1]: https://www.academia.edu/download/50754745/The_Interactive_E...
Sorry about the PDF link to [1]. The APA link has a paywall otherwise I'd link there.
On lower levels leetcode style stuff gives more quantifiable signals per minute/session.
Plus, do you really want to fully explain something like paxos or raft in interview context?
My personal pet peeves are
1. “event sourcing” and derivatives. It really attracts people who love to talk but never built anything large using it
2. Adepts of uncle martin
Plus how do you expect to get good signals from a leetcode whiteboard interview for someone who spends most of their time designing systems and only writes code when it gets to the point where it's faster and less frustrating to pair program vs explaining what needs to get implemented?
To clarify, I don't have a good answer: I still participate in leetcode style interviews (though system design is another component) - but although I sing the song and can't come up with anything better, I don't think it's the best way to go
In my experience things like publications, online code repositories, and facts are more than irrelevant but not much more because people don’t know how to independently evaluate these things. Worse, attempting to evaluate such only exposes the insecurity of people not qualified to be there in the first place.
Far more important are numbers of NPM downloads and GitHub stars. Popularity is an external validation any idiot can understand. But popularity is also something that you can bullshit, so just play it safe treat everyone as a junior developer and leet code them out of consideration.
And intentionally excluding candidates is the whole point of designing an interview process, it's lazy to lift your arms up and declare that they're all biased.
[1] Starting to do math with non-base-10 numbers was already a pass, regardless of the number you reach, you'd normally use a computer for that. But it really isn't too hard to do in your head directly, for anyone who's dealt with binary data.
If you're coming from the position that one is
"understand and assess and write and communicate a deep set of domain-specific knowledge"
and the other is
"on the fly recursive algo regurgitation"
then it will be hard to change your mind about any of this.
---
You could have just as easily called system design interviews regurgitation and coding interviews "communicating deep domain-specific knowledge".
Distributed algorithm > Standard problems: https://en.wikipedia.org/wiki/Distributed_algorithm#Standard...
Notes from "Ask HN: Learning about distributed systems?" https://news.ycombinator.com/item?id=23932271 ; CAP(?), BSP, Paxos, Raft, Byzantine fault, Consensus (computer science), Category: Distributed computing
"Ask HN: Do you use TLA+?" (2022) https://news.ycombinator.com/item?id=30194993 :
> "Concurrency: The Works of Leslie Lamport" ( https://g.co/kgs/nx1BaB )
Lamport timestamp > Lamport's logical clock in distributed systems: https://en.wikipedia.org/wiki/Lamport_timestamp#Lamport's_lo... :
> In a distributed system, it is not possible in practice to synchronize time across entities (typically thought of as processes) within the system; hence, the entities can use the concept of a logical clock based on the events through which they communicate.
Vector clock: https://en.wikipedia.org/wiki/Vector_clock
> https://westurner.github.io/hnlog/#comment-27442819 :
>> Can there still be side channel attacks in formally verified systems? Can e.g. TLA+ help with that at all?
"Quantum watch and its intrinsic proof of accuracy" (2022) https://journals.aps.org/prresearch/abstract/10.1103/PhysRev... https://www.sciencealert.com/scientists-just-discovered-an-e...
> Paperback $49.99
what.the.hell.
Also, in order to be constructive: the kindle version was published 21-June and thus is available right now. Sadly the kindle preview does not include the table of contents
It’s one of the more easily approached resources on the design of distributed systems, and a good read.
Data center/cloud system clocks can be tightly synchronized now in practice. Still never perfect and race conditions abound.
But that doesn't mean you can't rely on a clock to determine ordering, Google popularized a different approach with TrueTime/Spanner: https://cloud.google.com/spanner/docs/true-time-external-con...
truetime doesn't provide precise timestamps, each timestamp has a "drift" window
timestamps within the same window have no well-defined order, applications have to take this into account when doing e.g. distributed transactions
ordering is a logical property which can be informed by physical timestamps, but those timestamps aren't accurately described as "the basis" of that ordering
You can if you consider "unknown/possibly concurrent" also a valid outcome: for any 2 events A and B, True Time can definitely answers whether A is before B, A is after B, or A is possibly concurrent with B.
The context to my first comment included "timestamps within the same window"
https://learning.oreilly.com/library/view/patterns-of-distri...
Patterns of Distributed Systems (2020) - https://news.ycombinator.com/item?id=26089683 - Feb 2021 (58 comments)