I haven't found many good sources for this kind of information. If you are aware of any, please cite them in a comment below.
I haven't found many good sources for this kind of information. If you are aware of any, please cite them in a comment below.
Queuing theory seems trivial and easy how it is introduced, but it has many open questions.
Performance metrics for a system with random arrival times, independent service times, with k servers (M/G/k) is still an open question as an example.
https://www.sciencedirect.com/science/article/pii/S089571770...
There are actually lots of open problems in queuing theory that one wouldn't expect.
The real world is way more complicated...
You can think about each host as its own markov chain: they may be serving, have hidden or diagnosed problems at various confidence levels, be on route to various remedial processes (reboot/reinstall/repairs I cited above, to simplify), require software/firmware changes, scheduled or opportunistic diagnostics (e.g. periodic deep-screening for otherwise silent problems).
Repair workflows are even more complicated and depend on the specific of the fault, and connection of components (e.g. a system diagnosed with a missing GPU may be because a cable or an interposer card is not seated correctly). Also, repair time is modulated by work shifts and some details of the logistics in datacenters.
Parts can become bottlenecks too. I remember one time in the early TPUv3 days: we had delays in recovering from a large incident because of fault-positive diagnoses of a $4 fan, that was widely used in all systems.
Add that nowadays systems are not one host and some attached cards, but have multiple nodes: e.g. the simplest combination is a main compute node and the smart nic / IPU, but some systems can be a lot more complicated.
So these alone are a hierarchical markov chains, with inhomogeneous arrival times and service times that are themselves time-dependent functions. There are also a lot of long-term memory effects, ensemble average is not necessarily the the time average. Chaotic behavior, in the mathematical sense, is common.
Systems are built with field replaceable units (FRU): e.g. in some generations, you can't swap one GPU in a server that has 8, you have the swap a whole block of them. You can choose when to repair, to maximize TCO/$ based on usage patterns (how many users want all GPUs vs. a smaller number).
Some systems (e.g. TPU pods) have links between accelerator trays, both within a physical rack and across them. So the usefulness of the neighbor system is reduced while a host/tray is being services, and you can have the equivalent of deadlock/livelock in repair dependencies.
Cluster scheduling and management is designed to maximize service levels, minimizing disruptions. This also implies that some disruptions (e.g. when you do some repairs) is driven by what workloads you run. Workloads are power-law distributed in size, so you have small-world network dynamics and the potential for supercritical behavior: both in disruptions (picture jobs preempting other jobs as a graph) but also in risks (add to the previous picture, one job that trigger an hardware problem).
Multiply these by several thousands to get the size of a datacenter cluster. Add supporting compute+storage, networking, power and thermal constraints. Multiply these by the number of clusters (Some hyperscalers have global scheduling systems, that make it possible to see the whole ML fleet as one). Add rare events, because at this scale you start to think about utility and electrical grid failures.
Figuring out what to do in control systems for the current ML fleet is one problem. Simulating what kind of service the future systems a few generations of hardware, software, datacenter design down the line should provide, so you can define what you can offer, influence the designs and make the right investments... it's more complicated. Both of these two problems are my current job.
Meta had a few good YouTube videos about the problems of dealing with this many GPUs at scale.
[0] https://www.youtube.com/playlist?list=PLBnLThDtSXOw_kePWy3CS...
[1]. https://techcommunity.microsoft.com/t5/microsoft-mechanics-b...
Solving these problems has basically been by job for the past 6 years, for Google's TPU systems and some of the GPU systems (the non-Cloud ones). It is a pity that, after the pandemic, it has been impossible to give similar conference talks at my company (at least if you are employed in Europe, due to restrictions on travel outside one's country).
The Meta sessions are very interesting. I wasn't aware of their work on Arcadia; Nvidia has also their own systems (Nvidia Air / Omniverse) in this area.
Meta also publishes a number of papers/blogs/OSS projects on their engineering site [2]
James Hamilton of AWS gives a talk most years on their infrastructure. Worth watching multiple years [3].
[1] https://youtu.be/69PrhWQorEM?si=u7vh_Um6SQNoyeFH
[2] https://engineering.fb.com/category/data-center-engineering/