TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning
arxiv.org
arxiv.org
> stretched our ML supercomputer scale .. to 4096 TPU v4 nodes
> The Google tradition is to write retrospective papers ... TPU v4s and A100s deployed in 2020 and both use 7nm technology
> The appropriate H100 match would be a successor to TPU v4 deployed in a similar time frame and technology (e.g., in 2023 and 4 nm).
> TPU v4 supercomputers [are] the workhorses of large language models (LLMs) like LaMDA, MUM, and PaLM]. These features allowed the 540B parameter PaLM model to sustain a remarkable 57.8% of the peak hardware floating point performance over 50 days while training on TPU v4 supercomputers
> Google has deployed dozens of TPU v4 supercomputers for both internal use and for external use via Google Cloud
> Moreover, the large size of the TPU v4 supercomputer and its reliance on OCSes looks prescient given that the design began two years before the paper was published that has stoked the enthusiasm for LLMs
It also has the advantage that when they move to multi-wavelenght, its performance will greatly exceed electrical packet-switching networks.
If you think of packets like snakes going through a network, a layer 1 switching network creates tunnels for the snakes that you choose ahead of time (and can reconfigure whenever you want). A packet switched network creates tunnels that are chosen by the snakes. If you run a packet switched network, you can do everything you do with a layer 1 switched network by simply restricting which peers you send data to. On a hardware level, you need to convert from optics to electricity to do this, but you don't strictly need to do any buffering (the use of large switch buffers on Ethernet switches is because of Ethernet, not because of packet switching). Low-latency switches don't buffer unless they need to, and basically just read the header as it's coming in to choose a route for the packet.
EDR Infiniband networks could certainly handle TPU v4 levels of bandwidth in a packet-switched fashion (at the time when TPU v4 was being built and deployed), particularly when the packets are doing something as tame as going around a torus. It also gives you the flexibility to do other things, though.
It certainly raises the complexity of the system, but I assume sometime around TPU v6 or v7, Google will rediscover packet switching for inter-TPU links.
40 year old patent. But I'm thrilled to see it applied at scale to reconfigurable supercomputers.
[1] Jupiter Evolving: Transforming Google’s Datacenter Network via Optical Circuit Switches and Software-Defined Networking. https://research.google/pubs/pub51587/
EDIT: Ahhh a bit below on the page it is written "13.5 million AI-optimized cores" and there it's plural. So it was probably just a mistake.
e.g.
"He commanded a ten thousand man army." (not men)
"Andromeda, a trillion star galaxy, is 2.5 million lightyears away." (not stars)
etc
> The Edge TPU ... supports only TensorFlow Lite models that are fully 8-bit quantized and then compiled specifically for the Edge TPU.
They even wrote "The products offered by Google are unrelated to the products offered under the CORAL trademarks" in the page footer.
Grayskull is supposed to be A100 performance for $1000, with some cool features (horizontally scalable by plugging them to each other over ethernet, C++ programmable, sparse computation, etc.)
Over the last couple weeks, I've seriously considered purchasing AMD Instinct MI50 (32GB HBM2, 1 Gbit/s) card which goes for under a $1000. I know it lacks Tensor Cores, and can only offer 53 TOPS which sounds silly compared to Grayskull's 600 TOPS. However, isn't it the case that for the most part these cores are idle, & waiting for memory still? At any rate, you're not going be able to run, say Llama 30B which won't fit already into 16GB but must fit comfortably in a 32GB system— on a single card. But perhaps most importantly, amdgpu driver is actually open source and allows PCIe passthrough unlike other vendors so considering all of the above, it almost seems like a no-brainer for a trusted computing setup.
I wonder if my logic is correct.
MI50: https://www.amd.com/en/products/professional-graphics/instin...
Transformers naturally have a big part of the network that is unused on any particular token flowing through the model. You could see that by how little RAM ended up being used in llama.cpp when they moved to mmaping the model.
So my understanding is that the tenstorrent cards are drastically more efficient, even on "dense" models like transformers because of the sparsity of any specific forward pass.
Also: I wouldn't bet on AMD accelerators for ML. They've disappointed every time. I would trust Jim Keller, whose every project in the last decade ended up being impactful.
I think the internet access is just to download the driver. It's not some sort of DRM setup where it needs to be always-online.
Presumably the architecture keeps the model in CPU RAM and shuffles it dynamically to the PCIe card network? I'm guessing here.
Whatever quirks come out of their hardware, I want to keep an eye on it.
I also think that comparing the current generation of models, which are built and trained to maximize GPU or TPU bandwidth, could be improved if someone architected models to maximize their model to greyskull advantages. Given PyTorch runs on it, I don't think it'd be too hard to do.
From what I've read that was just an error in reading memory consumption after switching to the mmap version and not more memory efficient at all in the end.
The author of the mmap patch chimes in here:
Dense transformers (GPT-3 AFAIK is dense) don't.
I spoke with a PM at TT and he told me an important idea is that we're spending a lot of electricity multiplying things with zeros.
It’s goal is only to keep the monopoly: appear bening, keep tech advance in house, share it’s spying network with government so the government don’t regulate them, win-win.
they put out way more R&D research than other big firms. GPT-4 would most likely not exist today without them.
So you don't have to be frustrated anymore.