Cerebras’ new monster AI chip adds 1.4T transistors
spectrum.ieee.org
spectrum.ieee.org
The philosophy here seems to be “if we build it, they’ll buy it.” But suppose you wanted to train a gpt model with this specialized hardware. That means you’re looking at two months of R&D minimum to get everything rewritten, running, tested, trained, and with an inferencing pipeline to generate samples.
And that’s just for gpt — you lose all the other libraries people have written. This matters more in GAN training, since for example you can find someone else’s FID implementation and drop it in without too much hassle. But with this specialized chip, you’d have to write it from scratch.
We had a similar situation in gamedev circa 2003-2009. Practically every year there was a new GPU, which boasted similar architectural improvements. But, for all its flaws, GL made these improvements “drop-in” —- just opt in to the new extension, and keep writing your gl code as you have been.
Ditto for direct3d, except they took the attitude of “limit to a specific API, not arbitrary extensions.” (Pixel shader 2.0 was an awesome upgrade from 1.1.)
AI has no such standards, and it hurts. The M1 GPU in my new Air is supposedly ready to do AI training. Imagine my surprise when I loaded up tensorflow and saw that it doesn’t support any GPU devices whatsoever. They seem to transparently rewrite the cpu ops to run on the gpu automatically, which isn’t the expected behavior.
So I dig into Apple’s actual api for doing training, and holy cow, that looks miserable to write in swift. I like how much control it gives you over allocation patterns, but I can’t imagine trying to do serious work in it on a daily basis.
What we need is a unified API that can easily support multiple backends — something like “pytorch, but just enough pytorch to trick everybody” since supporting the full api seems to be beyond hardware vendors’ capabilities at the moment. (Lookin’ at you, google. Love ya though.)
I'm now bullish on elixir-nx because I think there's an outside shot they will get it right.
Python is a lost cause.
WRT Nx, my biggest question is how they'll crack the problem of still needing big balls of C++ and the shims everywhere to get acceleration. Creating a compiler that generates efficient GPU or other accelerator code is a massive research project with no clear winners, never mind the challenge of reconciling the very mutation-heavy needs of GPU compute with a mostly immutable language model.
Supposedly Cerebras is already profitable, so it's hardly a situation where they are building something and hoping people buy it eventually.
> That means you’re looking at two months of R&D minimum to get everything rewritten, running, tested, trained, and with an inferencing pipeline to generate samples.
Again, based on the companies representations, Cebebras transparently supports Pytorch and Tensorflow, only requiring a few lines of changed code.
Source: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-... (Dr. Cutress's video on TechTechPotato is also good).
There is a very specific test for “supports pytorch / tendorflow”: show an MLPerf imagenet resnet benchmark. It’s impossible to fake that. If you come anywhere close to TPUs in tensorflow, people (like me) will leap: https://mlcommons.org/en/training-normal-06/
Till then, it’s a “proof, please” type of situation.
I have high hope for Graphcore. They stated "We will be participating in MLPerf in 2021, starting with the first training submission in the spring".
https://www.graphcore.ai/posts/graphcore-sets-new-ai-perform...
(I would ignore all benchmarks in that post except the sentence quoted above however.)
Cerebras itself commented that DL is not their first use case. They mentioned fluid dynamics instead. (I can find the reference if someone is interested).
Cerebras also has a weird memory/compute ratio, so very hard for some uses.
Somewhat off-topic. If u have standard keras/TF training code, the GPU to TPU is a smooth-ish transition. No?
A re-write will squeeze another 20% but not 200%?
There is also webGPU which is experimental on Safari and Chrome. But that doesn’t touch the ML accelerators on the M1.
> “if we build it, they’ll buy it.”
This is literally how any new thing is invented and commercialized.
Sometimes, though, it would be nice to simply focus on doing interesting ML work rather than wrestling with the intricacies of tensor slicing and memory access pattern optimization (to say nothing of the cursed inability to manually manage memory, resulting in explosions of allocations in the forward pass for unclear reasons).
Let’s put it this way. If you contract me to implement a full gpt model on this new hardware, two months full time work would be my minimum estimate. That’s 40 hours a week of focused effort, with no breaks and no other projects. I don’t know about you, but two months with nothing to show can demoralize most teams I’ve worked with.
The point was, we’re early in ML’s lifecycle. And, like React for webdev, people are slow to change their habits — for better or worse, pytorch and tensorflow are the APIs people think in. So if you want your hardware to be widely adopted, you need tooling that supports the workflows people have spent months learning.
Jax is on the horizon too. Supposedly they’re launching something soon that might tip the scale in their favor. Perhaps an ambitious hardware vendor could capture the future market by preemptively implementing Jax support. But perhaps not: at that point you’d be competing toe to toe with Google’s TPU offering, since Jax is Google’s dogfood.
Right now my money is on TPUs, partly because of their fantastic support staff. But maybe some other company will come along and offer a better integrated cloud experience.
But you’re right, I should write up something. That guide is pretty good, but it doesn’t walk you through anything specific; it shows you a map, but doesn’t take you on a trip, so to speak.
Some tips:
Use tf.name_scopes! You probably use them for variables, but there’s a different one for operations. If you make an @op_scope decorator, you can use it on all your ML functions and immediately get lots of insights as to where the XLA ops end up mapping to in your source code.
As much as it pains me to say this, avoid tensorflow 2 style code like the plague. Pretend that if you use eager execution, someone will jump out of the bushes and shoot you. I technically use TF2.4 now, but it’s still session / graph-based, not the new tf.function magic. The new magic pipeline seems to be slower, harder to use, harder to debug, and very likely to explode when you do anything even slightly different than the tutorial examples. YMMV, and maybe things are better now (or in a future release).
The profiler is magical; leverage it whenever you can. My workflow is to start up a TPU run, then ssh into my server and fire up a Tensorboard to that model dir. Then I manually navigate to <tensorboard_url>/#profile (because the “profile” button doesn’t seem to show up in the menu anymore) which then lets you “capture profile”.
Make sure your TPU version matches your Tensorboard version. At this point we use TPU version 2.4.0, Tensorboard 2.4.1. If you get mysterious errors, this is likely the reason things are going wrong. Even TPU version 2.3.0 wasn’t enough.
Happily, the new profiling tooling is totally badass. The trace viewer is great, the op profiling is ok (though I wish it would show me all the damn ops, instead of “helpfully” hiding all but the top N ops), and the memory viewer is incredible. You can see exactly where in your pipeline is causing “peak memory usage”, what the peak is (down to the kilobyte), and have at least some idea of what’s causing it.
It’s not effortless though. The XLA fusion ops sort of make it harder to track down what’s doing what. (TF compiler is very powerful, but the trade off is that you almost never have manual control over memory usage, which can be frustrating).
All in all (or all-to-all, ha) it’s a lot of fun if you like seeing expensive hardware go brrrr, as I do. It makes it all worth it when your loss drops from 11 to 3 overnight on a 430m parameter gpt model. :)
Detailed Question 1: do you have examples of fusion making things hard? Is there a way to nudge the compiler to not fuse or create a symbol table tracking fusion? Does fusion cause issues with the name_scopes?
Detailed Question 2: isn't Pytorch eager execution? Do you know how it compares to Tensorflow's eager execution?
General Request: I'm in a more theoretical position, writing papers on programming languages for accelerators https://aetherling.org/ and TAing courses on accelerators http://cs149.stanford.edu/fall20. So, I'm excited to see people's practical experiences using these accelerators in industry. It would be very enlightening (if you have time) to write up a comparison of tuning a model for an A100's tensor cores vs a TPU. This seems like the key trade-off in comparing architectures.
Here's a tensorboard URL that will probably stop working within a few days. http://bulma.tensorfork.com:31337/#profile
You can view the memory profiler by using the dropdown menus on the left. Here's a particularly chonky CrossReplicaSum: https://i.imgur.com/CcdJzLj.png
The game here is to keep that number at the top -- peak memory usage -- below 15GB. In practice, TPUv3-8's run out of memory at around 14.5GB, which immediately crashes (and hence you can't profile it). So we're always trying to get as close to 15GB as possible.
The first thing you immediately notice is that real-life training runs are very spiky. Different parts of the pipeline end up allocating wildly different amounts of memory. There's almost no such thing as a constant memory usage pipeline (which I was dismayed to discover).
In this profiling run, you can see that there's a big ass-spike at ~4000 on the X axis. The green bar marks the lifetime of the operation causing the highest peak memory usage. Different operations depend on each other, forming a chain of allocations. a + b takes 'a' and 'b' as inputs, and any temporary tensor reachable by either 'a' or 'b' cannot be freed until a + b is finished executing. Ditto for all other operations.
So you see, it's easy to accidentally build a "tower" of allocations, rather than a flat line. Thus, your total model parameter count is severely limited compared to what it could be, since in this situation the only way to reduce memory usage (without rewriting the code) is to scale down the model params.
Hovering over the big ass-orange allocation, we see that the shape is -- gosh, tensorboard is infuriating sometimes. I tried to copy-paste the shape, but whenever I move the mouse off of the allocation, the info on the left vanishes. Anyway, the shape is F32[32,2048,1,12608][1,3,0,2]. It means the cross replica sum is happening across TPU cores 1, 3, 0, and 2; it's a float32 sum; the batch size is 32; the hidden dimension is 2048, and the vocab dimension is 12,608. Since it's across four cores, multiply that dim by 4, and the total vocab size is 50,432, which is exactly right for a GPT model (https://nv-adlr.github.io/MegatronLM has details).
So right away, we can see that (a) the non-peak memory usage is around 4GB or so, and (b) the peak mem usage of the spike is around 12GB. That means if we eliminate the spike, we can scale up our model by more than 3x, if usage scales linearly. (Sometimes you get lucky and it's linear, other times something is superlinear. It's more or less linear in my experience.)
So how do we eliminate the spike? Heck if I know how the Google pros do it, but my way of doing it is to unstack along the batch dimension and perform each operation sequentially.
In other words, the total memory usage here is O(32 * 2048 * 12,608) which is quite hefty. By unstacking along the batch dimension, you get 32 tensors, each of size 2048 by 12,608. Therefore, if you do each operation sequentially, the temporary buffer is now O(2048 * 12,608), giving us a 32x savings.
Is this slower? Surprisingly, more often than not, it's as fast or faster. The reason is subtle: slowdowns occur due to memory bandwidth and network bandwidth. As long as the unstack is strictly a memory bandwidth effect, then it's just as fast, because you're trading CPU cycles for memory -- and you have tons of CPU cycles here, since it's a TPU core. (The TPU core utilization in our experience is always around 30%, and we've never seen it higher than 65%.) So you should always, always make this trade whenever possible.
Network traffic is trickier. This is a cross-replica sum, which means it's sending the tensors across the network to each TPU core. The TPU cores are connected via a high speed interconnect nexus thingie, but like all bottlenecks, this one has a limit. It's a very high limit, but it's not endless.
The only solution I've found is to think of an idea and then test that idea. Reasoning from first principles almost never works for me. I've seen others solve problems by reasoning from first principles, so it's possible that I'm simply stupid. But I find it's much more effective to try as many ideas as possible, as quickly as possible. You often end up surprised.
I'll type more stuff later if I feel like it, or you can ask more followup questions. Feel free to DM me on twitter if you'd like to chat in realtime sometime.
For AD, I am bullish for Enzyme, which does AD on LLVM IR, avoiding deep compiler integration: https://enzyme.mit.edu/
An API is always going to be necessarily behind state-of-the-art because often research depends on inventing things that don’t fit within an existing API.
A PyTorch/TensorFlow-like API is always going to be massive, complicated, and hard to port to new hardware targets. Additionally, the complexity and reliance on C alone will make integrating exotic new concepts like say, differential equation solvers, extremely laborious.
Android standardized NNAPI and on-device inference is in pretty good shape. As you said, training is different matter.
>A key to the design is the custom graph compiler, that takes pyTorch or TensorFlow and maps each layer to a physical part of the chip, allowing for asynchronous compute as the data flows through.
https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
You mean one of the most rapid periods of graphical improvement in history? Growth was just too breakneck to standardize for a bit.
If AI acceleration chips aren't able to offer similar API standardisation/stability then it's not the same as what happened with GPUs, which was sillysaurusx's point.
Their compiler is closed source, so it's almost impossible to tell what's going on. But, when I connect using tf.Session(), then .list_devices() shows only CPU available.
However, when I enable the TF_MLC_LOGGING=1 magic variable, it does seem to be printing out messages that indicates it's doing some kind of graph substitution under the hood. Therefore, I assume that this is the intended usage mode.
In other words, there seems to be zero difference between the "CPU" and the "GPU". Normally you can say "Do this on the CPU" while "do that on the GPU." But not with this.
Hopefully they'll open source the code sometime this century so that it's clearer what the heck it's doing. For now, though, it's reasonably fast in whatever this "CPU" mode is -- I only need to run unit tests on my laptop anyway, since all training happens on TPUs. So I ended up happy.
(For the first day or so, I was panicking that I was going to have no working tensorflow whatsoever on my M1 laptop, which would've necessitated a swift return + substitution.)
If you're looking for standardized hardware and software for AI, Apple is the wrong platform to be on. I'm fairly confident it won't be happening there. (I don't say that to be a hater: see the direction they've gone re: GPU APIs... it's just Apple being Apple)
XLA has already proven its value by allowing PyTorch to run on TPUs (shittily, but that appears to be more of a VM/GCP infra problem than an XLA problem). The work done for TPUs (and to a lesser extent for GPU optimization) has started to expose some of the major issues and so work can start on addressing them (the cost of dynamic XLA compilation as tensor shapes change and how lots of important code assumes that accelerator-to-CPU communication isn't tooo expensive, but it's is a huge issue when trying to compile the graph into machine-specific code with XLA or similar to because it forces you to only be able to compile small subgraphs).
It's early, but the rise of a really effective IR in XLA combined with the huge amount of resources that Google/NVIDIA can pour into XLA makes me very bullish on purpose-built hardware for AI training. It will take a while I admit.
I agree with you, but I think we differ on our timetables. I am bearish for the next two years, at which point I’ll awaken from my slumber and become a flaming bull. (It helps to remember that “we overestimate the impact of years, but underestimate the impact of decades.” I try to plan accordingly.)
In other words, if you’re bullish that two years from now we’ll start seeing portability implemented in the field across various HPC chips, then we fully agree. But that’s also a glacial pace; GPT-2 changed the world almost two years ago now, and DALL-E seems to be the next frontier for doing interesting generative work. So, we’ll split the difference and say that the bears and bulls will meet in two years for a deep learning hackathon. As a bonus, the pandemic will be over by then, so it can be an in-person meetup.
Updated my profile. I've been working on DL training platforms and distributed training benchmarking for a bit so I've gotten a nice view into the GPU/TPU battle.
Shameless plug: you should check out the open-source training platform we are building, Determined[1]. One of the goals is to take our hard-earned expertise on training infrastructure and build a tool where people don't need to have that infrastructure expertise. We don't support TPUs, partially because a lack of demand/TPU availability, and partially because our PyTorch TPU experiments were so unimpressive.
[1] GH: https://github.com/determined-ai/determined, Slack: https://join.slack.com/t/determined-community/shared_invite/...
Yes.
Longer term, new hardware will also make it practical to train large models in a fully parallelized, fully distributed manner -- i.e., without having to backpropagate gradients, which requires a lot of complex bookkeeping and plumbing for distributed training.
Recent progress suggests this will happen. See, for example:
https://arxiv.org/abs/2006.04182
https://arxiv.org/abs/2103.03725
https://arxiv.org/abs/2010.01047
I for one am excited to see what happens over the next decade as it becomes trivial to train/use models with 1K, 1M, or 1B times more dense connections than present state-of-the-art models.
For a hobbiest, sure, that's a problem.
But for a big company with a big ML research team already, that isn't an issue - they just assign a few people to work on it, and in a few months it's done. If you're running any model at scale anyway you probably want to rewrite everything to make it run efficiently on your hardware.
EDIT: Or GPU, or whatever it is.
I wonder what their software stack looks like. Can they support the sort of virtualization and sharing you'd want to keep this expensive beast fully utilized 24/7?
I like it!
In previous articles they've gone into some detail about how they deal with reticle limits, jumping over the scribe line area, and other n stuff. Between that, chiplets, HBM-style die stacks, etc... the developments here have been more interesting than I expected.
> Cerebras achieves 100% yield by designing a system in which any manufacturing defect can be bypassed – initially Cerebras had 1.5% extra cores to allow for defects, but we’ve since been told this was way too much as TSMC's process is so mature.
Also, there is no reason they cant have some redundancy throughout the design so you can fuse off bad parts. It all really depends on the nature of the anticipated vs actual defects, which is an extraordinarily deep rabbit hole to climb into.
It's hard to estimate a per-unit cost. But suffice to say it would cost similar to other datacenter compute solutions on a performance/$ level.
This last decade people soldered resistors onto cheaper NVIDIA cards, to make them behave as more expensive NVIDIA cards: https://www.eevblog.com/forum/general-computing/hacking-nvid...
So chipmakers have an incentive to bury the process in the chip and make it irreversible.
1. piece by piece
2. on die test circuits
"When we spoke to Cerebras last year, the company stated that they already had orders in the ‘strong double digits’."
And they cost 2- 2.5 million each!
https://www.anandtech.com/show/15838/cerebras-wafer-scale-en...
"The CEO Andrew Feldman tells me that as a company they are already profitable, with dozens of customers already with CS-1 deployed and a number more already trialling CS-2 remotely as they bring up the commercial systems".
Quotes from https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
I'm sure there's a compiler and low level primitives to really get the maximum performance out of it, but the trade-off maybe worth it in many cases to do it using an abstraction like the linear algebra approach.
The article mentions a year of engineering went into dealing with the entire wafer thermally expanding under load.
[1] https://www.eetimes.com/powering-and-cooling-a-wafer-scale-d...
Under what circumstances does the chip need to access external memory?
What type of communication interfaces does this chip have?
Also, if the chip is the size of a wafer, is it appropriate to call it a Chip?
(more than the IEEE)
This thing pulls 15-20kW of juice!
It may be a very large wafer but dissipating that heat is still very impressive.
That said, I'd be fascinated to see the cooling solution. Is it just a massive copper heatsink & a boatload of airflow? Typical approaches of using heatpipes to expand the heatsink won't really work with something this big after all. Or is it a massive waterblock with multiple inlets/outlets so it can hit up a stack of radiators? How do they get even mounting pressure across that large of an area?
"To solve the 70-year-old problem of wafer-scale, we needed not only to yield a big chip, but to invent new mechanisms for powering, packaging, and cooling it.
The traditional method of powering a chip from its edges creates too much dissipation at a large chip’s center. To prevent this, CS-2’s innovative design delivers power perpendicularly to each core.
To uniformly cool the entire wafer, pumps inside CS-2 move water across the back of the WSE-2, then into a heat exchanger where the internal water is cooled by either cold datacenter water or air."
If you look at the wafer she’s holding at the top, it’s seemingly segmented into a 12x7 grid of roughly chip-sized rectangles. That’s 84 “CPUs” at 200-240 watts each, which is pretty well in line with discrete server CPUs.
The amount of heat coming off this thing must be amazing, though.
The 40GB of SRAM probably has tremendous bandwidth (it could all be updated every few cycles!), but the memory size is very small compared to the amount of compute available.
However, maybe a different way of looking at it is that this chip will allow the training steps on deep learning models to take a fraction of the time as a GPU. Perhaps what takes 1s on a GPU could take 10ms on this chip.
So, this product may be effective at making training happen very fast, but without substantial model size or efficiency gains.
That's still ground breaking -- you can't acheive this result on GPUs. You can't achieve this result by any parallelization or distributed training, either. The large batch sizes in distributed training do not result in the same model or one that generalizes as well.
from: https://www.youtube.com/watch?v=yso2S2Svdlg
@ 25:14
James Wang: "If a model doesn’t fit into a GPU’s HBM, is it smaller when it’s laid out in the Cerebras way relative to your 18 gigabytes?"
Andrew Feldman: "It is — it’s smaller in that we hold different things in memory than they do. One can imagine a model that has more parameters than we can hold — one can posit one, but remember our memory is doing different things. Our memory is basically holding parameters. That’s not what their memory is doing. Their memory is holding the shape of the model, their model is holding the results of the batches. We use memory rather differently. We haven’t found models that we can’t place and train on a chip. We expect them to emerge, that’s why we support clustering of chips and systems, that’s why we do that in whats called a “model parallel” way, where If you put two chips together you get twice the memory capacity. That’s not what you get when you put multiple GPUs together. When you put multiple GPUs together you get two versions of the same amount of memory, you actually don’t get twice the memory. I see you smiling here because you know that’s a problem… …With us if we support 4 billion parameters and you add a second wafer scale engine, now you support 8 billion parameters ,and if you add a third you can support 12 billion. That’s not the way it works with GPUs. With GPUs you just support two chips, each with a few million - tens of millions of parameters."
I'm curious how the problem effectively gets sliced.
> Also, if the chip is the size of a wafer, is it appropriate to call it a Chip?
Good question. I think it is. I mean the word "chip" isn't really that well defined (is HBM one chip?), but given that they sell it as a single unit and you can't really cut it in half I think it's one chip.
It's a really cool picture too.
Bubbles popping aren’t always evaluations of the objective quality of something; it could just be about its assumed value relative to other plausible options. Homes mostly don’t become uninhabitable when a real estate bubble pops.
Is the 20,000 amp number not an error?
Comparison table shows nvidia A100 which has max power consumption of 400W. This has roughly 50 times the transistors and 50 surface area so having 50 times the power consumption 50*400=20000W and same order of magnitude of current seems possible.
Of course 20kA is pretty insane, would be interesting to see just how they feed the power to this thing. Bond wires going to the middle of the chip? And the local regulator must be a beast, or rather lots of beasts.
Well at least on relevance* to active problems.
Incidentally, FPGAs are also insanely expensive at the high end.
According to google calculations, this is just under 6 square feet.
Though that is unimaginably large to me for a silicon chip. I can't comprehend it - the engineering behind it would be incredible.
> 0.49756176 square foot
If we have multiple human experts annotate a NLP task and measure inter-annotator agreement, it will be far from 100%; part of that will be genuine disagreements or fuzzy gray area, but part of the identified differences will be simply obviously wrong answers given by the experts - everyone makes mistakes. The same applies for many other domains - business process automation, data entry, etc; no employee will produce error-free output in a manual process, no matter how simple and unambiguous the task is.
And for simpler tasks the computer can easily make less mistakes than a human - especially if you measure the human reliability not for a few minutes of focus, but for a whole tedious working day.
I'd be glad to learn I'm wrong.
AI has "crossed human expert performance" on extremely narrow NLP/CV tasks.
AI is still light years away from human-level performance.
Those networks didn't match "Detective", a crappy story written by a 12yo.
with AI chips burning 15KW? No chance for the winter in sight. Some chances for AI hell though.