GigaGPT: GPT-3 sized models in 565 lines of code
cerebras.net
cerebras.net
Even with a wafer scale chipset this approach has limits. You eventually will still need to shard to fit more parameters / use different training modalities / etc. I'd look at this more as a proof of concept for the ergonomics of what LLM training can look like when you have access to a much larger compute primitive versus a new state of the art in feature-equivalent clean code.
Disclaimer: I'm a small investor in Cerebras.
That is: a lot of distributed big data processing tasks don't need to be distributed. Perhaps with beefy enough matrix multiplication chips, a lot of "big ML" tasks won't need to be distributed either.
At the end of the day I do believe ergonomics are going to win out. I think that's in large part why pytorch won out over tensorflow and jax; it provided the just-in-time computation that would allow people to more easily find bugs & visualize results without having to `compile()` everything down to a static computation graph. Hardware seems like a natural place for that abstraction layer - but maybe the silver bullet will really be on the software side, since we already have too many "low-RAM" equivalent ML devices in the wild. Cheaper to string things together after the fact vs. shipping net new hardware.
For instance, right now there's a new crop of "linear RNNs" (RWKV, Mamba, retnet, etc.) claiming to be as good as Transformers for language modeling but with two advantages: their compute cost is O(n) instead of O(n²), and they don't need to keep past context in memory.
I don't know if these linear RNNs will actually supplant Transformers, but I do think hardware requirements are likely to change over time.
However, the very next question a researcher will ask once a model fits on one device is “can I make it twice as fast/big if I use two?”
Or more directly, perhaps one should ask such researcher, "if your team was to double in head count, would you do this project twice as fast?".
If everybody's job was just to do dot products all day I'd hope the answer would be yes.
I think the broader point is that the last x0 years of ML research show that more compute is better, both for iteration speed and for resulting performance. Distribution is just the natural outgrowth of that imperative once it reaches the limit of a single device/node. If Cerebras can address models at today's scale on one device, the immediate next step is "what can N of these devices do together to build models at tomorrow's scale"
I still think work on improving single-core/device performance is worthwhile, as distribution will always strictly not-better, and almost always strictly worse, due to coordination costs reducing efficiency. If two Cerberas can be glued together and achieve roughly 2x of their performance, it's still going to be more efficient than achieving equivalent performance from many more regular GPUs. Getting the hardware fast enough so that you need just one device for your problem - that's a special case that will yield extra win.
We don't have overcomplicated distributed training infra. They are pretty much as complicated as needed.
Scaling vertically (having a bigger chip) is very hard. There are tons of tradeoff when making the chip, and it's overall an insanely complex problem. That's why Cerebras is a 8 years old company and yet you would be hard pressed to find anyone using them still.
And even if you give me a Cerebras chip that works perfectly, it will still be much easier for me to buy two of those chips and link them together in distributed training mode, than it will be for Cerebras to build a chip that is 2x the size.
The scale of the current generation of clusters to train models the size of GPT-4 are in the range of 25,000+ GPUs with 80GB of memory each, so no matter your chip size, complicated distributed infra is a necessity. Even assuming everything on Cerebras' marketing page is fully accurate, you would still need to distribute the training over 500+ of those massive chips to replicate a 25k GPU cluster.
Cerebras is trying to show how easy it is to on-board single ICs and demo their pytorch integration.
But yeah, where's the wallclock time comparison?! Surely they did one during development, and surely the Sales team knows (or they do once the article was published), yet not even a hint of what their throughput is like. For Cerebras to be this far and not be plastering benchmarks everywhere is a bad sign. Maybe they're going to just die off like Graphcore.
Hey investor, we want chatgpt 5. The issue is not lines of code.
Lines of code is not strictly speaking a bottleneck for the next generation of models, but it ties with other objectives that are: researcher productivity, hardware efficiency, and model verifiability. GPT-5 might be another case of simply scaling up the existing transformer models, but the next step-function change of model quality are going to involve a lot more R&D about the right architectural primitives to take before scaling it up. And in that case lines of code do matter - because they allow simple concepts to be robustly tested, optimized, and iterated against. Doing that against a 50k monolith is a much harder task.
AFAIK you can just increase the layer parameters of a 1B model to whatever you want? Like, the difference between a 1B and 175B model can be just changing a few numbers, and not adding any LOC at all?
LOC has never been a limitation for large models, it's been the compute+training data required.
Most of the LOC is spent on optimization, and they don't address MoE or anything fancy like that?
They can watch different layers train and find out how to optimize training or quantization, etc.
It feels like they kinda missed the forest for the trees here. The article should have focused on model architecture optimization due to the small LoC and the system having ridiculous training capacity.
Cerebras' point still stands though, even if you can get the LoC count down significantly nowadays, it's still a major PITA to debug those systems, deal with node crashing, tweak the architecture and the data-loading pipeline to have high GPU utilization, optimize network bottlenecks etc. Scaling vertically first like Cerebras is doing surely makes that much easier.
On a tangentially related note, this is imho where OpenAI has built it's moat: training and inference stack that they have refined over the last 6 years. They have good researchers, but so does MS, Google and Meta. But no one else has the ability to train such large models with such ease. Same for the inference stack, being able to run GPT-3.5/4 in prod at the scale at which they are doing it is no joke, and I'm 100% convinced this is why Gemini is still not widely available a year after 3.5 came out.
On windows you can get it via lmstudio.ai for example
Here is ur running on a MacBook M2 Air, we have smaller, more performant models coming
I recommend downloading and running OpenHermes inside LM Studio. https://lmstudio.ai/
In LM Studio, search for OpenHermes. Pick the Q5_K_M version (this is the best quality/speed trade off). Then go to the chat tab.
On the chat tab, set the context length to 4096 (or up to 16k if you want longer context) and set the number of CPU cores you have under "Hardware Settings."
Select the model from the drop down and start chatting!
Do you have recommendations for a different framework for training? Accelerate seems fantastic for scaling up once I need to
Inference is where it falls short really, and solutions like vLLM are much much faster.
Like LangChain, they were at the right place at the right time. That doesn’t make them good.
It's certainly an order of magnitude easier to use something like transformers or diffusers than the original implementations provided by the original model trainers, and has a few good optimizations out of the box.
That's different from LangChain which is complex for the sake of being complex.
I can't speak to the efficacy of it at large scale, though.
There are lots of "just one line of python" type frameworks that are fine if you want to do the one thing in the demo but are more complicated than just writing it yourself if you have to change something.
You can read more here https://docs.cerebras.net/en/latest/wsc/tutorials/custom-opt...
> ... write 20k LOC of complex code ...
Vs the all important:
> ... write _and maintain_ 20k LOC of complex code ...
The answer to the latter, for me, is a hard no.
Furthermore, how important is the breadth of data in the dataset to getting the desired results? I was under the impression that the main reason these LLM work is based on massive data sets.
As such, is there data-breadth metrics to validate whether training on a given dataset is even worthwhile? (ie: avoid sunk cost on a dataset that will yield a poorly performing LLM)
Andrew Karpathy is truly a gem and super grateful he still publishes videos showing his art.
Cerebras showing their distributed architecture on that same piece of code is impressive.
All of AI is search for a god algorithm. An algorithm so simple it could be written on an A4 piece of paper in 12px font size - but with enough data and compute it can more intelligent than entire cities of humans combined.
NanoGPT is a glimpse of that.
In that sense, 565 LoC is a perfectly fair number. It doesn't count PyTorch, numpy, the Python interpreter, or any of the library modules that are imported, but I don't think anyone was touting it as anything more; for instance, Mo Gawdat has said that GPT-4 is probably ~4500 LoC. And, yes, that certainly involves much more infrastructure, and doing that dance of going from GPU to CPU to a completely other node, etc.
I don't see the novelty/interesting bit in this article, personally.