Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
cerebras.net
cerebras.net
The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0].
It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the edge.
But. BUT. The things that we can do on the edge are incredible now. Just imagine a year from now, or two. These earth-shattering models, which seem to be upending a whole industry, will soon have equivalents that run on the edge. Without services spying on your data. Without censorship on what the model can/cannot say. Because it's all local.
When was the last time this happened? There will be players who publish weights for models that are free to use. The moment that torrent magnet link is published, it's out in the wild. And smart people will package them as "one click installers" for people who aren't tech-savvy. This is already happening.
So every time you're amazed by something chat-gpt4 says, remember that soon this will be in your pocket.
[0] the "confetti" idiom brought to you by chat-gpt4.
I thought running something on the edge referred to running it in close network proximity to the user, rather than users having control and running things themselves.
It is kind of cyclical then is not?
By that I mean computers used to be shared and to log into it through a terminal.
Then the PC came around.
Then about 15 years ago Cloud computing became the rage (really an extension or more sophisticated system than the first time shared computers)
Now we're back to local computing. I even see more self hosting and moving away from cloud due to costs.
All that rant is to say is it's interesting.
Side note, getting this AI to be localized as much as possible I imagine will be really useful in the medical industry because it helps alleviate HIPAA requirements.
> By that I mean computers used to be shared and to log into it through a terminal.
> Then the PC came around.
> Then about 15 years ago Cloud computing became the rage (really an extension or more sophisticated system than the first time shared computers)
There's a really neat article called "The Eternal Mainframe"[1] that you might be interested. It explores this idea in greater depth.
---
1. http://www.winestockwebdesign.com/Essays/Eternal_Mainframe.h...
I wonder if the author's perspective has changed with regards to freedom to compute.
Social Media is often used as an example of privacy invasion though I've failed to see why concerns over Facebook handling your private data is worrying when they don't have a product you need to have.
Email on the other hand, is pretty much a necessity today so privacy concerns are vital there imo. Of course you can host your own server whereas you can't host your own Facebook.
What I find more interesting is that in the classic "close network proximity", some parts of the world may not have benefited as much from that trend since the closest nodes of a global delivery network could be several countries away.
[0] https://github.com/ggerganov/llama.cpp/discussions/205
[1] https://medium.com/sort-of-like-a-tech-diary/consumer-ai-is-...
I don't like the connotations this carries. This is almost openly talking about reaching all the way into peoples' hardware to run your software, for your benefit, on them, without their knowledge, consent or control...
What I think is important in this AI Spring is that we make it possible for people to run their own models on their own hardware too, without having to submit anything to a large, centralised model for inference.
Somewhat; its consistent with, e.g., Google’s “Edge TPU” designation for its client-side neural processors.
> I thought running something on the edge referred to running it in close network proximity to the user
Typically, but on the client device is the limit-case of “close network proximity to the user”, so the use is consistent.
[1] from https://www.oblomovka.com/wp/2007/08/ at least
I imagine this "cat out of the bag" situation, the democratization and commodification of powerful technology accessible and affordable to the public, is similar to what's happening with single-board computers and microcontrollers like Raspberry Pi, Arduino, ESP32.
It might be similar to what happened with mobile phones, but there the power was quite restricted. The (mostly) duopoly of iOS and Android, with devices and apps locked down in various ways. Sure we can "jail break" and "root" our phone, but that's not for the general public.
Maybe solar energy production is going through a similar process, with panels and batteries becoming more efficient and affordable every year.
Certainly, it reminds one of the history of personal computers, the way such a powerful general-purpose tool became ubiquitous and local.
These local models are absolutely critical for that to happen though. I'm hitting daily situations where I have to reconsider my use of ChatGPT because what I'm asking would leak very private personal information or somehow trip its morality filter.
Just swapped in a 2TB nvme for a separate /home mount and reinstalled Arch just so I could have enough disk space to test a dozen models locally. I'm so ready!
The speed of recent advances may be due to picking low-hanging fruit.
Soon there may not be much low-hanging fruit left.
My guess: we will have a competent JARVIS(minus the holograms) this year.
Empowering everyone equally, very low barrier of entry. The spread of AI will make some skills become less important and other skills more important. There will be painful social changes. Memorising stuff, being exact, being careful, diligent - all these things can be done by LLM. But also creative recombination of ideas. For now we still hold supremacy in long form content, complex software projects and detecting inconsistencies, but they could be emergent qualities in GPT-5.
Integer-quantized LLaMA can fit on a 4090, and isn't quite human-level. If a fully aligned, reliably non-hallucinary turingrade LLM can run on the next generation 5 series GPUs then expect their market price to quadruple or more.
> No results found for "confetti has left the cannon".
I'm amazed that a "stochastic parrot" can come up with such a beautiful idiom.
Not yet. I'm sure somebody is taking a hard look already.
ME: What's another saying similar to: "The cat is out of the bag" or "The genie is out of the bottle"?
chatgpt: Another similar saying is "the toothpaste is out of the tube." Like the other two expressions, this phrase conveys the idea that something has been revealed or unleashed, and it cannot be undone or put back into its original state.
ME: Can you invent a new phrase that has the same meaning, which is completely new?
chatgpt: Certainly! Here's a new phrase with a similar meaning: "The confetti has left the cannon." This phrase suggests that once something is out in the open or has happened, it's impossible to reverse or contain it, much like confetti that has been shot out of a cannon during a celebration.
ME: What's the name for this kind of expression? Aphorism? Fable?
chatgpt: This kind of expression is known as an idiom. Idioms are phrases or expressions with a figurative meaning that is different from their literal meaning. They are commonly used in everyday language to convey a specific message or sentiment in a more colorful or memorable way.
Right now, their ability to learn is severely limited. And, yet, they outcompete us quite easily in a lot of different tasks.
If humans have some special sauce different from the computer, then it’s crazy that ChatGPT can emulate human writing so well. If humans are also just statistical models, then of course you can throw a big training set at some GPUs and it’ll do the same thing. Why should we be surprised or impressed by idioms?
To explain the other way to your thinking: human intelligence is the same; holy crap they cracked robotic 'human' intelligence, it works exactly the same way.
There are many mentions of confetti cannons on the web, along with explanations of how they work (saying something like confetti shoots out of the cannon). Chat-GPT just picked a random thing (confetti) and completed the pattern "X out of Y" with the thing confetti comes out of.
It's easy. The cereal is out of the box. The helium is out of the balloon. The snow is out of the globe.
And it's exactly the one thing LLMs are trained to do: reproduce patterns of words. They are (perhaps) already better than humans at that one specific skill (another win for AI research) but I don't think it's a sign of general intelligence.
Is Stockfish intelligent?
Is a system with A* pathfinding intelligent?
I would define intelligence as the ability to solve a wide variety of novel problems. A system built to be excellent at a single task may be better than humans at that task but still lack "intelligence".
We still don't know what that even is exactly, but historically people consistently underestimate how difficult it is.
Not even knowing how to approach it, researchers work on solving single specialized problems instead and make little progress on whatever "intelligence" is.
(If you'd prefer a different definition of intelligence under which Stockfish and GPT are intelligent, then what would you call the ability to solve a wide variety of novel problems? Feel free to substitute that word for "intelligence" above if you'd like to understand what I'm saying.)
> Is a system with A pathfinding intelligent?*
I'm not sure if we should get stuck on definitions of intelligence.
The fact is that these tools are useful, as are the currently existing AI's. The latter can also pass for humans, in many ways, while the algorithms you mentioned can only pass for humans in very narrow domains. Both can exceed human performance in some ways.
Eventually, AI's may be indistinguishable from human or convince humans that they should be treated differently from "mere" programs and algorithms, and at that point we will have entered a new era, call it what you will.
To do this we do not need to consider whether they're intelligent at all.
We have multiple example pairings of X and Y but the common components are that putting X back in Y would be impossible or extremely difficult, and also that X is in some way meant to no longer be contained inside Y as part of either a desired outcome, it’s normal function or the natural expected and thus inevitable result. Cats want to escape, helium leaks, confetti is expelled to have the desired effect, and toothpaste is squeezed out to use it…
For the snow to come out of a snow globe you have to smash it which is not normal usage, not normally meant to happen, and shouldn’t happen by itself. Making the idiom “The snow is out of the globe” not a proper member of this “family” of idioms. (Also I’m not sure if there’s an agreed upon collective noun for idioms)
- I regularly return cereal to its box.
- "helium" and "balloon" have a more awkward rhythm than "confetti" and "cannon". It also loses the connotations of sudden, explosive and exciting change.
- Snow & globe I'm not even sure what that means in practice. It has poor prospects as an idiom. Is the snow even known for leaving globes?
Not only that, but "the confetti has left the cannon" is an alliteration, which makes the phrase even more poetic.
But I do think it's cherry picking the most impressive example. I repeated the dialog (and some variations), each time asking for a completely new idiom, and ChatGPT responded with several phrases that aren't new at all:
"The toothpaste is out of the tube"
"The lid has been lifted"
"The secret has been spilled"
"The arrow has left the bow"
Yes, but don't we cherry pick from what humans have said, too? I'm sure there have been many dumb and obvious proverbs that didn't survive.
> ChatGPT responded with several phrases that aren't new at all
Are you using the GPT-4 version of ChatGPT? That's what GP used.
> And it's exactly the one thing LLMs are trained to do: reproduce patterns of words. They are (perhaps) already better than humans at that one specific skill (another win for AI research) but I don't think it's a sign of general intelligence.
while completely missing why the machine did a better job.
There was an argument made by Raphaël Millière in a recent Mindscape Podcast [1] with Sean Carroll that finally landed for me. He used the example that human beings are driven to eat and reproduce, so by that argument all humans are just eating and reproducing machines. "Ah! But we developed other capabilities along the way to allow us to be good at that!" And that's the point.
GPT-4, for example, is very very good at producing pleasing and useful output for a given input. It uses a simulated neural net to do that. Why would one assume that on the way toward becoming excellent at that task that a neural net wouldn't also acquire other abilities that we associate with reasoning or cognition? When we test GPT-4 for these things (like Theory of Mind) we actually find them.
"Ah hah!" you say, "Humans are set up to learn from the get go, and machines must be trained from scratch." However if you consider the entirety of our genetic legacy together with our childhoods, those are our equivalent "training" from scratch.
I don't think it can be easily dismissed that we're seeing something significant here. It's not human-level intelligence yet. Part of the reason for that is that human brains are vastly more complex than any LLM at the moment (100s of trillions of "parameters" in LLM-speak, along with other advantages). But we're seeing the emergence of something important.
Human beings evolved to eat and reproduce and yet here we are, building computers and inventing complex mathematical models of language and debating whether they're intelligent.
We're so far from the environment we evolved to solve that we've clearly demonstrated the ability to adapt.
ChatGPT doing well at a language task isn't demonstrating that same ability to adapt because that's the task it was designed and trained to do. ChatGPT doing something completely different would be the impressive example.
In short: I don't categorically reject the possibility that LLMs might become capable of more than being "statistics-based text generators", I simply require evidence.
We're seeing those other capabilities emerge; like being able to play chess though it's not been trained to do so. That is, these LLMs are displaying emergent abilities associated with reasoning.
These LLMs aren't R. Daneel Olivaw or R2D2 (which is what I think of when I think of the original term for AI, and what we took to calling AGI). We're closer to seeing the just-the-facts AIs we encounter in Blindsight. Intelligence without awareness.
Funny that we still have to use science fiction to make our comparisons because our philosophy of intelligence, mind, and consciousness are insufficient to speak on the matter clearly.
https://ar5iv.labs.arxiv.org/html/2210.13382
PS: More research has been done since that confirmed and strengthened the conclusion.
""" Here are a few more examples of idioms with meaningful and provable atomic originality:
"The kite has touched the stars" - This phrase could mean that someone has achieved a seemingly impossible goal or reached a level of success that was thought to be unattainable.
"The paint has mingled on the canvas" - This idiom might convey the idea that once certain decisions are made or actions taken, the resulting outcome can't be easily separated or undone, similar to colors of paint that have blended together on a canvas.
"The clock has chimed in reverse" - This expression could be used to describe a situation where something unexpected and unusual has occurred, akin to the unlikely event of a clock chiming in reverse order.
"The flower has danced in the wind" - This phrase could signify that someone or something has gracefully and nimbly adapted to changing circumstances, just as a flower might sway and move in response to the wind. """
It is a good expression though -- evocative but not gross or violent. You could imagine many less successful analogies to something ejecting something else.
While making up "what ifs" can be fun, it doesn't merit either of the words "conspiracy" or "theory".
https://www.instagram.com/p/CQdBiVyh5C2/?hl=en
Now that the cat is out of the bag, or, should I say the confetti is out of the… can?
> I like to think (right now please!) of a cybernetic forest filled with pines and electronics where deer stroll peacefully past computers as if they were flowers with spinning blossoms.
The mirror has shattered.
The poop has hit the propeller.
Pandora's box has opened.
Is the curve of what this class of algorithms can provide sigmoid? If so, then yeah, eventually researchers should be able to democratize it sufficiently that the choice to use versions that can run on private hardware rational. But if the utility increases linearly or better over time/scale, the future will belong to whoever owns the biggest datacenters.
That is, a real OpenAI with a open government body.
Any projects I can follow? Because I haven't seen any one click installers yet that didn't begin with "first install a package manager on the command line"
smh.
well, very close! its interesting that they just got Mac M1/M2 support at all, 2 weeks ago.
for SD I've been using DiffusionBee since maybe October last year.
I expect LLMs to have something in a few weeks, just want to know about it so I can tell other people that need it that way.
Though I have not tried those 1-click installers, instead I have been manually running it.
That project is based on the concept of this Stable Diffusion project: https://github.com/AUTOMATIC1111/stable-diffusion-webui
Which is a few months ahead (because the Stable Diffusion tech happened a few months earlier) and is definitely at a point where anyone can easily run it, locally or on a hosted environment.
I expect this "text-generation-webui" (or something like it) will be just as easy to use in the near, near future.
Wouldn't that be nice? It would also be contrary to all experience of the outcomes and pulls of corporations in modern society. The "local" LLMs will be on the fringe more than at the edge, because the ones that work the best and attract the most money will be the ones controlled by walled-garden "ecosystems."
I really hope it's different. I really hope there are local models. Actual personal assistants actually designed to assist their users and not the people that provide the access.
I want to believe you, but I'm ignorant of the hardware requirements for these things. How soon do you think we'd be able to run something reasonably gpt4-like on, say, a 4090?
The focus instead should be on running expansive models locally on your desktop system at home. This is a better focus in two ways: one the power issue is a non-concern (and AMD has already stated power usage will reach 500-700W average by 2025), two once you have it running locally the avenues of use open up such as being able to access your local model through other devices without the heavy burden on those devices.
We'll see if fine tuning can improve this, but e.g. alpaca-30b is still inferior to ChatGPT-3.5.
Model BoolQ PIQA SIQA HellaSwag WinoGrande ARC-e ARC-c OBQA
LLaMa 13B 78.1 80.1 50.4 79.2 73 74.8 52.7 56.4
Cerebras-GPT 13B - 76.6 - 51.3 64.6 71.4 36.7 28.6You're talking apples to oranges. The "Alpaca method" is a dataset generation method. Nothing about Alpaca's training method is novel, interesting, or efficient. Alpaca used the same standard training method everyone else uses, A100 clusters.
If you mean LoRA/PEFT training which people used to replicate Alpaca then that is also apples to oranges because LoRA/PEFT is a finetuning method not a pre-training method.
Presumably…
The Pythia models are also worth checking out, they might be better than or matched to CerebrasGPTs at each size (although they warn it is not intended for deployment).
Conclusion: the landscape of top open models remains unchanged.
> It would be interesting to know why you chose those FLOPS targets, unfortunately it looks like the models are quite under pre-trained (260B tokens for 13B model)
> We chose to train these models to 20 tokens per param to fit a scaling law to the Pile data set. These models are optimal for a fixed compute budget, not necessarily "best for use". If you had a fixed parameter budget (e.g., because you wanted to fit models on certain hardware) you would train on more tokens. We do that for our customers that seek that performance and want to get LLaMA-like quality with a commercial license
Which is the point made elsewhere in these comments, e.g. https://news.ycombinator.com/item?id=35344192, and also usefully shows how open Cerebras are. They're pretty open, but not as much as they would be if they were optimising for filling in other companies' moats.
https://www.anandtech.com/show/16626/cerebras-unveils-wafer-...
Each individual "chip" has 40GB of SRAM vs ~76MB for the Nvidia H100, and networked pools of external RAM, SSDs and such. Thats why the training architecture is so different.
There's a comparison picture there of one of their chips alongside a regular GPU chip. Effectively they use up the entire wafer.
A 16x CS2 cluster like they describe is like a huge Nvidia cluster in terms of throughput, but more like a single Nvidia node structurally.
Mostly teasing but my guess would be $500k+ since they'll likely price it so that it is the same $ as the equivalent NVIDIA cluster (or very close to it).
A single Nvidia H100 costs somewhere around $30,000 each, so a GPU server with every slot populated costs about $300,000.
https://www.servethehome.com/graphcore-celebrates-a-stunning...
Not sure about the H100, but it seems to be more supply constrained (hence pricier) atm.
Now, the real question is how many HGX nodes "equals" a single CS2 node. The math here is extremely fuzzy, as the benefit to such extreme node consolidation depends on the workload, and the CS-2 takes up less space, but the HGX cluster will have more directly accessible RAM and better turnkey support for stuff since its Nvidia.
This is getting so real so fast.
One example where Cerebras systems perform well is when a user is interested in training models that require long sequence lengths or high-resolution images.
One example is in this publication, https://www.biorxiv.org/content/10.1101/2022.10.10.511571v2, where researchers were able to build genome-scale language models that can learn the evolutionary landscape of SARS-CoV-2 genomes. In the paper mentions, researchers mention "We note that for the larger model sizes (2.5B and 25B), training on the 10,240 length SARS-CoV-2 data was infeasible on GPU clusters due to out-of-memory errors during attention computation."
Cerebras Second-Gen Wafer Scale Chip: 2.6 Trillion 7nm Transistors, 850,000 Cores, 15kW of Power - https://www.tomshardware.com/news/cerebras-wafer-scale-engin...
Trying to cool 15kW to 20kW of power is also rather impressive. https://www.cerebras.net/cs2virtualtour - the engine block and cooling manifold
> The challenge of extracting more than 20 kW of heat from the wafer was solved by having the wafer "float" on a cold plate. The wafer is allowed to expand and contract while remaining in contact with the polished front side of the cold plate, despite the different coefficients of thermal expansion of copper and silicon. The cold plate is much more than a a slab of metal: advanced computational fluid dynamics modelling was used to design a labyrinth of coolant channels capable of maintaining a precise, stable, thermal environment even as 850,000 Al-optimized cores swing into action.
> The power density of the CS-2 is too high for direct air cooling, so liquid cooling is used instead. The internal manifold transfers heat between the CS-2 system's internal coolant and facilties water. Separating these two fluids ensure that the CS-2 system is not affected by changes in the quality of facilities water and that the very highest-quality coolant circulates through the cold plate.
> The two pump modules plug into the upper four dry-break connectors. The lower two are for the air-cooling or water-cooling heat exchanger.
I used to work for a competitor with a more flexible architecture and even our compile times were bad (significant fractions of a day in some cases). And we didn't have to do place and route!
I just googled it and it's apparently bad enough that they had to implement incremental place and route.
(it's all blurry)
Really, really bad mark on whoever is in charge of their web marketing. Images should never look that bad, not even in support, but definitely not in marketing.
edit: so this post is more useful, 4k res using Edge browser
But I'm not sure anymore that it wasn't initially blurry... Perhaps I'm hallucinating, like large language models.
Current image displayed is https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-... , will see if it changes.
https://www.cerebras.net/wp-content/uploads/2023/03/Downstre...
https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-...
https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-...
EDIT: Looks like it scores better with less training - up until it matches GPT-J/Pythia/OPT and doesn't appear to have much benefit. It maybe scores slightly better then GPT-J which is pretty "eh", I'm not sure if GPT-J level performance is really useful for anything? NeoX 20B outperforms it in everything if you don't care about the amount of training needed.
Does the better performance for less training matter if that benefit only applies when it's only performing a lot worse then GPT-J? It appears to lose it's scaling benefits before the performance is interesting enough to matter?
edit: scratch that, it seems the AJAX endpoint returns 504 more often that not.
Which makes me wonder if Nvidia is doing anything with LLMs too?
But it is an issue that their chips are hardly optimized for LLMs.
StableDiffusion plus a whole chain of imagenets can make any visual imagery imaginable in 2GB of RAM. Meanwhile 2GB of RAM barely runs a basic tiny text completion NN that can't do anything intelligent. Text requires a lot more parameters (and more memory/RAM) than images.
Honestly, all the AI ASIC makers drastically underestimated the RAM requirements of future models. Graphcore's 4GB and Tenstorrent's 8GB per IC is kinda laughable, and it takes them longer to adjust than Nvidia. And Cerebras' original pitch was "fit the entire model into SRAM!"
I'm confused as to why 111 million parameter models are trained with the Chinchilla formula. Why not scale up the training data? If you're training smaller models, surely optimizing performance is better than optimizing total compute.
Seems like a silly misunderstanding of the Chinchilla paper, but I'm sure I'm missing something
Money quote for those who don't want to read the whole thing:
'''
When people talk about training a Chinchilla-optimal model, this is what they mean: training a model that matches their estimates for optimality. They estimated the optimal model size for a given compute budget, and the optimal number of training tokens for a given compute budget.
However, when we talk about “optimal” here, what is meant is “what is the cheapest way to obtain a given loss level, in FLOPS.” In practice though, we don’t care about the answer! This is exactly the answer you care about if you’re a researcher at DeepMind/FAIR/AWS who is training a model with the goal of reaching the new SOTA so you can publish a paper and get promoted. If you’re training a model with the goal of actually deploying it, the training cost is going to be dominated by the inference cost. This has two implications:
1) there is a strong incentive to train smaller models which fit on single GPUs
2) we’re fine trading off training time efficiency for inference time efficiency (probably to a ridiculous extent).
Chinchilla implicitly assumes that the majority of the total cost of ownership (TCO) for a LLM is the training cost. In practice, this is only the case if you’re a researcher at a research lab who doesn’t support products (e.g. FAIR/Google Brain/DeepMind/MSR). For almost everyone else, the amount of resources spent on inference will dwarf the amount of resources spent during training.
'''
I'm not so convinced, especially if people are doing multiple training runs for hyperparameter tuning, cleaning data, fixing bugs, etc.
I would be very interested in knowing what portion of OpenAI's compute budget is training. I would not be surprised if it was a significant minority.
That's only true for general-mass-consumer models.
Companies may want to fine-tune/train their own models, which don't have that many users for their narrow use cases (possibly only internal staff), will find that training cost is a substantial chunk of the TCO
As an example the BERT/RoBERTa family were trained for much longer than Chinchilla, you do get diminishing returns though.
There is a point of overtraining where downstream performance is impacted but that’s pretty high.
I think part of the answer to this is also that xxx million parameter decoder-only models don’t seem to be that useful so it may not be worthwhile to optimize them for performance?
They want you to think it's reasonable that because the line is so straight (on a flops log scale) for so long, it could be tempting to extrapolate the pile-loss consequences of continuing compute-optimal training for larger models beyond their largest 13B one, with the obvious caveat that the extrapolation can't continue linearly much further if for no other reason than the test loss isn't going to go below zero (it will flatten out sooner than that).
If you trained beyond compute-optimality on smaller models, it would mess up their straight line and make it look like we are sooner hitting diminishing returns on test loss.
Isn’t the test loss logarithmic? If so it sure can go below zero.
So yes the test loss can be seen as a log, but no it's not allowed to go below zero.
The intuition is that the test loss is the number of bits that the model would need on average to encode each next token in the test part of the pile, given that you have seen the preceding parts.
[1] https://www.hpcwire.com/2021/09/16/cerebras-wafer-scale-engi....
EDIT: Ok, looks like I've missed the hugging face repo. The language they use is a bit confusing.
Nonetheless, it's exciting to see all these open models popping up, and I hope that a LLM equivalent to Stable Diffusion comes sooner than later.
For a small organization or individual who is technically competent and wants to try and do self-hosted inference.
What open model is showing the most promise and how does it’s results compare to the various openAI GPTs?
A simple example problem would be asking for a summary of code. I’ve found openAI’s GPT 3.5 and 4 to give pretty impressive english descriptions of code. Running that locally in batch would retain privacy and even if slow could just be kept running.
Sadly, there's no open model yet that acts like a Swiss knife and gets good-enough results for multiple use cases.
Tldr; Chinchilla isn’t wrong, it’s just useful for a different goal than the llama paper.
There’s 3 hyper parameters to tweak here. Model size (parameter count), number of tokens pre trained on, and amount of compute available. End performance is in theory a function of these three hyperparameters.
You can think of this as an optimization function.
Chinchilla says, if you have a fixed amount of compute, here’s what size and number of tokens to train for maximum performance.
A lot of times, we have a fixed model size though though, because size impact inference costs and latency. Llama operates in this territory. They choose to fix the model size instead of the amount of compute.
This could explain gaps in performance between Cerebras models of size X and llama models of size X. Llama models of size X have way more compute behind them
First, it only holds for a given architecture and implementation. Obviously, a different architecture will have a different training slope. This is clear when comparing LSTM with Transformers, but is also true between transformers that use prenorm/SwiGLU/rotary-positional, and those that follow Vaswani 2017.
In terms of implementation, some algorithms yield the same result with fewer operations (IO, like FlashAttention and other custom CUDA kernels, and parallelism, like PaLM, which both came after Chinchilla), which unambiguously affect the Tflops side of the Chinchilla equation. Also, faster algorithms and better parallelization will yield a given loss sooner, while less power-hunger setups will do that cheaper.
Second, even in the original Chinchilla paper in figure 2, some lines are stopped early before reaching Pareto (likely because it ran out of tokens, but LLaMA makes it seem that >1 epoch training is fine).
> For instance, although Hoffmann et al. (2022) recommends training a 10B model on 200B tokens, we find that the performance of a 7B model continues to improve even after 1T tokens.
Cerebras says:
> For instance, training a small model with too much data results in diminishing returns and less accuracy gains per FLOP
But this is only of concern when you care about the training cost, such as when you are budget limited researcher or a company who doesn't deploy models at scale. But when you care about the total cost of deployment, then making a small model even better with lots of data is a smart move. In the end it matters more to have the most efficient model in prediction, not the most efficient model in training.
I wish rather than stopping training early they would have run more data through a small model so we could have something more competitive with LLaMA 7B.
"We chose to train these models to 20 tokens per param to fit a scaling law to the Pile data set. These models are optimal for a fixed compute budget, not necessarily "best for use". If you had a fixed parameter budget (e.g., because you wanted to fit models on certain hardware) you would train on more tokens. We do that for our customers that seek that performance and want to get LLaMA-like quality with a commercial license"
I'd chip in!
There are plenty of such efforts, but the organizer needs some kind of significance to attract a critical mass, and a AI ASIC chip designer seems like a good candidate.
Then again, maybe they prefer a bunch of privately trained models over an open one since that sells more ASIC time?
This is really weird to hear out loud.
I still think of Discord as a niche gaming chatroom, even though I know that (for instance) a wafer scale IC design company is hosting a Discord now.
Edit: The huggingface page has 0-shot benchmarks which you can compare against the llama paper
Here are some values but I don't know what they mean. LLama 60B on the left, Cerebras 13B on the right.
PiQA: 82.8 / 76.6 WinoGrade: 77.0 / 64.6 ARC-e: 78.9 / 71.4
Table format: Benchmark, Cerebras 13B, LLama 7B, LLama 13B, LLama 60B
HellaSwag, 51.3, 76.1, 79.2, 84.2
Piqa, 76.6, 79.8, 80.1, 82.8
Wino-Grande, 64.6, 70.1, 73.0, 77.0
Arc-e, 71.4, 72.8, 74.8, 78.9
Arc-c, 36.7, 47.6, 52.7, 56.0
OpenBookQA, 28.6, 57.2, 56.4, 60.2
From the Cerebras blog post: "Trained using the Chinchilla formula, these models provide the highest accuracy for a given compute budget."
From the LLaMA paper: "The focus of this work is to train a series of language models that achieve the best possible performance at various inference budgets, by training on more tokens than what is typically used."
The first one is whether they would actually sue. The optics would be terrible. A similar situation occurred in the 90s when the RC4 cipher’s code was leaked. Everyone used the leaked code pretending that it was a new cipher called arc4random, even though they had confirmation from people that licensed the cipher that its output was identical. Nobody was sued, and the RSA company never acknowledged it.
The second one is related to the terms. The LLaMA weights themselves are licensed under terms that exclude commercial use:[0]
> You will not […] use […] the Software Products (or any derivative works thereof, works incorporating the Software Products, or any data produced by the Software), […] for […] any commercial or production purposes.
But the definition of derivative works is gray. AFAIK, if LLaMA is distilled, there is an unsettled argument to be had that the end result is not a LLaMA derivative, and cannot be considered copyright or license infringement, similar to how models trained on blog articles and tweets are not infringing on those authors’ copyright or licensing. The people that make the new model may be in breach of the license if they agreed to it, but maybe not the people that use that new model. Otherwise, ad absurdum, a model trained on the Internet will have content that was generated by LLaMA in its training set, so all models trained on the Internet after Feb 2023 will break the license.
IANAL, but ultimately, Meta wins more by benefiting from what the community contributes on top of their work (similar to what happened with React), than by suing developers that use derivatives of their open models.
[0]: https://docs.google.com/forms/d/e/1FAIpQLSfqNECQnMkycAp2jP4Z...
That was the largest that had inference enabled - I'd really like to try this one: https://huggingface.co/cerebras/Cerebras-GPT-13B
This is called a silver lining for some (in case you were worried about gpt taking your job). Privacy requirements alone will in the near term force major companies to run their own inference (if not training). The expertise required are nearly identical to that of running large scale distributed computational graphs.
This is an interesting diveragence from what happened with web. The backends started out simple before map-reduce and before deconstructing databases and processing distributed logs. With ML, we'll jump right into the complex backends in tandem with easy-picking early stage edge applications (which we see daily on HN).
Alternatives are good!
I remember seeing news about the enormous chip Cerebras was/is selling (pdf https://f.hubspotusercontent30.net/hubfs/8968533/WSE-2%20Dat...).
Has there been any indication that the LLMs released in the last few months use exotic hardware like this, or is it all "standard" hardware?
And, at least officially, OpenAI's outputs can't be used to train other AI models.
Otherwise, if GPT-4 outputs were used to finetune these models, they may become much more interesting.
I don't understand why they describe them as GPT-3 models here as opposed to calling them GPT models. Or even LLMs - but I guess that acronym isn't as widely recognized.
IIRC most open source models to date - including the semi-open LLaMAs - have GPT-3-like performance. Nothing gets close to GPT-3.5 and beyond.
The problem with .ckpt is that it executes arbitrary code in your machine(very unsafe). While .safetensors was made by huggingface in order to have a safe format to store the weights. I've also seen people load up the llama 7B via a .bin file.
Every new variation of model gets some new name, just like every library gets a new name. There were all kinds of BERTs before - DistilBert, Roberta, SciBERT, Schmobert, Schmuber, etc. Many hundreds of them, I think.
With 96GB of memory you should also be able to fine-tune it (possibly some tricks like gradient accumulation and/or checkpointing might be needed), but you have to be ready for many days of computation...
I was thinking since we have API prices in tokens and now it looks like self hosted inference on high end GPUs for similar models. Then based on electricity prices there will be a self-hosted prices in tokens. Then how close are these already? What is the markup today from roughly the raw electricity cost that OpenAI has.
I kind of want 3d marble statues and baroque art of a future reinasance everywhere. But wonder if we will turn minimalistic as a response.
Yay there will be a paper let's gooooooo!