How Meta trains large language models at scale
engineering.fb.com
engineering.fb.com
> All of these hardware-related changes were challenging because we had to find a solution that fit within the existing resource constraints, with a very small degree of freedom to change and meet a tight schedule.
Seems like the time constraints put into the team impacted the overall quality of the model.
The last tech team to have no budget and time constraints to pursue their vision? I don’t know, the Xanadu team? Romero’s original Daikatana team?
I love how they built two completely insane clusters just to learn. That's badass.
When you find good content depends on when the algo judges you're already primed to a colorful dopamine intake.
Which data sources? How much of Meta users data (fb, instagram… etc). How do they sanitize PII?
I can't comment on how things like faces get used, but in my experience, PII at Meta is inaccessible by default.
Unless you're impersonating a user on the platform (to access what PII they can see), you have to request special access for logs or database columns that contain so much as user IDs, otherwise the data simply won't show up when you query for it. This is baked into the infrastructure layer, so I doubt the GenAI teams are using something else.
Things that aren't PII aren't "convenient" definitions. Doesn't mean everything that isn't PII is fine to share. It's like saying a kidnapping isn't a murder. That's not a convenient definition of murder; it's just a different thing. We shouldn't start talking like witch hunters as soon as we encounter a situation that we haven't memorised a reasonable response to. We should be able to respond reasonably to new situations.
Obvious examples: data that easily identifies a person (Photo, name, number, UUID, etc)
Thats trivial to block. Where it gets harder is stuff that on it's own isn't PII, but combined with another source, would be
For example, aggregating public comments on a celeb's post. (ie stripping out usernames and likes and assigning a new UUID to each person.) For a single post, thats good enough. You're very unlikley to be able to identify a single person.
But over multiple posts, thats where it gets tricky.
As with large companies, the process for getting permission to use that kind of data is righty difficult, so it often doesn't get used like that.
In the paper covering the original Llama they explicitly list their data sources in table 1 - including saying that they pretrained on the somewhat controversial books3 dataset.
The paper for Llama 2 also explicitly says they don't take data from Meta's products and services; and that they filter out data from sites known to contain a lot of PII. Although it is more coy about precisely what data sources they used, like many such papers are.
By literally opting everyone into their training data set and making it very cumbersome to opt out: https://threadreaderapp.com/thread/1794863603964891567.html
Think of everything being connected to a "Home Computer" in those "Future House of 2020" videos that were out there in 70s or what not.
Another example (very rough) would be something like "Weather data gets to a small model via an API, model looks at it, updates the home dashboard, also sees if there's any alerts, if so, adds x or y to home dashboard appropriately as to what it thinks best."
We can probably achieve the latter example today. (without any significant 'coding' on anyone's part except the API owner)
I want to believe, but I'm still yet to see this kind of set up being anywhere near GPT-4 level.
The weather example seems quite contrived. Why not just display the alerts for your area? Why is a complex system of smaller models reporting up to a slightly larger model necessary?
Because "Flood warning on roads" is very different than "hey there is a possible tornado/aftershocks/*" and I'd like the 2nd one to take up the entire home-assistant dashboard.
I can either code it myself, or let the model figure it out after passing it the YAML file for the dashboard.
>Why is a complex system of smaller models reporting up to a slightly larger model necessary?
because as the other poster said, cost and speed. thousands of queries every day (potentially) at 30 cents per million tokens is very different than 15 dollars per million tokens.
Literally the only thing that can stop Google now is the fact they keep bringing Microsoft and Oracle flunkies into leadership positions.
https://youtu.be/b87I1plPeMg?si=T4XSFUzXG8BwpphR
Ecosystem support for GPU is very strong and so smaller teams might find TPUs to have a steeper learning curve. Sophisticated users can definitely get wins today on TPU if they can get capacity.
And it’s not as if nvidia is standing still, so how the strengths of each change over future generations isn’t set either. Ie TPU are “simple” matrix multipliers and also optimized for operation at scale, GPU are more general purpose and have strong ecosystem power.
Disclaimer - work on GKE on enabling AI workloads.
Hell, Microsoft is renting GPUs from Oracle Cloud to get enough capacity to run Bing.
RTX 4090s are terrible for this task. Off the top of my head:
- VRAM (obviously). Isn't that where the racks come in? Not really. Nvidia famously removed something as basic as NVLink between two cards from the 3090 to the 4090. When it comes to bandwidth between cards (crucial) even 16 lanes of PCIe 4 isn't fast enough. When you start talking about "racks" unless you're running on server grade CPUs (contributing to cost vs power vs density vs perf) you're not going to have nearly enough PCIe lanes to get very far. Even P2P over PCIe requires a hack geohot developed[0] and needless to say that's umm, less than confidence inspiring for what you would lay out ($$$) in terms of hardware, space, cooling, and power. The lack of ECC is a real issue as well.
- Form factor. Remember PCIe lanes, etc? The RTX 4090 is a ~three slot beast when using air cooling and needless to say rigging up something like the dual slot water cooled 4090s I have at scale is another challenge altogether... How are people going to wire this up? What do the enclosures/racks/etc look like? This isn't like crypto mining where cheap 1x PCIe risers can be used without dramatically limiting performance to the point of useless.
- Performance. As grandparent comment noted 4090s are not designed for this workload. In typical usage for training I see them as 10-20% faster than an RTX 3090 at much higher cost. Compared to my H100 with SXM it's ridiculously slow.
- Market segmentation. Nvidia really knows what they're doing here... There are all kinds of limitations you run into with how the hardware is designed (like Tensor Core performance for inference especially).
- Issues at scale. Look at the Meta post - their biggest issues are things that are dramatically worse with consumer cards like the RTX 4090, especially when you're running with some kind of goofy PCIe cabling issue (like risers).
- Power. No matter what power limiting you employ an RTX 4090 is pretty bad for power/performance ratio. The card isn't fundamentally designed for these tasks - it's designed to run screaming for a few hours a day so gamers can push as many FPS at high res as possible. Training, inference, etc is a different beast and the performance vs power ratio for these tasks is terrible compared to A/H100. Now lets talk about the physical cabling, PSU, etc issues. Yes miners had hacks for this as well but it's yet another issue.
- Fan design. There isn't a single "blower" style RTX 4090 on the market. There was a dual-slot RTX 3090 at one point (I have a bunch of them) but Nvidia made Gigabyte pull them from the market because people were using them for this. Figuring out some kind of air-cooling setup with the fan and cooling design of the available RTX 4090 cards sounds like a complete nightmare...
- Licensing issues. Again, laying out the $$$ for this with a deployment that almost certainly violates the Nvidia EULA is a risky investment.
Three RTX 4090s (at 9 slots) to get "only" 72GB of VRAM, talking over PCIe, using 48 PCIe lanes, multi-node over sloooow ethernet (hitting CPU - slower and yet more power), using what likely ends up at ~900 watts (power limited) for significantly reduced throughput and less VRAM is ridiculous. Scaling the kind of ethernet you need for this (100 gig) comes at a very high per-port cost and due to all of these issues the performance would still be terrible.
I'm all for creativity but deploying "racks" of 4090s for AI tasks is (frankly) flat-out stupid.
A few years ago, if you wanted a lot of GPU power you would buy something like [1] - a 4/5U server with space for ten dual-slot PCIe x16 cards and quadruple power supplies for 2000W of fully redundant power. And not a PCIe riser in sight.
I share your scepticism about whether it's common to run >2 4090s because nvidia have indeed sought to make it difficult.
But if there was some sort of supply chain issue that meant you had to, and you had plenty of cash to make it happen? It could probably be done.
Some of the more value-oriented GPU cloud suppliers like RunPod offer servers with multiple 4090s and I assume those do something along these lines. With 21 slots in the backplane, you could probably fit 6 air-cooled three-slot GPUs, even if you weren't resorting to water cooling.
[1] https://www.supermicro.com/en/products/system/4U/4028/SYS-40...
You seem to be trapped in the delusion that this was anyone's first, second, or third choice.
There is workload demand, you can't get H100s, and if you don't start racking up the cards you can get the company will replace you with someone less opinionated.
If needed these other companies have the $$$ to buy the best chips money can buy from Nvidia. Better chips than Google could ever produce.
If anything, this is why IMO Google will fail.
No one will beat them at their game. However if there are any major breakthroughs that might render those processing capacities unneeded, or the major players hitting a wall regarding AI spending, then they will take a massive hit. It will come eventually because the chip business is always in boom/bust cycles.
Wheres others are playing the llm race to the bottom.
And we'll be writing case studies of how they squandered billions of R&D to help found other companies (kinda like Xerox Parc)
Almost every interesting paper after transformers has had it's authors leave to commercialize their own companies.
https://www.yitay.net/blog/training-great-llms-entirely-from...
GPU vs TPU, and good software managing large clusters of them across all sorts of failure.
the funny bit from the above article is the incident when someone forgot about a training job at google, and month later had the model fully trained without an alert of any kind. "outrageously good infra"
If Google had a better chip, or even a chip that was close, they would sell it to anyone and everyone.
From a quick search I can see Google's custom chips are 15x to 30x slower to train AI compared to Nvidia's current latest gen AI specific GPU's.
Maybe they are less than 1/18ths the cost, so google technically have a marginally better unit cost but i doubt it when you consider the R&D cost. They are less bad at inference, but still much worse than even an A100.
Despite various details I don't think that this is an area where Facebook is very different from Google. Both have terrifying amounts of datacenter to play with. Both have long experience making reliable products out of unreliable subsystems. Both have innovative orchestration and storage stacks. Meta hasn't published much or anything about things like reconfigurable optical switches, but that doesn't mean they don't have such a thing.
I would bet money that TPUs are at least better at doing AI research than anything Nvidia will sell you. That alone might be enough for Google to keep getting some new ones fabbed each year. The TPUs you can rent on Google Cloud might very well just be hardware requisitioned by the AI team, for the AI team, that they aren't always using to capacity, and so is "earning out" its CapEx through public rentals.
TPUs are maybe also better at other things Google does internally, too. Running inference on YouTube's audio+video-input timecoded-captions-output model, say.
"Deployed since 2020, TPU v4 outperforms TPU v3 by 2.1x and improves performance/Watt by 2.7x. ... For similar sized systems, it is ~4.3x--4.5x faster than the Graphcore IPU Bow and is 1.2x--1.7x faster and uses 1.3x--1.9x less power than the Nvidia A100. TPU v4s inside the energy-optimized warehouse scale computers of Google Cloud use ~2--6x less energy and produce ~20x less CO2e than contemporary DSAs in typical on-premise data centers."
Here is a link to the paper: https://dl.acm.org/doi/pdf/10.1145/3579371.3589350
Which sure makes the H100 sound both faster and more efficient (per unit of compute) than the TPU v4, given what was in your quote. I don't think your quote does anything to support the position that TPUs are noticeably better than Nvidia's offerings for this task.
Complicating this is that the TPU v5 generation has already come out, and the Nvidia B100 generation is imminent within a couple of months. (So, no, a comparison of TPUv5 to H100 isn't for a future paper... that future paper should be comparing TPUv5 to B100, not H100.)
[0]: https://developer.nvidia.com/blog/nvidia-hopper-architecture...
I think it's aimed at medium to large corporations.
Massive corporations such as Meta and OpenAI would build their own cloud and not rely on this.
The GPU really is a shovel, and can be used without any subscription.
Don't get me wrong, I want there to be competition with Nvidia, I want more access for open source and small players to run and train AI on competitive hardware at our own sites.
But no one is competing, no one has any idea what they're doing. Nvidia has no competition whatsoever, no one is even close.
This lets Nvidia get away with adding more vram onto an AI specific GPU and increase the price by 10x.
This lets Nvidia remove NVLink from current gen consumer cards like the 4090.
This lets Nvidia use their driver licence to prevent cloud platforms from offering consumers cards as a choice in datacenters.
If Nvidia had a shred of competition things would be much better.
jonathan [at] tensordock.com
It should be possible to hook up your idle devices to Nautilus (https://nationalresearchplatform.org/), which is a cluster set up to support researchers at a bunch of universities. I can't guarantee anything since I'm not involved in the cluster management itself, but I can put you in contact with those who are if you're interested.
[1] https://mlcommons.org/benchmarks/training/
[2] https://cloud.google.com/blog/products/compute/introducing-t...
They do sell them - but through their struggling cloud business. Either way, Nvidia's margin is google's opportunity to lower costs.
> I can see Google's custom chips are 15x to 30x slower to train AI
TPUs are designed for inference not training - they're betting that they can serve models to the world at a lower cost structure than their competition. The compute required for inference to serve their billions of customers is far greater than training costs for models - even LLMs. They've been running model inference as a part of production traffic for years.
(I work for Google, but the above is public information.)
Nvidia is in a vaguely unique position in that their products have great tooling support and few companies sell silicon at their scale.
While I do not actually think Google's chips are better or close to being better, I don't think this actually holds?
If the upside of <better chip> is effectively unbounded, it would outweigh the short term benefit of selling them to others, I would think. At least for a company like Google.
In AI, Google (TPU) and Intel (Gaudi) each have chips they push in cloud offerings. The cloud offerings have cross selling opportunities. That by itself would be a reason to keep it internal at their scale. It might also be easier to support one, or a small set, of deployments that are internal vs the variety that external customers would use.
https://www.theverge.com/2023/11/15/23960345/microsoft-cpu-g...
I'm not saying it's a bad buy, either, but if and when they turn the screws, there will be a mass exodus. Solid long term play perhaps, but not going to see Nvidia like price action.
The idea is obvious to everybody in the industry, it's a question of money and motivation.
I understand, especially amongst googlers, there’s a belief there are no others smarter than a googler. But it’s simply not the case. Nvidia is excellent at their core competencies and business, which is making absurdly parallel compute platforms with absurdly powerful interconnects. I’m saying google or meta won’t beat Nvidia at hardware. I’d also point to the fact Nvidias ability to raise capital is the best on earth now, so even money isn’t a barrier.
The advantage CUDA gives is in the tool chains, libraries, research, and all that that tens of thousands of people are contributing to as part of their jobs, research, and hobbies. This is almost -more valuable- than getting top performance. Getting top techniques, top software, top everything by having everyone everywhere working to make the ecosystem of your stuff is invaluable. Google won’t have that. They will just have the hubris of googlers who believe they’re smarter.
I would also note that at this phase of a cycle in tech trying to save billions takes your eye off the prize. Cost optimization comes much later after the market has been fully explored and directions are clean and diminishing returns on R&D kick in. Any company that doesn’t recognize that is run by CPA and deserves the ignominy they’ll face.
> I would also note that at this phase of a cycle in tech trying to save billions takes your eye off the prize
Big Tech companies are conglomerate-ish and can multitask. The search engine folk aren't pushing stuff back onto the backlog to put out fires delaying chip tape-out, and I bet the respective CEOs aren't burning braincycles micromanaging silicon development either; directors 2-3 rungs below the C-suite can motivate and execute on such an undertaking. The answer to "I need a budget of $300M in order to save the company $5-15B over 3 years" is "How soon can you start?"
For training Llama3 Facebook set up two clusters, one using fancy InfiniBand and one just using RoCE over Arista cards: https://engineering.fb.com/2024/03/12/data-center-engineerin... . The latter ended up doing fine, suggesting that all that Mellanox stuff isn't necessary for large-scale training (apparently at a large enough scale ethernet scales better than InfiniBand).
Apple, AWS, Google, Meta, Microsoft all have custom AI-centric silicon.
Rubbish they will fail because the product didn't fit the market, if they're successful they'll have money to buy servers and colo then drive down cost. If they succeed it will be in large part due to the fact they spent thier capital and more importantly time on code/engineers rather than servers.
Right now companies are searching for a use of AI that will add hundreds of billions to thier market cap. Once they find that they can make TPUs, right now only one thing matters; getting there first.
For any given mobile app startup, AWS is effectively infinite. The more money you throw at it the more doodads you get back. Nvidia's supply chain is not infinite and is the bottle neck for all the non-Google players to fight over.
Bad hiring practices aren’t exclusive to them, but from all accounts it seems like their internal focus is on optimizing ad revenue over everything else. I could be wrong or misinformed, but it seems to me like they are playing the finite game in the AI space (DeepMind group aside) while FAIR are playing the infinite game.
*meanwhile MSFT are simply trying to buy their way to relevance (e.g. OpenAI investments, etc) and carve out future revenues (Recall) and Jobs-less Apple is building their trademark walled-garden (AppleIntelligence?). Although the use of unified memory in Apple silicon poses some interesting possibilities for enabling the use of sizable models on consumer hardware.
Overall it seems like “big-tech” is by-and-large uninspired and asleep at the wheel save specific teams like those led by Lecun, Hassabis, etc. not sure where that leaves OpenAI now that Karpathy is gone.
What company do you think has better hiring practices, and subsequently a higher talent pool? Meta is pretty similar to Google's (though with an emphasis on speed over creativity). Microsoft is certainly worse at hiring than the two aforementioned...
They're using advanced cards meant for data centers and machine learning -- almost effectively "custom silicon"
You mean the thing that's already stopped them? If they had seriously invested into the TPU ecosystem in 2015, they would already have "won" AI.
1. Google's stock didn't siginificantly outperformed Meta, Microsoft, etc, in thet past two years.
2. Meta and Microsoft are trying to make their own chips as well.
3. They're not using "consumer video cards" to train AI. I don't even know if you can call these beasts video cards any more. H100 doesn't have HDMI port.
they can't get away with having scraped people's owned work forever. You can't steal things from workers and then undercut them by selling that hard work for pennies, and not expect everything to collapse. I mean, I know that the folks in charge of this aren't really known for their foresight, especially when stock numbers and venture capital are the entire point, but... surely I hope people can recognize that this can't go on unimpeded.
[0] https://www.reuters.com/technology/artificial-intelligence/h...
Yet here we are. Will the consumer video cards get cheaper and better faster or will Google's directors' infighting stop first?
Yes, but you are buying access to tested, supported units that are proven to work, don't require custom software, and are almost plug an go. When its time to upgrade, its not that costly.
Designing, fabricating and deploying your own silicon is Expensive, creating software support for it, also more expense. THen there is the opportunity cost of having to optimise the software stack your self.
You're exchanging a large capex, for a similar sized capex plus a fuckton of opex as well.
google is the biggest loser in all of this.
at consumer level, npu are becoming a useful accelerator, and here google can choose to become the platform of choice.
but nobody comes close for training workflow but nvidia as it stands. imho it is currently possible thanks to community efforts afforded by cuda being the only realistic option.
politics and leadership aside, it would be nice to sustian a market for highly efficient matrix multipliers. also a s/w ecosystem that finally makes multiprocessing workloads easy to integrate for dummies like me.
> Efficient scheduling helps ensure that our resources are used optimally. This involves sophisticated algorithms that can allocate resources based on the needs of different jobs and dynamic scheduling to adapt to changing workloads.
Wow thanks for that, captain obvious. So how do you do it?
I have anecdotally also heard that some are using k8s, but I've not seen that myself. Slurm [1] is basically built for this stuff; that's definitely what I would use!
at my company we use slurm "directly" for static compute we rent or own (i.e. not in a public cloud), but are considering using Kubernetes because that's how we run the rest of the company, and we'd rather invest more effort into being better at k8s than becoming good slurm admins.
¹: https://openai.com/index/scaling-kubernetes-to-2500-nodes/
²: https://www.coreweave.com/blog/sunk-slurm-on-kubernetes-impl...
RDMA and GPUDirect capable. Coordinates over MPI or (hi)redia.
Edit: I had a brain freeze or something... gloo is not CPU only but for whatever reason I don't see it outside of CPU-comms
And at that point having a lot of LLM based avatars that can help "fill in the space" will be valuable.
2. virtually free moderation for their existing (facebook, instagram, threads) and future social media services. likewise, their volume is too insane to even consider paying someone else to process it.
the models they do release are probably toys in comparison to their internal models.
I'm not sure if that's true. They did say they're going to make Llama 3 400B public and if they follow through, that's no toy and I doubt they have something significantly better internally.
Top HW failure modes:
* GPU falling off the bus
I honestly thought "do they mean GPUs falling off a bus entering the data center" and then realized its actually the connectivity, as they mention in the next line GPUs falling off: In this case, GPUs are not detected by the host on PCIe.There is a lot of interesting yet unpublished work on 'data center' scale compute complexes. It was a rabbit hole I fell into several times while at Google.
[0] https://research.google/pubs/the-datacenter-as-a-computer-an...
https://www.theverge.com/2021/11/6/22767046/someone-stole-sh...
I'm wondering if we could prompt llama3 with the above statement. What kind of response would it give?
The infamous "GPU falling off the bus" issue!
This problem typically occurs when a graphics processing unit (GPU) is not properly seated or connected to its expansion slot, such as PCIe, on a motherboard.
Here are some troubleshooting steps to help resolve the issue:
(numbered list of steps or options follows)
Tested on Llama 3 Instruct 7B Q8_0, because that one fits entirely on my GPU.
You are a helpful, smart, kind, and efficient AI assistant. You always fulfill the user's requests to the best of your ability.
And then the query was:
GPU falling off the bus
And yes, I imagine it read that query as ending with an implied "pls help!".
--
Anyone of the super computers listed here https://en.wikipedia.org/wiki/TOP500 suffers from the same issues.
Think about it. While the national labs use these systems to model serious stuff -such as climate or nuclear weapons- Meta uses them to train LLMs. What a joke, honestly!