JetMoE: Reaching LLaMA2 performance with 0.1M dollars
research.myshell.ai
research.myshell.ai
They want you to read this as "we spent $100k compared to Meta's spending billions", but that's not actually what this says. It says that they spent $100k and Meta has the resources to spend billions if they wanted to.
We don't know what Facebook spent on training LLaMA 2, but they say that it took them 184320 A100-80GB GPU-hours to train the 7B model [0]. AWS charges $14.46/hour for an instance that has 8 of those [1], which amounts to $1.81/GPU/hr.
At that rate and assuming they paid something resembling AWS's list price, LLaMA 2 7B cost ~$333k. That's more than $100k, but not by orders of magnitude, and it's likely that Facebook wasn't paying the full price AWS is charging today.
[0] https://github.com/meta-llama/llama/blob/main/MODEL_CARD.md#...
I am pretty damn sure I could build a 8 GPU Intel Xeon E5-2686 v4 (Broadwell) (that's what Amazon uses - it's $30 to $75 on eBay) server for less than that and come out ahead on electricity even at full throttle. RTX 4090 are just under $2000 each on eBay.
8 GPU × $2000 (RTX 4090) + $1000 (for the rest of the computer) = $17,000
If pulling 2kW continuously at 15 cents per kW*hr for 1 year that's 2000 watts × 365 days × (0.15/(kW×hr)) or $2,628
In total the computer will cost $19,628 if you throw it in the dumpster at the end of each calendar year of using it.
If you stack internet cost of $200 a month on top, that's $2400 a year, which raises your annual cost to: $22,028
This is still $124,334 cheaper per year than one AWS 8-GPU server if you fully depreciate your own hardware at the end of year 1 to $0.
I could hire an engineer in America to babysit it with the money left over.
This is inconsequential when you're playing Overwatch for a few hours a night and a frame drops now and again. If you're training an iteratively developed LLM though, physical defects could propagate into huge deficiencies in the final model.
Having said that, switching to something like the Tesla V100-SXM2-16GB wouldn't cost that much more.
TBH, I'm shocked at how many people treat Amazon as the first choice for this stuff. Much of it isn't even what most would consider a "production" workload. You are paying for a lot of enterprise-readiness that you don't need for training.
You can thank Amazon's legions of salespeople for that, particularly the end of year junket in Las Vegas where attendees are so pampered that about the only thing they won't do is suck your dick
Oh, yeah, they'll also yell at you on stage if you complain about their UI
I still think it would be impractical at scale because they are so much more hot and power hungry than the datacenter cards, and you would be lucky to score one or two if you’re on a wait list.
I'm actually really surprised that you can still buy 4090s for under $2,000 (cheapest available I saw was $1,800 new and I only took 30 seconds to look), but you can usually sell certain models for quite a bit more. For example, my used 4090 FE is currently worth more than I paid for it.
I've played with AI, and while admittedly I've not done anything super serious, I can tell you that both the 3090 and 4090 are more than capable of performing. Tie them with a power efficient AMD CPU and you have something that can be competitive with enterprise (somewhat).
I've seen the pricing of "cloud" offerings and I've toyed with the idea of creating an "AI Cloud" because I have access to really fast internet and super cheap electricity, but I haven't executed because I'm most certainly not a salesperson. I do, however, know enough about marketing that one should not target price, so there is that...
New Tyan more costly but great case layout.
Don't get me wrong - I've frequently argued that AWS is price gouging and relying on peoples lack of understanding of how the devops costs of running your own works out, but it doesn't take a huge budget before this calculation will look very different (still cheaper to own your own, though).
Add in lots of credits, and if you pay list price, you're being taken to the cleaners..
I've done contract work for clients to be ready to migrate both as part of maximising credits and as part of negotiating posture, and the savings can be enormous (though it'd still usually be cheaper to use managed servers).
For instances specifically, any planned usage will be using either reserved instances or at minimum a compute savings plan (CSP) that drops the hourly rate dramatically in exchange for a committed number of instance hours, with or without an upfront payment.
Finally, there may be a negotiated rate for specific instance types built into the contract. Again, common for very large customers.
source: I was on one of the cloud-related infrastructure teams (left in early 2022). I have no idea about their spend (or discounts) today, but two years ago it was enough that Andy Jassy would meet 1:1 with Mark to "discuss the relationship".
The Mistral guys specifically mention that training speed (due to not needing as much compute) was one of the reasons Mixtral was released so soon after Mistal 7b.
- People hear "mixture of experts" and they think "N specialists" - but ex. think how much know you need to know to autocomplete "Two plus two is "
- Fundamental thing of ML is you define functions and give it data, and the more data you give it to the better. Once youre at "I will simply give it the training data needed to be good enough at the task and wall off that part of the implementation" you're outside ML and have a chicken and egg problem
- We don't know GPT-4 is MoE
- MoE in practice is fundamentally about trading off runtime vs. static size properties to gain inference speed. I.e. 7x8 stored and picking 7x2 at runtime means youre somewhere between 7x2 and 7x3 in quality, inference at 7x2 speed, and have to train and store and load 7x8. You don't reach for it to increase quality, you reach for it to increase inference speed at the expense of inference ram and total model size.
Didn't Yampleg's tweet / leak confirm this one? I mean, he could be wrong about this, but I thought the consensus was on it being true by now.
(Copy of the removed tweets at https://www.reddit.com/r/mlscaling/comments/14wcy7m/gpt4s_de... )
* I used to crusade against this rumor because the only source is that article, and people repeating that source. But I imagine it's a no-brainer given they have enough users that they essentially get a throughput bump 'for free' even if the model weights are huge, i.e. better to utilize as much GPU ram as you can muster, the cost of needing more GPU ram is offset by the cost of being able to run multiple inference against the model all the time anyway
These kind of things are usually only selectively shared.
More important, we let an awful lot of self-promotion from the big guys slide around here.
I can live with the guys and gals doing this on a shoestring getting a little of that sweet hype love. This seems pretty legit.
What is the training data anwyay? Books? (Is it legal to use non public domain, wait whom am I kidding) Reddit posts scrapped from the site? (Without clean up there are lots of bad subreddits) Wikipedia?
Of course I understand why labour costs + employee cost related costs are ignored.
ChatGPT used a crawl of the internet and patches things up with alignment and DPO. Big boys like Microsoft might have deals with publishers to get textbooks in bulk.
Contents from sites with moderation can be filtered using the platform's mechanism, e.g., only include text with a certain length and count of upvotes.
LLMs can be used to generate and filter data as well. Humans have been used to do this, they might have to do this less in the future. Mostly to review what the LLMs are suggesting.
If they had tried to replicate Llama 2 identically with their hardware setup, it'd cost a little bit less than twice their MoE model.
[1] https://github.com/meta-llama/llama/blob/main/MODEL_CARD.md#...
[2] https://blog.ori.co/choosing-between-nvidia-h100-vs-a100-per...
> If two apples cost 2 dollars and 1 apple costs 1.20, what is the discount you're getting
> To calculate the discount, you need to compare the original price of the apples to the price after the discount.
> The original price for two apples is: 2 apples * $1.20/apple = $2.40
> The price for one apple after the discount is: 1 apple * $2.00/apple = $2.00
> Now, let's calculate the discount for one apple: > Original price for one apple - Discounted price for one apple = Discount amount > $2.40 - $2.00 = $0.40
> The discount for one apple is $0.40. Since you're buying two apples, you'll get the discount on both: > Discount amount for two apples = 2 * $0.40 = $0.80
> So, the discount you're getting for two apples is $0.80.
It's good to see MoE being attempted at the smaller sizes, and it may scale well downwards as well given their results. But regardless, 1.25T is very little training data compared to the 6T that Mistral 7B received and even that makes it barely usable and likely not yet saturated. Before it, the sub-13B size class was considered basically an academic exercise.
Interestingly Llama2:13b objects that there is no discount until I clarify: "the discount you're getting [with 2 apples]"
It's not just math though it's any kind of complex reasoning and ambiguity. Comparing to humans is always complex, but humans for the most part wouldn't balk at me asking what discount you're getting without specifying that it's the 2 apples that have the discount in this example. A more advanced model often states the assumptions.
There are lots of nuances in this question as well. I'm still paying 80c more than buying one apple, so I should only buy two apples if I would use two apples.
I mean, you're an idiot of a PI if you have $500K/year of grants and spend it on 5 students and no compute.
> Sally (a girl) has three brothers. Each of her brothers has two sisters. How many sisters does Sally have?
Sally (a girl) has three brothers. Each of her brothers has two sisters. How many sisters does Sally have?
Sally has two sisters. Here's the reasoning: Sally has three brothers. Each of her brothers has two sisters. Since her brothers are the ones who have two sisters each, and we know there are three brothers, we can conclude that each brother has two sisters. Therefore, there are two sisters in total (one for each brother). So, Sally has two sisters.
ollama run llama2:13b
>>> Sally (a girl) has three brothers. Each of her brothers has two sisters. How many sisters does Sally have?
This is a classic lateral thinking puzzle that requires you to think outside the box and challenge your assumptions. The answer is not what you might expect at first glance.
To solve this puzzle, we need to break it down step by step:
1. Sally has three brothers. 2. Each of her brothers has two sisters.
So, let's count the number of sisters that each of Sally's brothers has:
Brother 1: 2 sisters Brother 2: 2 sisters Brother 3: 2 sisters
Now, let's add up all the sisters that Sally has:
Sister 1 + Sister 2 + Sister 3 + Sister 4 + Sister 5 = 2 + 2 + 2 + 2 + 2 = 8
Therefore, Sally has 8 sisters.
Looking at the repo, there is no training or data processing code.
I guess it is good that they mentioned some of it, but yeah, that isn't exceptionally helpful when making claims of it being 100% open source.
I'm not sure why they feel the need to be so secretive if all of the sources are open.
To preempt other queries, maybe just paste the `set` output.
FAANG AI Lab
Lambda Reserved Cloud [1] starts at $1.89 per H100 per hour.
It could be possible to get the cost down to a lower amount:
$1.89 * 96GPUs * 24hours * 14days = ~$61k
Just had a quick chat about kubernetes, very simple discussion. It held well
The Facebook budget includes money to pay off people they've ripped off (in private settlements) and money for lawyers to shield the developers so they can feel free to rip off copyrighted content without having to pay personal penalty or be imprisoned for infringement. It also includes the price of buying lobbyists to alter laws to let this practice continue.
Also, unless the authors work inside Facebook, they have no idea how much Facebook spent on training that model specifically.