Meta AI releases Code Llama 70B
twitter.com
twitter.com
Not to diminish the value of the contribution, but "commitment" is an interesting word choice.
When you send something interesting to thousands of people without vetting their credentials, you'd expect the stuff to get "leaked" out eventually (and sooner rather than later).
I'd say it's more appropriate to say the weights were "pirated" than "leaked".
That said, you're probably correct that the community that quickly formed around the "pirated" weights might have influenced Zuckerberg to decide to make llama2's more freely accessible.
https://ollama.ai/library/codellama:70b https://x.com/ollama/status/1752034686615048367?s=20
Just need to run `ollama run codellama:70b` - pretty fast on macbook.
If you want to try it out, this blog post[3] shows how to do it step by step - pretty straightforward.
[1] https://huggingface.co/docs/optimum/concept_guides/quantizat...
[1] https://asciinema.org/a/fFbOEfeTxRShBGbqslwQMfJS4 Note: This recording is in real-time speed, not sped-up.
1.1 GB/ 38 GB 24 MB/s 25m21s
For running 4 bit quantized model, with 70B parameters you will need around 35G Ram to load it in the memory. So I sould say a Mac with at least 48G memory. That is M3 Max.
And instructions on how to change the provider to use Ollama w/ whatever model you want:
Install and run Ollama - Put ollama in your $PATH. E.g. ln -s ./ollama /usr/local/bin/ollama.
- Download Code Llama 70b: ollama pull codellama:70b
- Update Cody's VS Code settings to use the unstable-ollama autocomplete provider.
- Confirm Cody uses Ollama by looking at the Cody output channel or the autocomplete trace view (in the command palette).
- Update the cody settings to use "codellama:70b" as the ollama model
One issue, though: I took a look at the Cody website and it looks like one can't have unlimited completions even when self-hosting a LLM.
I understand you guys have a business model and need to make money out of it. I'm just asking because I work as a teacher and I have students who can't pay an extra subscription and/or students who want to hack into stuff.
A pull/merge request is being worked on: https://github.com/continuedev/continue/pull/758
ugh, not so easy.
Was testing with Codellama-70b this morning and it’s clearly a step up from other OS models
If those rent-seeking bastards at NVidia hadn't killed NVL on the 4090, you could do it on two linked 4090s for only $4k, but we have to live under the thumb of monopolists until such time as AMD 1. catches up on hardware and 2. fixes their software support.
I don’t _believe_ that either of these lets you bypass that restriction (although I’d love to be proven wrong), so if you don’t want to sign up for a subscription you’ll need to use something like Continue.
With careful prompt engineering, you can get a lot out of free Bard except when its censored.
But even if you don't have a faster rig, you can still leverage it for slower tasks to generate docs or tests.
Twinny should really be more popular, didn't find a more powerful no-bullshit plugin for VSCode.
As for small models, Microsoft has been making noise with the unreleased WaveCoder-Ultra-6.7b (https://arxiv.org/abs/2312.14187).
https://huggingface.co/spaces/bigcode/bigcode-models-leaderb...
I highly recommend watching it.
I think Copilot is already highly subsidized by Microsoft.
Let's say you use Copilot around 30% of your daily work hours. How much kWh does an opensource 7B or 13B model use then in a month on one 4090?
EDIT:
I think for a 13B at 30% use per day it comes around 30$/no on energy bill.
So probably with a even more smaller but capable model can beat the Copilot monthly subscription.
So really you're looking at using the GPU for around 10 minutes a day.
Monthly cost is pennies.
46b Mixtral q4 (26.5 gb required) with around 75% in vram: 15 tokens/s - 300w at the wall, nvtop reporting GPU power usage of 70w/30w, 0.37kWh
46b Mixtral q2 (16.1 gb required) with 100% in vram: 30 tokens/s - 350w, nvtop 150w/50w, 0.21kWh.
Same test with 0% in vram: 7 tokens/s - 250w, 0.65kWh
7b Mistral q8 (7.2gb required) with 100% in vram: 45 tokens/s - 300w, nvtop 170w, 0.12kWh
The kWh figures are an estimate for generating 64k tokens (around 35 minutes at 30 tokens/s), it's not an ideal estimate as it only assumes generation and ignores the overhead of prompt processing or having longer contexts in general.
The power usage essentially mirrors token generation speed, which shouldn't be too surprising. The more of the model you can load into fast vram the faster tokens will generate and the less power you'll use for the same amount of tokens generated. Also note that I'm using mid and low tier AMD cards, with the mid tier card being used for the 7b test. If you have an Nvidia card with fast memory bandwidth (i.e., a 3090/4090), or an Apple ARM Ultra, you're going to see in the region of 60 tokens/s for the 7b model. With a mid range Nvidia card (any of the 4070s), or an Apple ARM Max, you can probably expect similar performance on 7b models (45 t/s or so). Apple ARM probably wins purely on total power usage, but you're also going to be paying an arm and a leg for a 64gb model which is the minimum you'd want to run medium/large sized models with reasonable quants (46b Mixtral at q6/8, or 70b at q6), but with the rate models are advancing you may be able to get away with 32gb (Mixtral at q4/6, 34b at q6, 70b at q3).
I'm not sure how many tokens a Copilot style interface is going to churn though but it's probably in the same ballpark. A reasonable figure for either interface at the high end is probably a kWh a day, and even in expensive regions like Europe it's probably no more than $15/mo. The actual cost comparison then becomes a little complicated, spending $1500 on 2 3090s for 48gb of fast vram isn't going to make sense for most people, similarly making do with whatever cards you can get your hands on so long as they have a reasonable amount of vram probably isn't going to pay off in the long run. It also depends on the size of the model you want to use and what amount of quantisation you're willing to put up with, current 34b models or Mixtral at reasonable quants (q4 at least) should be comparable to ChatGPT 3.5, future local models may end up getting better performance (either in terms of generation speed or how smart they are) but ChatGPT 5 may blow everything we have now out of the water. It seems far too early to make purchasing decisions based on what may happen, but most people should be able to run 7b/13b and maybe up to 34/46b models with what they have and not break the bank when it comes time to pay the power bill.
Working out what copilot models perform best has been a deep exercise for myself and has really made me evaluate my own coding style on what I find important and things I look out for when investigating models and evaluating interview candidates.
I think three benchmarks & leaderboards most go to are:
https://huggingface.co/spaces/bigcode/bigcode-models-leaderb... - which is the most understood, broad language capability leaderboad that relies on well understood evaluations and benchmarks.
https://huggingface.co/spaces/mike-ravkine/can-ai-code-resul... - Also comprehensive, but primarily assesses Python and JavaScript.
https://evalplus.github.io/leaderboard.html - which I think is a better take on comparing models you intend to run locally as you can evaluate performance, operability and size in one visualisation.
Best of luck and I would love to know which models & benchmarks you choose and why.
Wow, just realized, in the future employers will mostly interview LLMs instead of people.
I wonder why they didn't use DeepSeek under the "senior" interview test. I am curious to see how it stacks up there.
I honestly don't know what benchmarks to look at or even what questions to be asking.
I would be uncomfortable recommending less than 4-bit quantization on a non-MoE model, which is ~40GB on a 70B model.
No… that’s not such a great thing. Helpful in a pinch, but if you’re not running at least 70% of your layers on the GPU, then you barely get any benefit from the GPU in my experience. The vast gulf in performance between the CPU and GPU means that the GPU is just spinning its wheels waiting on the CPU. Running half of a model on the GPU is not useful.
> Then again, one could have grabbed 2x 3090s for the price of a 4090 and ended up with 48gb of VRAM in exchange for a very tolerable performance hit.
I agree with this, if someone has a desktop that can fit two GPUs.
Water cooling can get you down to 2x slot height, with all of the trouble involved in water cooling. NVIDIA really segmented the market quite well. Gamers hate blower cards, but they are the right physical dimensions to make multi-GPU work well, and they are exclusively on the workstation cards.
Then you can just run it entirely on CPU. There is no point to buy an expensive GPU to run LLMs to be bottlenecked by your CPU in the first place. Which is why I do not get so excited with these huge models, as they gain less traction as not as many people can run them locally, and finetuning is probably more costly too.
Microsoft accidentally leaked that ChatGPT-3.5-Turbo is apparently only 20B parameters.
24GB of VRAM is enough to run ~33B parameter models, and enough to run Mixtral (which is a MoE, which makes direct comparisons to “traditional” LLMs a little more confusing.)
I don’t think there’s a clear answer of what hardware someone should get. It depends. Should you give up performance on the models most people run locally in hopes of running very large models, or give up the ability to run very large models in favor of prioritizing performance on the models that are popular and proven today?
The downside of Apple's hardware at the moment is that the training ecosystem is very much focused on CUDA; llama.cpp has an open issue about Metal-accelerated training: https://github.com/ggerganov/llama.cpp/issues/3799 - but no work on it so far. This is likely because training at any significant sizes requires enough juice that it's pretty much always better to do it in the cloud currently, where, again, CUDA is the well-established ecosystem, and it's cheaper and easier for datacenter operators to scale. But, in principle, much faster training on Apple hardware should be possible, and eventually someone will get it done.
I was curious if some kind of summary or compression of old exchanges flagged as such might allow the app to remember stuff that had been discussed but fallen outside the token limit.
But possibly request key details lost during summary to bring them back into the new context.
I had thought chatgpt was doing something like this but haven’t read about it.
Anything that saves me time writing “boilerplate” or figuring out the boring problems on projects is welcome - so I can expend the organic compute cycles on solving the more difficult software engineering tasks :)
Cool nonetheless
Anybody have $10 billion sitting around to deploy that gigantic open source set-up for millions of users? There's your moat and only a relatively few companies will be able to do it.
One of Google's moats is, has been, and will always be the scale required to just get into the search game and the tens of billions of dollars you need to compete in search effectively (and that's before you get to competing with their brand). Microsoft has spent over a hundred billion dollars trying to compete with Google, and there's little evidence anybody else has done better anywhere (Western Europe hasn't done anything in search, there's Baidu out of China, and Yandex out of Russia).
VRAM isn't moving nearly as fast as the models are progressing in size. And it's never going to. The cost will get ever greater to operate these at scale.
Unless someone sees a huge paradigm change for cheaper, consumer accessible GPUs in the near future (Intel? AMD? China?). As it is, Nvidia owns the market and they're part of the moat cost problem.
Models of any given quality are declining in size (both number of parameters, and also VRAM required for inference per parameter because quantization methods are improving.)
Raw FLOPs may increase each generation but VRAM becomes a limiting factor. And fast VRAM is expensive.
I do expect to see incremental innovation in reducing the size of foundational models.
for end users, yes. For small companies that want to finetune, evaluate and create derivatives, it reduces the cost by millions.
You don't have to run them locally.
It is in at least 2025. AMD (and Intel, maybe) will have M-Pro-Esque APUs that can run a 70B model at very reasonable speeds.
I am pretty sure Intel is going to rock the VRAM boat on desktops as well. They literally have no market to lose, unlike AMD which infuriatingly still artificially segments their high VRAM cards.
Could you please expand on what else would be capable of running such models locally?
How about a linux laptop/desktop with specific hardware configuration?
Compute is actually not that big of a deal once generation is ongoing, compared to memory bandwidth. But the initial prompt processing can easily be an order of magnitude slower on CPU, so for large prompts (which would be the case for code completion), acceleration is necessary.
For example both the RTX 4090 and the RTX 6000 Ada Generation use the AD102 chip. The RTX 6000 Ada though, would be able to run 70b models due to the larger memory pool despite having the same memory interface width.
It also really depends on what you consider "beginner level". Fine-tuning is really easy these days and many people do it, but you really want CUDA for that.
People are recommending Macbooks because they're a relatively cheap and easy way to get a very large amount of RAM hooked up to your accelerator.
Note that these are quantized versions of the model, so they're not as good as the original 70B model, though people claim their performance is really close to original performance. To run without quantization you'd need about 140GB of VRAM. Which would only be possible with an NVidia H100 (don't know the price) or two A100's (at $18,000 each).
I run 33B parameter models on my RTX 3090 (24GB VRAM) no problem. 70B should easily fit into 64GB of RAM.
Mixtral runs at about 43 tokens/s at q3_K_S with all layers offloaded. I normally avoid going below 4-bit quantization, but Mixtral doesn’t seem phased. I’m not sure if the MoE just makes it more resilient to quantization, or what the deal is. If I run it at q4_0, then it runs at about 24 tokens/s, with 26 out of 33 layers offloaded, which is still perfectly usable, but I don’t usually see the need with Mixtral.
Ollama dynamically adjusts the layers offloaded based on the model and context size, so if I need to run with a larger context window, that reduces the number of layers that will fit on the GPU and that impacts performance, but things generally work well.
There are a few things to keep in mind: no programmer that I know is sitting there typing code for hours at a time without stopping. There’s a lot more to being a developer than just typing, whether it is debugging, thinking, JIRA, Slack, or whatever else. These CoPilot-like tools will only activate after you type something, then stop for a defined timeout period. While you’re typing, they do nothing. After they generate, they do nothing.
I would honestly be surprised if the GPU active time was more than 10% averaged over an hour. When actively working on a large LLM, the RTX 3090 is drawing close to 400W in my desktop. At a 10% duty cycle (active time), that would be 40W on average, which would be 320Wh over the course of a full 8-hour day of crazy productivity. My electric rate is about 15¢/kWh, so that would be about 5¢ per day. It is absolutely not running at a 100% duty cycle, and it’s absurd to even do the math for that, but we can multiply by 10 and say that if you’re somehow a mythical “10x developer” then it would be 50¢/day in electricity here. I think 5¢/day to 10¢/day is closer to reality. Either way, the cost is marginal at the scale of a software developer’s salary.
Note that M1/M2 Ultra is quite a bit faster than M3 Max, mostly due to 800 Gb/s vs 400 Gb/s memory
In traditional software, the same program compiled for 32-bit and 64-bit architectures won’t be able to handle all of the same inputs, because the 32-bit version is limited by the available address space. It’s still the same program.
If we’re not willing to declare that you are a completely separate person when you’re tired, or that 32-bit and 64-bit versions are completely different programs, then I don’t think it’s worth getting overly philosophical about quantization. A quantized model is still the same model.
The quality loss from using 4+ bit quantization is minimal, in my experience.
Yes, it has a small impact on accuracy, but with massive efficiency gains. I don’t really think anyone should be running the full models outside of research in the first place. If anything, the quantized models should be considered the “real” models, and the full fp16/fp32 model should just be considered a research artifact distinct from the model. But this philosophical rabbit hole doesn’t seem to lead anywhere interesting to me.
Various papers have shown that 4-bit quantization is a great balance. One example: https://arxiv.org/pdf/2212.09720.pdf
The question of whether I am still me after a traumatic brain injury is philosophically unclear, and likely depends on specifics about the extent of the deficits.
It’s far more similar to the model being perpetually tired than it is to a TBI.
You may nitpick the analogy, but analogies are never exact. You also ignore the other piece that I pointed out, which is how we treat other software that comes in multiple slightly different forms.
For example, take a look at the GGUF file sizes here: https://huggingface.co/TheBloke/Llama-2-70B-GGUF
Instructions here: https://github.com/facebookresearch/llama/pull/947/
Or how you deal with context length? I.e. do you send anything other than the current file? How is the prompt constructed?
(Please don't say "commoditize your complement" without explaining what exactly they're commoditizing...)
If Meta can help prevent there from being an AI monopoly company, but rather an ecosystem of comparable products, then they avoid having another threatening tech giant competitor, as well as preventing their own AI work and products from being devalued.
Think of it like Google releasing a web browser.
It's akin to a Great Filter, if such an analogy helps. If Meta's open models make a company's closed models uneconomical for others to consume, then the business case for those models is compromised and the odds of them growing to a size where they can compete with Meta in other ways is mitigated a bit.
On the advertiser side, they're commoditizing the ability for companies to write more persuasively-targeted ads. Higher click-through rates = more money.
[edit]: For models that generate code instead of content (TFA), it's obviously a different story. I don't have a good grip on that story, beyond "they're using their otherwise-idle GPU farms to buy goodwill and innovate on training methods".
They've gained an incredible amount of influence and mindshare.
They get free R&D and suppress competition, while looking like they have principles. Yann is clueless about open source principles, or the models would have been Apache or some other comparably open license. It's all ruthless corporate strategy, regardless of the mouth noises coming out of various meta employees.
Just because certain entities can't profitably use a product or obtain a license doesn't make it not-open. AGPL is open, for an extreme example.
This argument is also subjective, and not new - "Which is more open BSD-style licenses or GPL?" has ben a guaranteed flameware starter for decades.
It's shitty when other companies do it. It's shitty when Broadcom does it. It's shitty when Meta does it.
It's never a not shitty thing to do.
But sure, sounds more reasonable
For shareholders, this subpar performance has destroyed value. Disney stock has underperformed the stocks
of Disney’s self-selected proxy peers and the broader market over every relevant period during the last
decade and during the tenure of each non-management director. Furthermore, it has underperformed since
Bob Iger was first appointed CEO in 2005 – a period during which he has served as CEO or Executive
Chairman (directing the Company’s creative endeavors in this role) for all but 11 months. Disney shareholders
were once over $200 billion wealthier than they are now
Which is radically different from previous 90 yearshttps://trianpartners.com/wp-content/uploads/2023/12/Trian-N...
Disney has steamrolled Hollywood for the last decade, bringing in by far the biggest global box office revenue in 7 consecutive years out of 8. They have more billion dollar box office movies than every other studio co mbined. This kind of dominance was unheard of in the history of Hollywood.
Setting box office aside, Disney revenue has tripled since Iger took over and is twice as much as it should be adjusted for inflation.
The idea that the company has underperformed for the last 10 years or that they spend millions "on a whim" is a joke. And using share price as some justification is even more absurd, share price was double what it was today just in 2021.
Alone, it was unlikely they would become a major player in a field that might be massively important. With a large community building upon their base they have a chance to influence the direction of development and possibly prevent a proprietary monopoly in the hands of another company.
Maybe now their leadership wants to push for practicality so they don't end up like Google (also a research powerhouse but failing to convert to popular advances) so they are publicly pushing strong LLMs.
1. They become an attractive place for AI researchers to work, and can bring in better staff. 2. They make it less appealing for startups to enter the space and build large foundation models (Meta would prefer 1,000 startups pop up and play around with other people's models, than 1000 startups popping up and trying to build better foundational models). 3. They put cost pressure on AI as a service providers. When LLAMA exists it's harder for companies to make a profit just selling access to models. Along with 2 this further limits the possibility of startups entering the foundational model space, because the path to monetization/breakeven is more difficult.
Essentially this puts Meta, Google, and OpenAI/Microsoft (Anthropic/Amazon as a number four maybe) as the only real players in the cutting edge foundational model space. Worst case scenario they maintain their place in the current tech hegemony as newcomers are blocked from competing.
Mistral is right up there.
Hopefully they can prove me wrong though!
Then Ai sprung to the front pages and any CEO who stood up and said "Ai" was rewarded with a 10x stock price. The unloved stepchild that was the ML team became the A team and the metaverse team have been sent to the naughty step. Facebook/Meta have no actual customer facing use for Ai unlike Microsoft/Google/GitHub but they like a good stonk price rise and so what we see is their stategy to stay in the ai game and relevant.
It turns out it is pretty good for the rest of us (possibly the first time facebook has given something positive to humanity) as we get shinny toys to play with.
It costs them nothing to open it up, so why not. Kinda like all the rest of their GitHub repos.
Meta releases model. Joe builds a cool app with it, earns some internet points and if lucky a few hundred bucks. Meta copies app, multiply Joes success story with 1 billion users and earn a few million bucks.
Joe is happy, Meta is happy. Everybody is happy.
Meta sees this as the way to improve their AI offerings faster than others and, eventually, better than others.
Instead of a small group of engineers working on this inside Meta, the Open Source community helps improve it.
They have a history of this with React, PyTorch, hhvm, etc. All these have gotten better as OS projects faster than Meta alone would have been able to do.
Essentially, you mitigate IP claims and reduce vendor dependency.
https://eightify.app/summary/technology-and-software/the-imp...
(My theory: if there's an AI pot of gold, what megacorp can risk one of the others getting to it first?)
Most likely, they work for your competitors. They may not be working to improve your system for free.
> No company can replicate innovation from open source internally.
Lot of innovation does come from companies.
Of course, i am not arguing that. But when it comes to software as general as code generation, or text generation, the possible applications are so broad, that a team of A.I. researchers in a company, however talented and productive they are, cannot possibly optimize it for every possible use case.
That's what Yan Le Cunn is referring to, and i agree with him. There are a lot of companies which push deep learning forward, and do not release their code or weights freely.
(I try to train myself to say it right ..)
(Look at “tags” to see the different quantizations)
[0] https://github.com/facebookresearch/codellama [1] https://ai.meta.com/research/publications/code-llama-open-fo...
Who would have thought that Meta, that has been chucking billions on the metaverse is on the forefront of Open Source AI.
Not to mention their stock is up and they are worth $1TN, again.
Not sure how I feel about this given the fact of all the scandals that have plagued them and the massive 1BN fine from the EU, Cambridge Analytica, and last of all caused a genocide in Myanmar.
Goes to show that nobody cares about all of these scandals and just moves on onto the future, allowing Facebook to still collect all this data for their models.
If any other startup or mid sized company had at least two of these large scandals, they would be dead in the water.
But I do wonder in the back of my mind why. And I should be suspicious of their angle and I will keep thinking about it. Is it paranoid to think that maybe their angle is putting almost some kind of metadata by style of code being unique to different machines that they can trace generated code to different people? Is that their angle or am I biased in remembering who they have been for the past decade?
Which is open source.
I do not know anybody important in the AI space apart from Google using TensorFlow.
Even if you want to reproduce the model and they give you the data, you would need to do this at Facebook scale, so you and the GP are just making moot points all around.
https://about.fb.com/news/2023/05/metas-infrastructure-for-a...
https://www.theregister.com/2024/01/20/metas_ai_plans/
The fact that these models are coming from Meta in the open rather than Google which releases only papers with no model tell's me that Meta's models is open enough for everyone to use.
Besides, everyone using the Pytorch framework benefits Meta in the same way they were originally founded as a company:
Network effects
It's relevant.
The thing that bothers me more is that it's not actually an open-source licence; there are restrictions on what you can do with it, and whatever you do with the model is subject to those restrictions. It's still very useful and I'm not opposed to them releasing it under that licence (they need to recoup the costs somehow), but "open-source" (or even "open") it is not.
1. Similar to other initiatives (mainly opencompute but also PyTorch, React etc), community improvements help them improve their own infra and helps attract talent.
2. Helping people create better content ultimately improves quality of content on their platforms (Both FoA & RL)
Sources:
[1]Interview with verge: https://www.theverge.com/23889057/mark-zuckerberg-meta-ai-el... . Search for "regulatory capture right now with AI"
> Zuck: ... And we believe that it’s generally positive to open-source a lot of our infrastructure for a few reasons. One is that we don’t have a cloud business, right? So it’s not like we’re selling access to the infrastructure, so giving it away is fine. And then, when we do give it away, we generally benefit from innovation from the ecosystem, and when other people adopt the stuff, it increases volume and drives down prices.
> Interviewer: Like PyTorch, for example?
> Zuck: When I was talking about driving down prices, I was thinking about stuff like Open Compute, where we open-sourced our server designs, and now the factories that are making those kinds of servers can generate way more of them because other companies like Amazon and others are ordering the same designs, that drives down the price for everyone, which is good.
Multiple of their major competitors/other large tech companies are trying to monetize LLMs. OpenAI maneuvering an early lead into a dominant position would be another potential major competitor. If releasing these models slows or hurts them that is in and itself a benefit.
What benefit is there to grabbing market share from your competitors... in a business you don't even want to be in?
By that logic you could justify any bizarre business decision. Should Google launch a social network, to hurt their competitor Facebook? Should Facebook, Amazon and Microsoft each launch a phone?
* https://www.lifewire.com/whatever-happened-to-the-facebook-p...
I mean, Google did launch a social network, to hurt their competitor Facebook. It was a whole thing. It was even a really nice system, eventually.
There's far more risk if Meta were to try to directly compete with OpenAI and Microsoft on this. They'd have to manage the infra, work to acquire customers, etc, etc on top of building these massive models. If it's not a space they really want to be in, it's a space they can easily disrupt.
Meta's late game realization was that Google owned the web via search and Apple took over a lot of the mobile space with their walled garden. I suspect Meta's view now is that it's much easier to just prevent something like this from happening with AI early on.
Its a smart move IMO
Put another way - if OpenAI were the only game in town how much would they be charging for their product? They’re competing on price because competitors exist. Now imagine the price if a hypothetical high quality open source model existed that can customers can use for “free”.
That’s the future Meta wants. They weren’t getting rich selling shovels like cloud providers are, they want everyone digging. And everyone digs when the shovels are free.
Content generation is complementary to most of meta's apps and projects
If Meta were at the forefront, these models would not be openly available.
They are scrambling.
“No one can compete with us, but it’s cute to try! Make applications though” —almost direct quote from Sam Altman.
I have 64gb and an RTX 3090 and a macbook M3, and I already can’t run a lot of the newest models even in their quantized form.
The business model requires this to be a subscription service. At least as of today…
So. Maybe we could help other people figure out why VRAM is maxing out. I think it has to do with various new platforms leaking memory.
In my case, I suspect ollama and diffusers are not actually evicting VRAM. nvidia-smi shows it in one case, but I haven’t figured it out yet.
Hey, my point remains. The models are going to get too expensive for me, personally, to run locally. I suspect we’ll default into subscriptions to APIs because the upgrade slope is too steep.
I used copilot to refactor that, and it just didn’t put no_grad back, and I did not notice.
I was uselessly recalculating all of my weights to /dev/null and waste heat.
I daemonized my process to control it more tightly, but I still see way above expected vram allocation.
Just keep in mind. It’s not like the models are going to get smaller, past the quantization limit. What, is there a quantization of retrievable information to 0 bits? ;)
This is the exact opposite of bait and switch. The current model couldn't be un-opensourced and over time it will just become easier to run it.
Also unless there is reason to believe that prompt engineering of different model families is very different(which honestly I don't believe), there is no effect of baiting. I believe it will always be the case that best 2-3 models would be closed weights.
Of course, you can buy quite a lot of hosted model API access or cloud GPU time for that money.
You can buy a 24GB gpu for $150-ish (P40).
SUPPORTED
=========
* Ada / Hopper / A4xxx (but not A4000)
* Ampere / A3xxx
* Turing / Quadro RTX / GTX 16xx / RTX 20XX / Volta / Tesla
EOL 2023/2024
=============
* Pascal / Quadro P / Geforce GTX 10XX / Tesla
Unsupported
===========
* Maxwell
* Kepler
* Fermi
* Tesla (yes, this one pops up over and over, chaotically)
* Curie
Older don't really do GPGPU much. The older cards are also quite slow relative to modern ones! A lot of the ancient workstation cards can run big models cheaply, but (1) with incredible software complexity (2) very slowly, even relative to modern CPUs.
Blender rendering very much isn't ML, but it is a nice, standardized benchmark:
As a point of reference: A P40 has a score of 774 for Blender rendering, and a 4090 has 11,321. There are CPUs ($$$) in the 2000 mark, so about dual P40. It's hard for me to justify a P40-style GPU over something like a 4060Ti 16GB (3800), an Arc a770 16GB (1900), or a 7600XT 16GB (1300). They cost more, but the speed difference is nontrivial, as is the compatibility difference and support life. A lot of work is going into making modern Intel / AMD GPUs supported, while ancient ones are being deprecated.
I find that my hosts using 9x P40 do inference on 70b models MUCH MUCH faster than a e.g. a dual 7763 and cost a lot less. ... and can also support 200B parameter models!
For the price of a single 4090, which doesn't have enough ram to run anything I'm interested in, I can have slower cards which have cumulatively 15 times the memory and cumulatively 3.5 times the memory bandwidth.
Technically, P40 is rated at an impressive 347.1GB/sec memory bandwidth, and 4060, at a slightly lower 272GB/sec. For bandwidth-limited workloads, the P40 still wins.
The 4090 is about 3-4x that, but as you point out, is not cost-competitive.
What do you use to fit 9x P40 cards in one machine, supply them with 2-3kW of power, and keep them cooled? Best I've found are older rackmount servers, and the ones I was looking at stoped short of that.
You can get gpu server chassis that have 10 pci-slots too! for around $2k on ebay. But note that there is a hardware limitation on the PCI-E cards such that each card can only directly communicate with 8 others at a time. Beware, they're LOUD even by the standards of sever hardware.
Oh also the nvidia tesla power connectors have cpu-connector like polarity instead of pci-e, so at least in my chassis I needed to adapt them.
Also keep in mind that if you aren't using a special gpu chassis, the tesla cards don't have fans, so you have to provide cooling.