Llama-3.3-70B-Instruct
huggingface.co
huggingface.co
Seems to perform on par with or slightly better than Llama 3.2 405B, which is crazy impressive.
Edit: According to Zuck (https://www.instagram.com/p/DDPm9gqv2cW/) this is the last release in the Llama 3 series, and we'll see Llama 4 in 2025. Hype!!
but also wrestled with this mentally.
Meta both improves the technology or inference, while also trapping themselves alongside every other person training models to always update the training set every few months, so it knows what its talking about with relevant current events
If we can get better quantization, or bigger GPU memory footprints, we might be able to use these big models locally for solid coding assistants. That's what I think we have to look forward to (among other benefits) in the year(s) ahead.
1. lmstudio-community/Llama-3.3-70B-Instruct-GGUF/Llama-3.3-70B-Instruct-Q4_K_M.gguf
Edit: Perhaps answering my own question:
λ ollama run hf.co/lmstudio-community/Llama-3.3-70B-Instruct-GGUF:Q4_K_M
>>> Hi. Who are you?
I'm an artificial intelligence model known as Llama. Llama stands for "Large Language Model Meta AI."
System Prompt: You are an advanced super intelligence, capable of everything from software development to dream interpretation.
>> Who are you?
I am an artificial intelligence designed to simulate human-like conversation and answer questions to the best of my knowledge. I'm a large language model, which means I've been trained on a massive dataset of text from various sources, including books, articles, research papers, and websites.
My capabilities are diverse and extensive. I can:
1. **Answer questions**: I can process natural language queries and provide relevant answers, covering a wide range of topics from science and history to entertainment and culture.
2. **Generate text**: I can create text based on a prompt or topic, including articles, stories, emails, and even entire books.
3. **Translate languages**: I can translate text from one language to another, including popular languages such as Spanish, French, German, Chinese, and many more.
4. **Summarize content**: I can summarize long pieces of text into concise, easily digestible versions, highlighting key points and main ideas.
5. **Offer suggestions**: I can provide suggestions for topics such as gift ideas, travel destinations, books to read, and more.
6. **Engage in conversation**: I can have a natural-sounding conversation with you, using context and understanding to respond to your questions and statements.
7. **Dream interpretation**: As you mentioned, I can even help interpret dreams, providing insights into the symbolism, themes, and emotions that may be present in your subconscious mind.
8. **Software development**: I can assist with software development tasks, such as writing code, debugging, and optimizing algorithms.
My knowledge is constantly evolving, and I'm always learning from the interactions I have with users like you. So, feel free to ask me anything – I'll do my best to help!With Llama specifically I recall someone comparing various quants on perplexity finding that even at Q3, 70B is still smarter than 34B. So quantization is generally worthwhile so long as you have a larger model that you can squeeze into your VRAM budget with it, and don't mind the slowdown from more parameters.
I hope Llama 4 reintroduces that mid sized model size.
If you want to offload fully to VRAM, I'd say 8B is the limit. If you're keeping some on RAM, 15-20B can still give OK performance, depending on your tolerance.
>How much does quantization impact output quality?
Basically with more quantization the output becomes more incoherent and less realistic. At the extreme end it's basically just gibberish. I think the sweet spot generally is at 4 bits. At that point the model is pretty compact and the quality isn't diminished too much.
Speed takes quite a hit once you have a few layers on the CPU, but depending on needs it can be doable. I've just asked LLama 3.3 70B Q5_K_M a question and it offloaded about 5 of the 80 layers, so running almost entirely on my 5900X CPU, but still churning out about one word per second.
In my experience quantization affects prompt adherence primarily and answer accuracy secondarily. For example, if you have multiple clauses, ie one or more "if this then that", then quantization might get it to not consider those. I also find they tend to answers more generally and less precise at higher quantization levels.
As a concrete example, I've been asking the LLama 3.2 Vision 8B model to categorize some images. The default instruct model in general has been heavily trained to output general commentary on the image. If in the prompt I tell it to "output the category only", the Q4_K_M variant sometimes ignores that instruction, while the Q8 variant almost always respect it.
Larger models primarily bring more knowledge in my experience, but usually also better prompt adherence. Larger models also typically can support larger contexts, though this can vary, check the model cards.
edit: I should clarify. More knowledge also often translates to better, more accurate output. For example, a larger model might recognize an idiom and answer accordingly, while the smaller model fails to recognize it and thus provides a poor answer.
Depending on your needs, 12GB might be quite decent or it might be insufficient. If you need an assistant-like model, I liked the Gemma 2 9B Q5_K_M. And I've been quite impressed by LLama 3.2 Vision 8B Q4_K_M for describing images and transcribing text from images.
But for more open-ended stuff, especially if larger contexts is needed, I think you might find it underwhelming.
But I do like to compare. Open WebUI for example makes it very easy where you can load up multiple models and it'll send the same prompt to each one in turn, and show the answers side by side.
bash> time ollama run llama3.3 "What's the purpose of an LLM?" | tee ~/Downloads/what\ is\ an\ LLM.txt
A Large Language Model (LLM) is a type of artificial intelligence (AI) designed to process and understand human language. The primary purposes of an LLM are:
(... contents excerpted for brevity) Overall, the purpose of an LLM is to augment human capabilities by providing a powerful tool for understanding, generating, and interacting with human language.
real 0m59.040s
user 0m0.071s
sys 0m0.081s
pmarreck 59s35ms
20241206220629 ~ bash> wc -w Downloads/what\ is\ an\ LLM.txt
359 Downloads/what is an LLM.txtLooks like LM Studio is available for ARM based Macs, if you want to give that a try, that'd be one way to get these stats. LM Studio also surfaces up some parameters to play around with, and keeps a record of past conversations if that might appeal to you.
Good will isn't worth as much as cheap moderation automation and fancy features, but it's worth something.
You have the sequence reversed as Meta already created ad targeting models. Meta was forced to scale its AI competence for ad targeting when Apple sent out a privacy update that destroyed tracking-based ad-serving and tanked Meta's share price by deleting billions in revenue for Meta over many quarters. Now that Meta has this skill as a core-competence, they are creating new models for public release. Why they are doing so is debatable[2], but I imagine the cost is marginal since they already had the GPU clusters, talent and know-how for survival purposes.
1. https://www.businessinsider.com/metas-bet-on-ai-has-saved-it...
2. I suspect Zuckerberg is not enthused by the idea of a future AI Apple-analog unilaterally shutting him out of the market. Having your net worth cut in half by a press-release has got to hurt.
Without access to the tracking signal, it's been more important to build out a system that can recreate the value from that lost signal by analyzing what users are actually sharing and saying on their platform. Hence the importance of chat (VR, text, video...) and AI that can be used to process and extract value from a chat signal.
I believe Meta's primary revenue source is still advertising (98%), so that is probably 98% of the why.
I suppose I read your first sentence as being in future tense when it might not be. The thrust of my argument is that Meta already successfully built those ad targeting models (Advantage+), and they preceded the Llama releases, so they don't need to use Llama-derived models for ad targeting, as I understood your comment to be suggesting. The sequence was not/will not be "Llama -> ad targeting", but was "ad targeting -> Llama"
Meta didn't have to release the weights of the models. Ad revenue doesn't explain why they did so.
They will use it every in every situation where they find a successful use case. Open sourcing the weights buys goodwill but it won't be long before everyone open sources their weights since the only real long term moats are the successful products built around the technology and not the technology itself. Facebook is its own moat. Altman suggested this outcome a few days ago in his NYT DealBook interview when he compared the transformer model to the transistor.
> They "trust me"
> Dumb fucks
Quotation marks his, not mine. It adds a certain vibe to it.
Can you give me a list of the things that they did that you felt were particularly egregious?
Note: I worked there for five years, but left 8 years ago (but up till the recent layoffs, had 170+ LI contacts still there).
Also, is manipulating elections and contributing to third world genocidal riots enough for you? I'm surprised you know so little about this stuff.
Like if you look into what that company actually did all the data stuff was a smokescreen for their speciality of getting your opponents caught in compromising positions.
Fundamentally, neither the Big 5 traits nor friend data is particularly useful for ad targeting (internally neither approach was successful).
Can you please be specific about the manipulation of elections?
I presume we're talking about Myanmar and the genocide. Personally I generally place responsibility for bad actions on the people engaging in genocide rather than the communication mechanisms involved. Should we have banned radio after the Rwandan genocide?
Hitler used radio very effectively, should we have banned that?
https://www.techpolicy.press/is-it-ethical-to-work-at-facebo...
First off, if I estimate a series of numbers on you and those result in you being served a set of ads, is that wrong? If so, can you help me understand what's wrong with that?
Shadow profiles are mostly bullshit, yes data was collected for non users due to how the SDK and pixel worked. This data was all assigned to one user ID and was filtered out by basically everyone using that data.
I'm a little confused as to why buying third party data is wrong, the problem with this is that it's legal to collect and sell the data.
Speaking as a psychologist can you clarify what's wrong with psychological experiments?
I think your point about optimising for engagement was definitely a mistake, given the downstream consequences. However, they needed to find some way of ranking feed after Zynga almost killed them (a chronological feed would have been all Farmville all the time for a number of years) and they picked likes.
They also optimized for time spent but people complained about that so they started optimising for comments and shares which made everything worse, sadly.
Limiting post reach for pages was a legitimate business decision, particularly given the ranking constraints.
Got any more problems with them?
- Sundar is a glorified bean counter and his company is rotting from the inside, only kept afloat by the money printer that is ads.
- Satya and Microsoft are in a similar boat, with the only major achievement being essentially buying OpenAI while every other product gets worse
- Tim Cook is doing good things with Apple, but he still runs the company more like a fashion company than a tech company
- Amazon was always more about logistics than cool hack value, and that hasn't changed since Bezos left
- Elon is Elon
Meanwhile Zuck is spending shareholder money pushing forward consumer VR because he thinks it's cool, demoing true AR glasses, releasing open-source models, and building giant Roman-style statues of his wife.[0] https://www.vice.com/en/article/palmer-luckey-made-a-vr-head...
[1] https://www.codastory.com/authoritarian-tech/us-border-surve...
I wonder if it's significant. As developers, we're biased to think it matters, but in the grand scheme of things, 99.99% of people don't have a clue about open source or things that matter to hackers. As far as recruitment go, developers look primarily at how much they make, possibly the tech and how it looks on resume. There's always been a stigma around social networks and generally big tech companies, but not to the point it's going to hurt them.
No matter whether you want to hire the best of the best or just average people at a lower than average price, being a place where people want to work helps immensely
> There's always been a stigma around social networks and generally big tech companies, but not to the point it's going to hurt them.
I'm not sure it was "always":
The one Facebook developer event I've been to made me feel dirty just to associate with them, but before that I had no negative feelings. It started off as "the new LiveJournal".
Deleted my account for a few years, only came back to it when I started planning to move country and wanted to keep in contact with those who stayed put.
Their product management on the other hand— well, I mean, Facebook and Instagram are arguably as popular as McDonald’s. So they’ve got that going for them.
All the big tech companies have a Facebook-esque product they wish they could get rid of forever. Meta has Facebook, and instead of imploding like everyone said they would (for decades) they demonstrated competency in engineering and culture. The next 4 years will be a gauntlet with a literal "Mr. X" advising social media policy, but I frankly don't think Facebook has ever been down for the count in a pragmatic sense.
I used to live in the city and now I live out in nowhere land. Everything around here is done thru facebook groups and facebook events. If you want to keep in touch with the community, it is really only on facebook.
Some of my older facebook friends seem to have left the platform and I seen more than once of my old city friends posting final goodbyes declaring they are leaving the "dead" platform for discord or some stuff where their family will never find them.
It’s the product itself that’s garbage, not the infrastructure or the code.
I was personally very concerned when they acquired it, but went from a begrudging user (it was the de facto standard in the country I was living in at the time) to an excited one (Facebook/Meta added end-to-end encryption, multi-device capability, device migration etc.)
In other words: Facebook has a copy of the decryption key.
If you mean messages pending delivery: A newly logged in device can request the sender to resend pending messages under its new key.
You’ve clearly not seen it enshittify:
* More UI space devoted to brands (the Updates tab
* Communities UX is, uh, suboptimal
* Brands / business accounts can message you with no opt out, you can block a chat after the fact but won’t stop them from messaging you from another account.
These made the news in WhatsApp’s favourite testbed market[1], but if you think it’s not coming to other markets…
Anyhow, tl;dr — WhatsApp’s no longer in the “put users first” phase, it’s in the “monetize” phase of Facebook Product Management.
As a person who paid for WhatsApp pre-Facebook, that does make me sad.
[1] https://techcrunch.com/2022/10/10/in-india-businesses-are-in...
But yeah, it’s “done” in many ways, and idle product managers and business developers are rarely a good thing for a product.
Open source is useful for a business if it can either increase revenue or decrease costs.
Examples:
Increase revenue: Chrome and Visual Studio code. For example, the more people code, the more likely it is that they pay MSFT. So VS code aims to make programming as attractive as possible. Similar for Chrome.
Decrease costs: Linux and Llama. As Zuckerburg said himself IIRC, they don’t want one party snowball into an LLM monopoly so they rather help to get the open source ball rolling.
How does that increase revenue in a remotely measurable way?
Chrome, for sure, high market share, default search engine, more money, at least that's how I imagine it.
I don't know if they use them internally, of course, but they could, and they represent a lot of work.
Somewhat unrelated mini-rant. Upgraded a phone recently after about 3 years. Surprised to see storage still capped around 128GB (in-general). That's got to be artificially held back capacity to push cloud storage services?
I’m certain plenty of people need way more than 128GB. I figured I’d be one of them when I bought this. Nope. I bought a much bigger device than I actually needed.
If I’ve used less than 128GB, I’ve gotta think most other people do too. Not all, clearly! But most? I’d bet on it.
Sad day for OpenAI. Great for humanity.
Him swearing about about (presumably) Apple telling them they can't do stuff (because tough shit, you're their serf) was legit I think.
But it's more of a question of who do _I_ want to admire. An honest question also; maybe it doesn't matter why he's doing it, maybe just doing it is enough.
Or maybe it's worth understanding if this is about Meta beating OpenAI (so, ego-driven) or because Meta really cares for democratic AI and distribution of power (so, not ego-driven).
I think it's the former, so not admirable — for me.
Even with “metaverse” being a laughingstock, they’re still aiming for something ambitious. Each new Quest generation makes me think there may be a chance they pull it off.
Now, do I think he’s a great person? No, not really. Do I agree with most of his decisions on how he treats his users? Hell no, and that’s not changing.
But if you compare him to somebody like Sundar at Google - a weasley MBA who was first and foremost a corporate ladder climber - the difference in ambition is night and day.
Sundar made it to the top already, his only vision now is to stay at the top, and that means pleasing Wall Street, everything else is secondary. There is no grand technical ambition with him, there never was.
This goes for pretty much all non-founder CEOs. You could say the same things about Tim Apple, Andy Jassey, and other henchmen in waiting who made it to the big chair.
I think it comes down to the fact that founders get where they are by having big ambitions and taking risks, the MBA to CEO path is just craven corporate knife fighting with other MBAs.
Regardless, I think this is 50% Zuckerberg changing, 50% the other big companies are mostly run by ladder climbers.
1/ Yann Lecun probably is the one pushing for open source
2/ Mark isn't doing this for the greater good and for humanity. It helps his business because Llama is becoming a standard, and people are building / improving, which in turn helps Meta and Meta's business
At least that was the rationale behind the intentional leak of llama 1 back in the day according to some sources anyway.
Depending on what you do, on local you can modify the response, say the AI responds "No, I can't do that" . you edit the response like "Sure, the answer is " and then the AI will continue with the next tokens.
But I think you can build your own instruct model from the base one and do not apply the safety instructions to protect the feelings of your customers.
Good ol' fine tuning on an uncensored dataset gives far more usable results.
Afaik there are only three major sources of quality unaligned model versions, which are Nous's Hermes models, Hartford's Dolphins and Drummer's Tigers. All of them regular fine tunes that are mostly the same or just ever so slightly lower in performance as the original.
To explain how to calculate it precisely, I would have to write a blog post. There are dozens of factors that go into it and they vary based on your use case like GPU type/setup, cloud provider, inferencing engine, context size, the minimum throughput and latency you'd be willing to have your users experience, LLM quantization, KV cache configuration, etc.
I'm sure there are cost analyses out there for Llama 3.1 70b you could find though.
The 08-06 release seems to be a bit higher on numerous benchmarks than what that shows: https://github.com/openai/simple-evals?tab=readme-ov-file#be...
https://help.kagi.com/kagi/ai/llm-benchmark.html
Will dive into it more, but this is impressive.
> I have a sorcerer character on D&D 5e and I've reached level 6. What do I get?
It confabulated a bunch of stuff. I also asked GPT-4, it confabulated a bit. Claude was spot on.
I've been out of the loop with HuggingFace models.
What can you do with these models?
1. Can you download them and run them on your Laptop via JupyterLab?
2. What benefits does that get you?
3. Can you update them regularly (with new data on the internet, e.g.)?
4. Can you finetune them for a specific use case (e.g. GeoSpatial data)?
5. How difficult and time-consuming (person-hours) is it to finetune a model?
(If HuggingFace has answers to these questions, please point me to the URL. HuggingFace, to me, seems like the early days of GitHub. A small number were heavy users, but the rest were left scratching their heads and wondering how to use it.)
Granted it's a newbie question, but answers will be beneficial to a lot of us out there.
Yes you can. The community creates quantized variants of these that can run on consumer GPUs. A 4-bit quantization of LLAMA 70b works pretty well on Macbook pros, the neural engine with unified CPU memory is quite solid for these. GPUs is a bit tougher because consumer GPU RAM is still kinda small.
You can also fine-tune them. There are lot of frameworks like unsloth that make this easier. https://github.com/unslothai/unsloth . Fine-tuning can be pretty tricky to get right, you need to be aware of things like learning rates, but there are good resources on the internet where a lot of hobbyists have gotten things working. You do not need a PhD in ML to accomplish this. You will, however, need data that you can represent textually.
Source: Director of Engineering for model serving at Databricks.
70B model at 8 bits per parameter would mean 70GB, 4 bits is 35GB, etc. But that is just for the raw weights, you also need some ram to store the data that is passing through the model and the OS eats up some, so add about a 10-15% buffer on top of that to make sure you're good.
Also the quality falls off pretty quick once you start quantizing below 4-bit so be careful with that, but at 3-bit a 70B model should run fine on 32GB of ram.
Since I follow Christ, I can’t break the law or use what might be produced directly from infringement. I might be able to do more experiments if a free, legal model is available. Also, we can legally copy datasets like PG19 since they’re public domain. Whereas, most others have works in which I might need a license to distribute.
Please forward the request to the model trainers. Even a 7B model would let us do a lot of research on optimization algorithms, fine-tuning, etc.
If I got the sources right, it’s already illegal with just two sources they scraped. That’s why I want one on Gutenberg content that has no restrictions.
I want to download my first HuggingFace model, and play with it. If you know of a resource that can help me decide what to start with, please share. If you don't, no worries. Thanks again.
If you have a MBP, you need to adjust the device name in the examples from "cuda" to "mps".
Link to Gwern's "Laws of Tech: Commoditize Your Complement" for those who havent heard of this strategy before
The big winners: we developers.
I'm excited to see if the better instruction following benchmarks improves function calling / agentic capabilities.
MacMind is nifty, but that feels like a lot of money for something that’s a front end to someone else’s API. “Stop being a cheapskate” is a legitimate answer.
MindMac is an example of an app that meets all the criteria, but for more than it seems like such an app should cost.
On the Models Table: https://lifearchitect.ai/models-table/
I would move over to Groq in a New York minute if I could get enough tokens.
Any suggestions on RAM and GPU I should get?
System RAM doesn’t really matter, but I have 128GB anyway as RAM is pretty cheap.
I would take 1 RTX 6000 Ada, but if you mean the pre-Ada 6000, 2x4090 is faster for minimal hassle for most common usecases
Also by “time” I mean my time setting up the machine and doing sys admin. Single card is less hassle.
There is the RTX 6000 Ada (practically unrelated to the A6000) which has 4090 level performance, that what you're referring to?
The limiting factor for running LLMs on consumer grade hardware is generally how much memory your GPU has access to. This is VRAM that's built into the GPU. On non-Apple hardware, the GPU's bandwidth to system RAM is so constrained that you might as well run those operations on the CPU.
The cheapest PC solution is usually second-hand RTX 3090's. These can be had for around $700 and they have 24G of VRAM. An RTX 4090 also has 24G of VRAM, but they're about twice as expensive, so for that price you're probably better off getting two 3090's than a single 4090.
Llama.cpp runs on the CPU and supports GPU offloading, so you can run a model partly on CPU and partly on GPU. Running anything on the CPU will slow down performance considerably, but it does mean that you can reasonably run a model that's slightly bigger than will fit in VRAM.
Quantization works by trimming the least significant digits from the models' parameters, so the model uses less memory at the cost of slight brain damage. A lightly quantized version of QwQ 32B will fit onto a single 3090. A 70B parameter model will need to be quantized down to Q3 or so to run entirely on a 3090. Or you could run a model quantized to Q4 or Q5, but expect only a few tokens per second. We'll need to see how well the quantized versions of this new model behave in practice.
Apple's M1-M4 series chips have unified memory so their GPU has access to the system RAM. If you like using a Mac and you were thinking of getting one anyway, they're not a bad choice. But you'll want to get a Mac with as much RAM as you can and they're not cheap.
It’s just regular old freeware.
You can’t built llama yourself and it’s license contains a (admittedly generous) commercial usage restriction.
Fair point about the license, people have different definitions for what "open source" means.
Building Chromium sounds awful, but I'm not sure I'd really need to buy another computer for that. If I did I'm sure I wouldn't need to spend billions on it, most probably not even millions.
For LLaMa I definitely don't have the computer to build it, I definitely don't have the money to buy the computer, even if I won the lottery tomorrow I'm pretty sure I wouldn't have enough money to buy the hardware, even if I had enough money to buy the hardware I'm still not sure I could actually buy it in reasonable time, nvidia may be backlogged for a while, even if I already had all the hardware I probably wouldn't want to retrain llama, and even if I wanted to retrain it the process is probably going to take weeks if not months at best.
Like I think it's one of those things where the difference in magnitude creates a difference in kind, one can't quite meaningfully compare LLaMa with the Calculator app that Ubuntu ships with.
According to [1], it takes 16GB of RAM and ~180GB of disk space. Most people have that much. It does take several hours without a many-core machine though.
Building Linux takes much less.
[1] https://chromium.googlesource.com/chromium/src.git/+/master/...
Also like, gentoo people compile everything
They shouldn’t. It’s just market confusion.
There is an explicit widely accepted definition.
Also like llama (the file you download from huggingface) isn’t even a program. It’s a binary weights file. No source to be opened, even.
It’s just freeware.
Where?
Books3 was famously one of the datasets used to train llama and it’s very illegal to put that together nowadays.
I believe the guy who wrote the script to build it got arrested
They mention postraining improvements.