Seems to perform on par with or slightly better than Llama 3.2 405B, which is crazy impressive.
Edit: According to Zuck (https://www.instagram.com/p/DDPm9gqv2cW/) this is the last release in the Llama 3 series, and we'll see Llama 4 in 2025. Hype!!
Seems to perform on par with or slightly better than Llama 3.2 405B, which is crazy impressive.
Edit: According to Zuck (https://www.instagram.com/p/DDPm9gqv2cW/) this is the last release in the Llama 3 series, and we'll see Llama 4 in 2025. Hype!!
If we can get better quantization, or bigger GPU memory footprints, we might be able to use these big models locally for solid coding assistants. That's what I think we have to look forward to (among other benefits) in the year(s) ahead.
1. lmstudio-community/Llama-3.3-70B-Instruct-GGUF/Llama-3.3-70B-Instruct-Q4_K_M.gguf
Edit: Perhaps answering my own question:
λ ollama run hf.co/lmstudio-community/Llama-3.3-70B-Instruct-GGUF:Q4_K_M
>>> Hi. Who are you?
I'm an artificial intelligence model known as Llama. Llama stands for "Large Language Model Meta AI."
System Prompt: You are an advanced super intelligence, capable of everything from software development to dream interpretation.
>> Who are you?
I am an artificial intelligence designed to simulate human-like conversation and answer questions to the best of my knowledge. I'm a large language model, which means I've been trained on a massive dataset of text from various sources, including books, articles, research papers, and websites.
My capabilities are diverse and extensive. I can:
1. **Answer questions**: I can process natural language queries and provide relevant answers, covering a wide range of topics from science and history to entertainment and culture.
2. **Generate text**: I can create text based on a prompt or topic, including articles, stories, emails, and even entire books.
3. **Translate languages**: I can translate text from one language to another, including popular languages such as Spanish, French, German, Chinese, and many more.
4. **Summarize content**: I can summarize long pieces of text into concise, easily digestible versions, highlighting key points and main ideas.
5. **Offer suggestions**: I can provide suggestions for topics such as gift ideas, travel destinations, books to read, and more.
6. **Engage in conversation**: I can have a natural-sounding conversation with you, using context and understanding to respond to your questions and statements.
7. **Dream interpretation**: As you mentioned, I can even help interpret dreams, providing insights into the symbolism, themes, and emotions that may be present in your subconscious mind.
8. **Software development**: I can assist with software development tasks, such as writing code, debugging, and optimizing algorithms.
My knowledge is constantly evolving, and I'm always learning from the interactions I have with users like you. So, feel free to ask me anything – I'll do my best to help!With Llama specifically I recall someone comparing various quants on perplexity finding that even at Q3, 70B is still smarter than 34B. So quantization is generally worthwhile so long as you have a larger model that you can squeeze into your VRAM budget with it, and don't mind the slowdown from more parameters.
I hope Llama 4 reintroduces that mid sized model size.
If you want to offload fully to VRAM, I'd say 8B is the limit. If you're keeping some on RAM, 15-20B can still give OK performance, depending on your tolerance.
>How much does quantization impact output quality?
Basically with more quantization the output becomes more incoherent and less realistic. At the extreme end it's basically just gibberish. I think the sweet spot generally is at 4 bits. At that point the model is pretty compact and the quality isn't diminished too much.
Speed takes quite a hit once you have a few layers on the CPU, but depending on needs it can be doable. I've just asked LLama 3.3 70B Q5_K_M a question and it offloaded about 5 of the 80 layers, so running almost entirely on my 5900X CPU, but still churning out about one word per second.
In my experience quantization affects prompt adherence primarily and answer accuracy secondarily. For example, if you have multiple clauses, ie one or more "if this then that", then quantization might get it to not consider those. I also find they tend to answers more generally and less precise at higher quantization levels.
As a concrete example, I've been asking the LLama 3.2 Vision 8B model to categorize some images. The default instruct model in general has been heavily trained to output general commentary on the image. If in the prompt I tell it to "output the category only", the Q4_K_M variant sometimes ignores that instruction, while the Q8 variant almost always respect it.
Larger models primarily bring more knowledge in my experience, but usually also better prompt adherence. Larger models also typically can support larger contexts, though this can vary, check the model cards.
edit: I should clarify. More knowledge also often translates to better, more accurate output. For example, a larger model might recognize an idiom and answer accordingly, while the smaller model fails to recognize it and thus provides a poor answer.
Depending on your needs, 12GB might be quite decent or it might be insufficient. If you need an assistant-like model, I liked the Gemma 2 9B Q5_K_M. And I've been quite impressed by LLama 3.2 Vision 8B Q4_K_M for describing images and transcribing text from images.
But for more open-ended stuff, especially if larger contexts is needed, I think you might find it underwhelming.
But I do like to compare. Open WebUI for example makes it very easy where you can load up multiple models and it'll send the same prompt to each one in turn, and show the answers side by side.
bash> time ollama run llama3.3 "What's the purpose of an LLM?" | tee ~/Downloads/what\ is\ an\ LLM.txt
A Large Language Model (LLM) is a type of artificial intelligence (AI) designed to process and understand human language. The primary purposes of an LLM are:
(... contents excerpted for brevity) Overall, the purpose of an LLM is to augment human capabilities by providing a powerful tool for understanding, generating, and interacting with human language.
real 0m59.040s
user 0m0.071s
sys 0m0.081s
pmarreck 59s35ms
20241206220629 ~ bash> wc -w Downloads/what\ is\ an\ LLM.txt
359 Downloads/what is an LLM.txtLooks like LM Studio is available for ARM based Macs, if you want to give that a try, that'd be one way to get these stats. LM Studio also surfaces up some parameters to play around with, and keeps a record of past conversations if that might appeal to you.
but also wrestled with this mentally.
Meta both improves the technology or inference, while also trapping themselves alongside every other person training models to always update the training set every few months, so it knows what its talking about with relevant current events