Llama 4 Now Live on Groq
groq.com
groq.com
EDIT - Seems that Groq has stopped selling their chips and now will only partner to fund large build outs of their cloud [2].
0 - https://groq.com/the-groq-lpu-explained/
2 - https://www.eetimes.com/groq-ceo-we-no-longer-sell-hardware
It's not - it's absolutely a vanishingly small market.
https://www.nvidia.com/en-us/products/workstations/dgx-spark...
So, an Apple Mac Studio?
They stopped selling the hardware to the public, and it takes an extraordinary amount of it to run these larger models due to limited ram.
we have a pretty generous free tier and a dev tier you can upgrade to for higher rate limits. also, we deeply value privacy and don't retain your data. you can read more about that here: https://groq.com/privacy-policy/
All three of those can also be accessed via OpenRouter - with both a chat interface and an API:
- Scout: https://openrouter.ai/meta-llama/llama-4-scout
- Maverick: https://openrouter.ai/meta-llama/llama-4-maverick
Scout claims a 10 million input token length but the available providers currently seem to limit to 128,000 (Groq and Fireworks) or 328,000 (Together) - I wonder who will win the race to get that full sized 10 million token window running?
Maverick claims 1 million and Fireworks offers 1.05M while Together offers 524,000. Groq isn't offering Maverick yet
Notably the max sequence length in training was 256k, but the native short context is still just 8k. I'd expect the retrieval performance to be all over the place here. Looking forward to seeing some topic modeling benchmarks run against it (ill be doing so with some of my local/private datasets).
[1] https://github.com/meta-llama/llama-models/blob/eececc27d275...
EDIT: should be fair/complete and note they do claim perfect NIAH text retrieval performance across all 10M tokens for the Scout model on their blog post: https://ai.meta.com/blog/llama-4-multimodal-intelligence/. There are some serious limitations and caveats to that particular flavor of test though.
Would you mind expanding on this? Or point to a reference or two? Thanks! I am trying to understand it.
There is a wealth of literature to catch up on to understand the performance motivations behind those choices, but you can think of it as essentially a balancing act. They want to extend the context length, which is limited by conventional attention compute scaling. RoPE on the other hand is a trick that helps you to scale attention to longer context, but at the cost of poor retrieval across the entire context window. This approach is a hybrid of those two things. The recent Cohere models employ a similar methodology.
Otherwise yes there are lots of papers on this and related topics, a few dozen in fact. But here are some notable ones, a couple of them are linked in their blog post.
RoFormer: Enhanced Transformer with Rotary Position Embedding - https://arxiv.org/abs/2104.09864
Scaling Laws of RoPE-based Extrapolation - https://arxiv.org/abs/2310.05209
The Impact of Positional Encoding on Length Generalization in Transformers - https://arxiv.org/abs/2305.19466
Scalable-Softmax Is Superior for Attention - https://arxiv.org/abs/2501.19399
A colleague from a discord I spend time in threw together this video a year or so ago, might be helpful as a first watch before a deep dive: https://www.youtube.com/watch?v=IZYx2YFzVNc
Covers positional encoding as a general concept first, then goes into rotary embeddings.
Very few of the models supported on Groq/Together/Fireworks support function calling. And rarely the interesting ones (DeepSeek V3, large llamas, etc)
It places more of a "mental burden" on the model to output tool calls in your custom format, but it worked enough to be useful.
i'd actually love to hear your experience with llama scout and maverick for function calling. i'm going to dig into it with our resident function calling expert rick lamers this week.
$0.11 per 1M tokens, a 10 million content window (not yet implemented in Groq), and faster inference due to fewer activated parameters allows for some specific applications that were not cost-feasible to be done with GPT-4o/Claude 3.7 Sonnet. That's all dependent on whether the quality of Llama 4 is as advertised, of course, particularly around that 10M context window.
AMD MI300x has day zero support to run it using vLLM. Easy enough to rent them for decent pricing.
Is there some technical limitation on the context window size with LPUs or is this a temporary stop-gap measure to avoid overloading groq's resources? Or something else?
Out of curiosity, the console is letting me set max output tokens to 131k but errors above 8192. what's the max intended to be? (8192 max output tokens would be rough after getting spoiled with 128K output of Claude 3.7 Sonnet and 64K of gemini models.)
Will have Llama 4 Maverick running in 4bit quantization (typically results in only minor quality degradation) once llama.cpp support is merged.
Total hardware cost well under $50,000.
The 2T Behemoth model is tougher, but enough Blackwell 6000 Pro cards (16) should be able to run it for under $200k.
That said, I'm obviously biased but you're probably better off renting it.