EDIT: HN is rate-limiting me so I will reply here: In my opinion 1B and 3B truly shine on edge devices, if not than it's not worth the effort, you can have much better models for already dirt cheap using an API.
EDIT: HN is rate-limiting me so I will reply here: In my opinion 1B and 3B truly shine on edge devices, if not than it's not worth the effort, you can have much better models for already dirt cheap using an API.
Wouldn't they lower the costs compared to big models drastically?
I want an LLM, STT, or TTS model to run efficiently on a Raspberry Pi with no GPU and no network.
There is huge opportunity for LLM-based toys, tools, sensors, and the like. But they need to work sans internet.
It's just like spawning two copies of the same program, doesn't require that you have two copies of the program's text and data sections sitting in your physical RAM (as those get mmap'ed to the same shared physical RAM); it only requires that each process have its own copy of the program's writable globals (bss section), and have its own stack and heap.
Which means there are economies of scale here. It is increasingly less expensive (in OpEx-per-inference-call terms) to run larger models, as your call concurrency goes up. Which doesn't matter to individuals just doing one thing at a time; but it does matter to Inference-as-a-Service providers, as they can arbitrarily "pack" many concurrent inference requests from many users, onto the nodes of their GPU cluster, to optimize OpEx-per-inference-call.
This is the whole reason Inference-aaS providers have high valuations: these economies of scale make Inference-aaS a good business model. The same query, run in some inference cloud rather than on your device, will always achieve a higher-quality result for the same marginal cost [in watts per FLOP, and in wall-clock time]; and/or a same-quality result for a lower marginal cost.)
Further, one major difference between CPU processes and model inference on a GPU, is that each inference step of a model is always computing an entirely-new state; and so compute (which you can think of as "number of compute cores reserved" x "amount of time they're reserved") scales in proportion to the state size. And, in fact, with current Transformer-architecture models, compute scales quadratically with state size.
For both of these reasons, you want to design models to minimize 1. absolute state size overhead, and 2. state size growth in proportion to input size.
The desire to minimize absolute state-size overhead, is why you see Inference-as-a-Service providers training such large versions of their models (OpenAI's 405b models, etc.) The hosted Inference-aaS providers aren't just attempting to make their models "smarter"; they're also attempting to trade off "state size" for "model size." (If you're familiar with information theory: they're attempting to make a "smart compressor" that minimizes the message-length of the compressed message [i.e. the state] by increasing the information embedded in the compressor itself [i.e. the model.]) And this seems to work! These bigger models can do more with less state, thereby allowing many more "cheap" inferences to run on single nodes.
The particular newly-released model under discussion in this comments section, also has much slower state-size (and so compute) growth in proportion to its input size. Which means that there's even more of an economy-of-scale in running nodes with the larger versions of this model; and therefore much less of a reason to care about smaller versions of this model.
Not sure I follow. CoT and go over length of the states is a relatively new phenomenon and I doubt when training the model, minimize the length of CoT is an explicit goal.
The only thing probably relevant to this comment is the use of grouped-query attention? That reduces the size of KV cache by factor of 4 to 8 depending on your group strategy. But I am unsure there is a clear trade-off between model size / grouped-query size given smaller KV cache == smaller model size naively.
Pretend for a moment that Transformers don't actually have context-size limits (a "spherical cow" model of inference.) In this mental model, you can make a small, dumb model arbitrarily smarter — potentially matching the quality of much larger, smarter models — by providing all the information and associations it needs "at runtime."
It's just that the sheer amount of prompting required to get a dumb model to act like a smart model, goes up superlinearly vs. the marginal increase in intelligence. And since (for now) the compute costs scale quadratically with the prompt size, you would quickly hit resource limits in trying to do this. To have a 10b model act like a 405b model, you'd either need an inordinate amount of time per inference-step — or, for a more interesting comparison, an amount of parallel GPU hardware (VRAM to hold state, and GPU-core-compute-seconds) that in both dimensions would far exceed the amount required to host inference of the 405b model.
(This superlinear relationship still holds with context-size limits in place; you just can only do the "make the dumb model smarter with a good prompt" experiment on roughly same-order-of-magnitude-sized models [e.g. 3b vs 7b] — as a 3b really couldn't "act as" anything above 7b, without a prompt that far exceeds its context-size limit — and so, in practice, you can't calculate enough of the ramp at once to fit a curve to it.)
The obvious corollary to this, is that by increasing model size (in a way that keeps more useful training around, retains intelligence, etc), you decrease the required resource consumption to compute at a fixed level of intelligence, and that this decrease scales superlinearly.
This dynamic explains everything current Inference-as-a-Service providers do.
It explains why they they are all seeking to develop their own increasingly-large models — they want, as much as possible, to get their models to achieve better results with less prompting, in fewer inference steps, and in proportionately cheaper inference steps — as these all increase their economies of scale, by decreasing the compute and memory requirements per concurrent inference call.
And it explains why they charge users for queries by the input/output token, not by the compute-second. To them, "intelligent responses" are the value they provide; while "(prompt size + output size) x (number of inference steps)" is the overhead cost of providing that value, that they want to minimize. A per-token pricing structure does several things:
• most obviously, as with any well-thought-out SaaS business model, it pushes the overhead costs onto the customer, so that customers are always paying for their own costs.
• it therefore disincentivizes users from sending prompts that are any longer than necessary (i.e. it incentivizes attempting to "pare down" your prompt until it's working just well enough)
• and it incentivizes users to choose their smarter models, despite the higher costs per token, as these models will achieve the same result with a shorter prompt; will require fewer retries (= wasted tokens) to give a good result; can "say more" in fewer tokens by focusing in on the spirit of the question rather than rambling; and require less CoT-like "thinking out loud" steps to arrive at correct conclusions.
• it also incentivizes the company to put effort into R&D work to minimize per-token overhead, to increase profitability per token. (Just like e.g. Amazon is incentivized to optimize the per-request overhead of S3, to increase the profitability per call.)
• and, most cynically, it locks in their customers, by getting them to rely on building AI agents that send minimal prompts and expect useful + accurate + succinct output; where you can only achieve that with these huge models, which in turn can only run on the huge vertically-scaled cluster nodes these Inference-aaS providers run. The people who've built working products on top of these Inference-aaS providers can't meaningfully threaten to switch away to "commodity" hosted open-source-model Inference-aaS providers (e.g RunPod/Vast/etc) — as nobody but the few largest players can host models of this size.
(Fun tangent: why was it not an existential mistake for Meta to open-source Llama 3.1 405b? Because nobody but their direct major competitors in the Inference-aaS space have compute shaped the right way to run that kind of model at scale; and those few companies all have their own huge models they're already invested in, so they don't even care!)
In a way it also matters to individuals, because it allows them to run more capable models with a limited amount of system RAM. Yes, fetching model parameters from mass storage during inference is going to be dog slow (while NVMe transfer bandwidth is getting up there, it's not yet comparable to RAM) but that matters if you insist on getting your answer interactively, in real time. With a local model, it's trivial to make LLM inference a batch task. Some LLM inference frameworks can even save checkpoints for a single inference to disk and be cleanly resumed later.
When it's behind an API its just a standard margin/speed/cost discussion.