Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
github.com
github.com
What you are saying is like me losing 13kg is as impressive as having the potential to do so if only I would stop eating ice cream and too many snacks every day.
The potential to lose weight, in this context, is the original post because there was potential but it wasn't shown
This commenter is the person who lost 13kg.
The technology is not different but the demonstration is more impressive
What does tjis mean? Another generative model?
I don't think that makes it a dead end. Laya is designed to be fine-tuned for a specific task, and the site reports the fine-tuning gains. The 0.766 number comes from fine-tuning on the benchmark's train split, not from the base checkpoint. They also report that fitting a single temperature scalar per question type cuts expected calibration error from 0.466 to 0.081. That's a large gain, and it only shows up after you specialize the model.
"Peer" is doing a lot of work in that comment. Jev can take on a new task without retraining because it starts with far more knowledge; Laya trades that away to stay small and trainable for a fixed task. So I wouldn't compare base Laya to Jev and stop there. Compare Jev to fine-tuned Laya on the same task and test set, then look at accuracy, latency, cost, calibration, and robustness, depending on which of those matter for the deployment.
My main point of disagreement would be that I fundamentally see a different use for a jev sort of model (generalism is applealing), but if you're finetuning, a bert base is not bad.
Anyways, thanks for the vouching!
However at this point I talk to LLMs more than anyone except probably my wife. As a multiple times immigrant, I can absolutely believe I'm adjusting my speech patterns to its vernacular.
It's turning into pimp my llm...
For clarity: no, a 0.8B model is not gonna do that.
Boosted trees over embeddings work surprisingly well too.
I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.
Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.
I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes
The classifiers also run in <1ms, so they can be very fast and precise at the same time
But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)
For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data
Have I understood correctly that you trained only the logistic classifier, but didn't need to train the embedding model?
If so, I'm curious whether you compared that approach (A) with:
B) Jev only, with a single output.
C) Jev with multiple outputs fed into a logistic classifier.
Obviously C has cons (can't be self-hosted, needs some up-front work on deciding the shape of the output) but it might be somewhat more interpretable. (And I suppose it might have better performance?)
Here's a gist with code you can use to test the Banking77 dataset: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller
I haven’t compared different ways of sending requests to Jev
The data to train the classifiers comes from the datasets used to test them (not from Jev)
My understanding of Jev is that it’s a replacement for the LLM you’d necessarily need to use to identify reasoning-sensitive workloads in a heterogeneous mix, where Jev will be cheaper than an actual LLM and so act as an actual optimization / de-bottlenecking change.
Depending on how much overfit, you can go from routing deterministically based on features/shape of the input data, all the way to training a routing model (which could be a classifier too). I’ll need to experiment to find the best approach
For completely unseen/unexpected, I’ve also experimented routing to a local LLM: request comes in, if there’s a marching classifier, send it there, otherwise send to LLM+training. As the system learns more tasks, the % of requests that go to the LLM go down over time
You can already so that with classification models such as ModernBERT, at 0.4B.
Jev's value is its zero shot performance without having to fine-tune.
On top, most EMs wouldn’t take a risk on an exploration of something “unknown” (to them) and couldn’t get buy-in from a PM.
I say this as an EM. Interview hundreds of people and, while, yes, some people don’t interview well, you might be shocked at the level of creative thinking. Even when “creative” is narrowly scoped to “this is a solved problem in a related domain.
It's more likely the product is focused on something else, but a classification model could come in handy...
There is a fixed cost (and some maintenance) to e.g. fine tuning ModernBERT.
Maybe once you include all of that it might be a half-day to a day of engineering time to set everything up in a maintainable fashion.
For Jev, it takes all of 30 seconds of prompting. And it's not that much more expensive to deploy vs. a BERT model.
What I said will only make sense if you take yourself out of your current context and think entirely from the perspective of someone who knows little-to-nothing about ML.
It’s the same mistake folks on HN made when Dropbox launched, drawing comparisons to rsync and other Unix tools as if they were somehow equivalent.
In practice, that's enough of a barrier to not even try the approach on a number of cases where it might potentially be useful.
I wouldn't be surprised if Jev turned out to be a "gateway drug" that validates approach on a use case, the team gathers experience and labeled data, and switches to an in house locally tuned model to minimize costs.
I did side by side comparison with Gemini 2.5 Flash Lite, Jev, Jeff
I tried the 0.8B model, completely useless in classification. Qwen Jeff-Qwen3.5-2B was better, but still missed job type.
I suppose with larger model, this could be useful, but would require more ram and will be slower.
they all suck
HN's desire to pretend the data pipeline doesn't exist or isnt meaningful is silly.
I do note though, that is infinitely easier to 'deploy' a Jev/LLM based solution than a data/model pipeline.
Jev is basically a kind of FLAN-BERT, if you want, where it has built-in multi-task ability, but doesn't generate text. It only generates 255 floats all at once, making it much faster, and what those floats mean (if anything) depends on the prompt.
Eg, the following query is put in the encoder model:
{"question": "Rank these 5 things by increasing order of how big they are", "choices": ["truck", "cow", "mouse", "ant", "building"] }
The model returns [3., 2., 1., 0., 4.], and 249 other meaningless floats that are hidden from you by the UI.
The UI stitches the first 5 floats with the choices and returns something like:
{"rank": ["ant", "mouse", "cow", "truck", "tower"]}
There was a github repo but I have not checked it.
edit I see, thats an unaffiliated site that latched onto that, my bad, here's the GH that I should've linked: https://github.com/vosen/ZLUDA
(Though it could turn into "nvidia pay lots of people to use LLMs to make non-portable, tightly coupled backends to _even more_ open source projects")
It’s just that Nvidia’s stuff works, and available at scale, and includes the full stack with networking, cooling and such.
If you can make them work Arc cards are a massive bargain. On raw compute the silicon is not bad.
In fact NVIDIA wasn't even a close competitor to 3DFX in the graphics card game at that point
It's not like 3DFX was the first GPU vendor. Nvidia saw the opportunity to be the first true GPGPU vendor, and they beat their competitors.
ATI (AMD) in 1989 copied the unpatentable parts of the 8514/A, improved it, and went on to dominate 2D accelerators.
nVidia’s first 3D card was a complete flop. They did not achieve success for years.
Nvidia is in the right place at the right time.
You should look at the IA/A - it had its own C like compiler, CPU, etc which did things reminiscent of a modern GPU or SIMD.
The momentum is with CUDA because CUDA is the most broadly usable one. Especially with AI-driven optimization loops and similar APIs, competitors can more easily pick up momentum, if they'd actually try.
1M x $0.50 == 1B x $0.0005
For everyone else who is conscious of cost, you're already seeing this being built into harnesses.
Almost certainly, you'll see versions of this from all the Chinese labs as fast as humanly possible.
If I had to guess, Cursor/Grok or Google/Antigravity will be the first major players to natively support something like this to drive down cost, as they're primarily the budget conscious choices.
I would be astounded if Anthropic leads the way on a cost reduction.
OpenAI and Anthropic don't want to give out logprobs these days but could trivially add a dedicated classification API to their existing models if there was enough demand.
I mean, Jev is also probably cheaper because it's a rather small model (or at least, I suspect it is based on the overall level of intelligence it demonstrates) so that helps make it cheap too.
Thats also why it feels weird they say they don't "charge for output tokens" since its literally generating a single (or at most very few tokens).
It also returns confidence scores for all choices.
Granted, they are not stable. They fluctuate even when you reorder choices, but it still counts as an additional feature.
Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?
edit: Where I'm coming from, I can triage 10,000 support tickets with deepseek flash for less than $1, and latency is sub 1 second if it needs to be integrated into a live user flow. I don't need anything cheaper or faster than that.
As a closed source chat-API provider, you just need to find a way speak JSON correctly at API output.
But both LLMs and Jev-like models would need to prefill, the only optimization Jev does differently is the decode which can be emulated by reading off logprobs.
We don’t know the param size of Jev, to determine the most comparable model, but if I had to guess it’s sub-100B.
Orders of magnitude faster and cheaper answers.
Sure a MacBook Pro can control a servo motor but maybe an arduino or a cheaper microcontroller for deployment?
not sure what made you think you cant go cheaper than jev. it is being done for a long time.
Once you discover what the useful problem to solve is and how to utilize it in production, you can of course replicate it.
This is not alchemy taught by some wizard in secret.
Very often, the innovation is in identifying what to build, what has interesting use cases, what people will pay for.
Once this is established, there is further research optimizing it further, replicating it locally, etc.
The fact that anyone could have come up with it is irrelevant.
The same reason I want diffusion language models to be mainstream
The aspect that I like the most is the typesafe API that introduces new probabilistic concepts that are more sound than json schema and constrained decoding with quasi-confidence scores. Developers were asking LLMs to also emit confidences which made absolutely no sense whatsoever.
This looks pretty decent: https://www.aidancooper.co.uk/constrained-decoding/
No comment on this model itself, might be over fitted to the Jev benchmarks just to beat it.
Name Entity Recognition (NER) is one example.
So many of them... https://huggingface.co/models?language=ner&sort=trending
Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don't want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).
The nice thing about Jev is that people started taking about models that are not LLM text streams again.
This is also bogus unless you are talking about Bayesian inference. No classifier can output CI for a single point estimate. In every ML theory textbooks worth their $, it's always stressed not to treat these sigmoid'ed or softmax'ed numbers as probabilities or confidence scores, there is no such thing as CI for point estimate.
That said, I think the advantage of Jev style approaches is not their capabilities, but rather the capabilities that they have for a much lower resource requirement.
Nowadays any LLM and any harness you use will just do this for you.
But there are helper libraries like Instructor that have been around since like gpt3, which abstract away retries and stuff to make this super easy.
Why does it sound like that to you?
The Von numbers have led me on a rabbit whole of getting a classifier to play Doom
I got it to average 22 kills (the max is 26) on that same scenario that Jeff and Von are testing on (it’s called Defend Center)
Now I’m having it play a more advanced scenario, and it’s doing about 45 kills (SOTA is ~59 kills)
It’s amazing what you can do with small classifiers if you can collect some data. These models I’m testing train on CPU in seconds (what takes the longest is running the game, doing test runs and collecting data), they are <1MB in size and do inference in <1ms on CPU
Edit: after looking at Jeff's numbers more in detail, the 6.5 kills number is not that bad, but it can definitely be better ;)
(Specifically, there are two O(n^2) steps in an attention layer and KV caching makes the first one O(n) with caching - but the overall big-O is still O(n^2) because KV caching doesn't affect the second step.)
softmax(QK^T/sqrt(d_k))V
So without KV-caching, if suppose you have N tokens, then you have Q: N x q_dim (leaving the batch size and n heads out for simplicity, they are constants for our purposes here)
K^T: k_dim x N
So QK^T: N x N <- this is the attention matrix
The scaling and softmax are irrelevant here, since they are elementwiseThen
V: N x v_dim
(QK^T)V: N x v_dim
This gives us the normal quadratic complexity of prefills (of course, in an actual implementation an attention mask is used to ensure that tokens cannot attend to prior tokens and you may have things like ALiBi).Note that during decoding, the dimensions change. Since it is an autoregressive model, we do not need to recompute the values and keys of prior tokens, only of the token that we are currently decoding. Of course, the token that we are decoding still attends to all prior tokens and itself.
Q: 1 x q_dim
K^T: k_dim x (N+1)
QK^T: 1 x (N+1) <- Note that attention in this step is not quadratic anymore, since
we only need to compute how the current token attends to prior
tokens, the representations of prior tokens are frozen. K comes
from the KV-cache.
V: (N + 1) x v_dim <- V comes from the KV-cache
(QK^T)V: 1 x v_dim <- Also not quadratic, the value is only computed for the current token.
So attention during a decoding step is O(N), so when decoding N tokens, it is O(N^2) overall.The point-wise feed-forward layer does not matter, in decoding it only needs to be computed for the representation of the token that we are currently generating. We don't need the representations of the preceding tokens for the next layer, since we have already cached their keys and values for each layer.
Disclaimer: I was one of the developers of a widely-used inference engine and implemented several of these optimizations.
My hypothesis for Jev is that they simply generate many answers independently in parallel from your prompt, and then discard the duplicates (or train to avoid duplication with attention between the ). In that way the entire batch is 1 user's prompt.
Those are telltales of a Bert model.
IMO you going to get better results by doing a small tune of a tiny model. Making training dataset for it with LLMs is easy, serving it is going to be cheaper.
Edit: Running them for the masses.
I've been holding it in since I saw the title. I was surprised it wasn't all over repo!
When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware. So everything ran at home: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing, all monitored from my phone over Tailscale.
The caveat: the published Jev and AutoJev numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86-89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64-68% against Jev's 94%, and about 50% on JevBench's hard tier against 73%. That isn't surprising, and I don't think it matters: no 0.8B or 2B model reasons like a large one, and nobody should expect it to. These are extremely fast judgement-callers. In one of my apps I use the 0.8B for voice navigation; a quick fine-tune (about half an hour on one GPU) took it from 32% to 96% on held-out commands, at about 40 ms per decision.
The fun part: games, as a zero-shot test. Games aren't the ideal zero-shot test, but they're fun, and TypeSafe did it with Jev too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one; the options say what each move leads to, never which one is right. Over 20 episodes each:
- Doom (ViZDoom): Jeff 0.8B 6.55 kills per episode, the same as a hand-coded bot and as Jev's published run. Jev's prompt spells out the aiming rule and takes about 212 ms per call over its API; Jeff gets "the nearest monster is a little to your left" and decides in about 29 ms on my Mac.
- Frogger: 10.3 crossings, level with the hand-coded bot (10.25), and 10x the untrained base model (1.0).
- Pac-Man: 57 of 98 pellets, about 60% of the bot's score and 2x the untrained model.
Videos of every run are linked in the README.
Lessons learned:
- System 1 models are here to stay. Being able to process unstructured data at software speed inside an app is extremely powerful, and being able to do it locally is fantastic.
- A small model is a classifier, not a planner. Models of 0.8B-2B don't reason like Qwen3.8-27B or Jev, and they don't need to: present the options the right way and you get 40+ decisions per second, depending on your hardware.
- Fine-tune it if needed. If zero-shot isn't good enough for your task, a short fine-tune on your own examples is.
- Wording matters enormously. Giving Frogger's final step the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Before that, the frog just stayed on the last log.
- Bigger isn't better. The untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers turning away from the nearest monster), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.
- Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, but not reliably, and in a real-time loop the mistakes compound.
Would love to see what this run would cost from something like Verda. Just out of curiosity. I'm not going to be installing any sparks at home anytime soon.