Strands Decider 2B: a small, open-source, decision model
strandsagents.com
strandsagents.com
NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).
I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.
Btw, would love your opinion on this: pre trained classifiers that run and train on CPU https://github.com/nicobrenner/jeffy
Meet Jerry- it's literally a guy named Jerry answering your questions.
Jerry, "The Decider" https://www.youtube.com/watch?v=r8VbzrZ9yHQ
[0] https://news.ycombinator.com/item?id=49723267 (see parent for reference)
So "noul"--short for Bernoulli as in the Bernoulli distribution--"maps to if-statements."
I'd googled "bernoulli map" after reading that, and came across this[0], which I wasn't aware of and thought was somehow related so I entrenched my misunderstanding (I didn't dig in.)
Side note: I just wrote the sentence "you're completely right!", and almost changed it because it sounds like slop now. I wonder if we're gonna get an increase in these sentences in human written prose over time as people mimic these sentences, or a decrease as people shy away from them to not sound like AI. :)
It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?
On humans-for-humans, blog posts with my name on aren't written by AI (https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html) either for my personal blog or at work. I'm a heavy AI user, but this is something I think is best left to humans.
There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence.
We're releasing more of that functionality in strands, and I expect an explosion of innovation in that area.
It is just the right mix of size, capability and speed to make it generally useful for adhoc bulk classification tasks.
For those wanting to run it in browser: https://huggingface.co/alxnahas/strands-decider-2B-webgpu
Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:
$ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose
[...]
File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name
raise ValueError(f"Can not map tensor {name!r}")
ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'I haven't looked in depth, but it should be doable without modifications to llama.cpp. You can't just convert the LoRA, through - there's a whole pointer head and some custom layers that need to be correctly handled (in fact, the LoRA mostly exists to get the base model to behave the way the pointer head needs).
I appreciate what the JevBench folks are doing, but it's true the JevBench (like all benchmarks) isn't representative of every real-world workloads. I particularly don't love how JevBench handles confidence, and rewards being highly confident in certain cases. Benchmarking is hard, and JevBench is probably more representation of its target workloads than TPC-C is :)
As for actual intelligence, there's only so much you can expect from 2B without reasoning. One of the design challenges in training this model was to avoid forgetting too much, and the KL-to-frozen-base step partially exists for that purpose. Even then, Qwen3.5 2B base has fairly limited single-pass reasoning ability (which shows up for us in the performance on the JevBench hard set).
If I have a sizeable amount of labeled data and need decisions calibrated to that data should I
1. Ignore the hype and train a traditional classifier
2. Finetune an LLM based decision model
3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow
If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?
(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.
(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.
(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.
The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.
The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results
Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.
As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).
Everything you need to fine-tune, or even retrain from scratch, is in the github repo.
(Disclaimer, I work at Cloudflare, but not on models)
I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking.
Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.
v21 supports vision: https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...
To vaguely confirm that this is plausible, i tested the sample query from the post with differently arranged options (This should produce slightly different outputs since each option influences all subsequent hidden states, including those of the other options):
"billing,sales,retail" produces billing -> 0.843, retail -> 0.092, sales -> 0.065
"retail,billing,sales" produces billing -> 0.470, retail -> 0.468, sales -> 0.062
"retail,sales,billing" produces billing -> 0.517, retail -> 0.415, sales -> 0.068
"billing,retail,sales" produces billing -> 0.803, retail -> 0.146, sales -> 0.051
"sales,retail,billing" produces billing -> 0.647, retail -> 0.127, sales -> 0.225
So each of these arrangements favor billing, but with pretty different confidence scores, and it looks like there is a pretty big bias toward the first option listed - which makes me wonder how viable it would be to have the torso process these options independently instead
Their docs say:
> An option's score depends on what it says, not where it sits. There is no per-option parameter anywhere in the model, so nothing can learn that "the first option is usually the right one"
This isn't quite correct. There might not be an explicit "per-option" parameter, but the hidden states themselves "contain" all prior options. My gut feeling tells me they messed up data augmentation by incorrectly permuting options during training.
Edit: the above numbers were obtained with the v19 version/checkpoint, v21 is much less sensitive to ordering, but still shows a first position bias.
[0] https://github.com/strands-labs/strands-decider/blob/main/do...
On multiple tokens, we put each option (however many tokens it is) onto a line, and then read the hidden state at the end of that line to use at the option state. This is done option-by-option.
The hidden state at the end of the question is the query `q` (again, not sensitive to how many tokens the question is), each options end-of-line state is the key `k`, and the per-option logit is calculated as `q.k / sqrt(256)`.
{ "model": "strands-decider-2B-hobson-v19", "answers": { "is_urgent": { "type": "noul", "noul": 0.8287 } }, "usage": { "input_tokens": 86, "output_tokens": 1 }, "latency_ms": 1732.17 }
This is how I got it running - https://gist.github.com/2891eb0db9ea92c1a4e860d44f556292
There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.
I should let the development agent select a model and run it to improve token efficiency.
Is it being used this much these days?