Strands Decider 2B: a small, open-source, decision model

274 points by gmays 20 hours ago on hackernews | 77 comments

miguelspizza | 19 hours ago

This is a great model. I've been running it on device in chrome extension to filter things like email.

It is just the right mix of size, capability and speed to make it generally useful for adhoc bulk classification tasks.

For those wanting to run it in browser: https://huggingface.co/alxnahas/strands-decider-2B-webgpu

babelfish | 19 hours ago

keyle | 19 hours ago

Fantastically well written. It's rare for me to be able to understand what the AI gurus are talking about, and this was written by humans for humans.

It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?

nryoo | 17 hours ago

Picking lunch menu..?

mattstir | 14 hours ago

A use case I have in mind is creating a relatively simple home assistant app that would let me control smart home things like lightbulbs and playing music on some kitchen speakers. I can also imagine having it differentiate between commands like "set lights to 30% brightness" and more general questions like "what's the weather today" and piping the latter over to a cheap LLM to handle. It could probably all be handled by an LLM, but I like how fast these decision models end up being.

tmoreton | 4 hours ago

Have you implemented this for yourself? In theory it sounds like a perfect use case but I actually did this with the Laya model but didn't have great results. Going to try again with this model now
One of the authors here. Thanks!

On humans-for-humans, blog posts with my name on aren't written by AI (https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html) either for my personal blog or at work. I'm a heavy AI user, but this is something I think is best left to humans.

There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence.

We're releasing more of that functionality in strands, and I expect an explosion of innovation in that area.

adenta | 19 hours ago

At this point I can't wait for a comedian to release a decision model backed by humans.

Meet Jerry- it's literally a guy named Jerry answering your questions.

melvinram | 18 hours ago

Your comment made me chuckle.

Jerry, "The Decider" https://www.youtube.com/watch?v=r8VbzrZ9yHQ

bfung | 17 hours ago

Same, but my mind went straight to Rick and Morty… def don’t want Jerry deciding XD

throwawayffffas | 13 hours ago

Ah, easy to work around, you just iterate eliminating his choices one at a time, until you get the one he didn't pick (that's your correct pick).

Balooga | 17 hours ago

Oh! You probably want ChatTJB [1]

[1] - https://chattjb.org/about

jeef_berky | 17 hours ago

Don't worry, Jerry will rig everything up for you.

jeremycarter | 16 hours ago

RACE HIM JERRY!

spwa4 | 8 hours ago

davvie | 17 hours ago

Looks really nice, I think I could use it on my Mac mini for some smaller automations

teruakohatu | 17 hours ago

Any idea how well this would run on a CPU?

gopalv | 17 hours ago

On my M3 mac, it works okay inside a docker container with just CPU.

{ "model": "strands-decider-2B-hobson-v19", "answers": { "is_urgent": { "type": "noul", "noul": 0.8287 } }, "usage": { "input_tokens": 86, "output_tokens": 1 }, "latency_ms": 1732.17 }

This is how I got it running - https://gist.github.com/2891eb0db9ea92c1a4e860d44f556292

There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.

avereveard | 17 hours ago

About half a second per decision on six cores

davidwritesbugs | 15 hours ago

Isnt that a bit slow for these?

soltanov | 17 hours ago

Benchmark calibration does not establish reliability on unfamiliar production inputs.

yieldcrv | 17 hours ago

a strand type game

SubiculumCode | 17 hours ago

Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....

necubi | 16 hours ago

Cloudflare’s clef is multimodal (https://blog.cloudflare.com/clef-decision-models/)

(Disclaimer, I work at Cloudflare, but not on models)

sauhsoj | 16 hours ago

Strands Decider can take vision in. How does it go with that question?

Zopieux | 14 hours ago

This is not advertised on their page, did you make this up?

I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking.

Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.

Zopieux | 9 hours ago

Thanks!
One of the authors here.

v21 supports vision: https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...

stephantul | 16 hours ago

2B being called small is such a sign of the times

mattvr | 16 hours ago

Why is everyone calling binary choices `noul`? Does this have some meaning or is it just copying Jev’s API?

haarts | 16 hours ago

It's from Bernoulli.

jtfrench | 16 hours ago

Interesting. Is that a unit he invented or is it just a reference to his last name that stuck?

mijoharas | 16 hours ago

From Bernoulli maps apparently. I still don't understand why[0].

[0] https://news.ycombinator.com/item?id=49723267 (see parent for reference)

henrymerrilees | 14 hours ago

A bit of a garden path path sentence... I believe "maps" in "maps to" was being used as a verb.

So "noul"--short for Bernoulli as in the Bernoulli distribution--"maps to if-statements."

mijoharas | 12 hours ago

oh god, you're completely right! nice and obvious on a reread. (I also just learnt about "garden path" sentences.)

I'd googled "bernoulli map" after reading that, and came across this[0], which I wasn't aware of and thought was somehow related so I entrenched my misunderstanding (I didn't dig in.)

Side note: I just wrote the sentence "you're completely right!", and almost changed it because it sounds like slop now. I wonder if we're gonna get an increase in these sentences in human written prose over time as people mimic these sentences, or a decrease as people shy away from them to not sound like AI. :)

[0] https://en.wikipedia.org/wiki/Dyadic_transformation

fennecfoxy | 13 hours ago

Huh. And here I was thinking the obvious thing is that it's a way to have "null" without it accidentally being parsed as null.

WASDx | 14 hours ago

The new OpenAI Decisions API calls it "predicate". Also calling the API "decisions" rather than "system one". Usually I don't like inventing new standards but I hope the OpenAI schema takes over. We don't need this hype terminology.

dprkh | 14 hours ago

Claude generated it and it stuck.

tchalla | 14 hours ago

The same reason why they’re calling this System One thinking. Everyone wants to be seen doing different things and smart ones.

hiimkeks | 14 hours ago

My guess is if you pronouce this "nool" it sounds similar to "bool", and it's a slice of the name "Bernoulli" because the models generate Bernoulli distributions

lr1970 | 5 hours ago

> Why is everyone calling binary choices `noul`?

It is from BerNOULli distribution [0]

[0] https://en.wikipedia.org/wiki/Bernoulli_distribution

EDIT: formatting

hrpnk | 16 hours ago

clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption.

Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:

    $ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose
    [...]
      File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name
    raise ValueError(f"Can not map tensor {name!r}")
    ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'
It's very doable, but we haven't done it yet (although some great community folks did an ONNX version of v21).

I haven't looked in depth, but it should be doable without modifications to llama.cpp. You can't just convert the LoRA, through - there's a whole pointer head and some custom layers that need to be correctly handled (in fact, the LoRA mostly exists to get the base model to behave the way the pointer head needs).

real_faxenoff | 15 hours ago

As a regular user of a bunch of specialized micromodels, I'll tell you this: you won't be happy with such a model (and its JEV counterparts) running permanently in the background on your PC's CPU. You need to offload their processing to the NPU. There are many pitfalls along the way, but the result is worth it.

NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).

I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.

cyanydeez | 11 hours ago

I haven't seen any frameworks for running the NPU. my 395+ needs a buddy.

AbsurdCensor | 10 hours ago

I think under Lemonade you can do that, especially with AMD systems, using the models that allow for hybrid operation. Prefill happens on the NPU and token generation happens on the GPU. Not sure if this is faster or better than just doing it all on the GPU side.

shifto | 10 hours ago

It's slower but more power efficient. I think only some form of onyx models are available and I couldnt get anything to run on the first gen NPU's in the 7940HS cpu.

Infernal | 6 hours ago

FastFlowLM is what you're looking for - Lemonade does package it, as sibling comment says.

https://fastflowlm.com/

I wish they supported older gen XDNA though :(

ricardobeat | 11 hours ago

Most LLMs cannot run efficiently on current NPUs (except for prefill stage), the hardware was built for a different kind of ML workload.

mermerico | 9 hours ago

Decision models are prefill only

spwa4 | 8 hours ago

That is most of the explanation for their speed, even.
What models are you running on CPU? Any repos or gists you can share? Curious about the models and your use case. Are you doing multi-language or single language?

Btw, would love your opinion on this: pre trained classifiers that run and train on CPU https://github.com/nicobrenner/jeffy

woadwarrior01 | 15 hours ago

The ~2-week-old Intern-Decision family of models (0.8B, 2B and 4B) have the same Qwen3.5 base model family (albeit the instruction-tuned variants) and pointer head architecture.

https://huggingface.co/collections/internlm/intern-decision

mynti | 14 hours ago

Can someone explain this architecture a bit more in depth? They say the pointer head scores the hidden state at each option against the hidden state of the answer. But the LLM produces hidden states per token, so an option can span multiple tokens, no?

fxwin | 10 hours ago

as far as gpt-5.6 could help me understand the code in their repository (and as far as their architecture is accurate [0]), the "hidden state at each option" refers to the hidden state at the last token in each option, which is then scored by the (learned) pointer head against the hidden state at the <answer>-position.

To vaguely confirm that this is plausible, i tested the sample query from the post with differently arranged options (This should produce slightly different outputs since each option influences all subsequent hidden states, including those of the other options):

"billing,sales,retail" produces billing -> 0.843, retail -> 0.092, sales -> 0.065

"retail,billing,sales" produces billing -> 0.470, retail -> 0.468, sales -> 0.062

"retail,sales,billing" produces billing -> 0.517, retail -> 0.415, sales -> 0.068

"billing,retail,sales" produces billing -> 0.803, retail -> 0.146, sales -> 0.051

"sales,retail,billing" produces billing -> 0.647, retail -> 0.127, sales -> 0.225

So each of these arrangements favor billing, but with pretty different confidence scores, and it looks like there is a pretty big bias toward the first option listed - which makes me wonder how viable it would be to have the torso process these options independently instead

Their docs say:

> An option's score depends on what it says, not where it sits. There is no per-option parameter anywhere in the model, so nothing can learn that "the first option is usually the right one"

This isn't quite correct. There might not be an explicit "per-option" parameter, but the hidden states themselves "contain" all prior options. My gut feeling tells me they messed up data augmentation by incorrectly permuting options during training.

Edit: the above numbers were obtained with the v19 version/checkpoint, v21 is much less sensitive to ordering, but still shows a first position bias.

[0] https://github.com/strands-labs/strands-decider/blob/main/do...

I wrote a slightly deeper blog post here: https://brooker.co.za/blog/2026/09/28/engineering-system-one...

On multiple tokens, we put each option (however many tokens it is) onto a line, and then read the hidden state at the end of that line to use at the option state. This is done option-by-option.

The hidden state at the end of the question is the query `q` (again, not sensitive to how many tokens the question is), each options end-of-line state is the key `k`, and the per-option logit is calculated as `q.k / sqrt(256)`.

fxwin | 9 hours ago

Appreciate the transparent process! Is there a reason why you chose to process the entire list of options as one input to the LLM torso instead of splitting them and processing options separately? From what i can tell from my (limited) experiments, there is a lot of variation caused by simply reordering options, plus a heavy bias toward the first option listed [0], and my guess is that this is caused by hidden state of each option "leaking" into subsequent options

[0] https://news.ycombinator.com/item?id=49987076

No particularly principled reason, no. It's one of the (many) design variants we haven't had time to experiment with yet.

girvo | 14 hours ago

Does anyone know if it is worth fine-tuning one of these decision models on the shape of the questions you want it to work on, vs the more general versions? I'm using Jev pretty successfully at work at the moment, but am curious about what is doable

weinzierl | 13 hours ago

I'm interested in this as well and maybe to broaden the scope of the question a little:

If I have a sizeable amount of labeled data and need decisions calibrated to that data should I

1. Ignore the hype and train a traditional classifier

2. Finetune an LLM based decision model

3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow

If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?

All are viable options.

(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.

(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.

(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.

The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.

gxcsoccer | 13 hours ago

I think jev points to an interesting way to fine-tune open models

The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results

One of the authors here.

Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.

As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).

Everything you need to fine-tune, or even retrain from scratch, is in the github repo.

Ujj-001 | 13 hours ago

is jev commoditized now ?

kimseungyong | 13 hours ago

I wait for this open-source model.

I should let the development agent select a model and run it to improve token efficiency.

Is it being used this much these days?

ContinuityLab | 12 hours ago

Swapping out the text-generation head for a dedicated pointer head on a small footprint model is a pragmatic approach for low-latency local decision pipelines.

ricardobeat | 11 hours ago

I wish these would stop using JevBench. It focuses way too much on text classification tasks, and some of the models perform very poorly on tasks that need actual intelligence.
One of the authors here.

I appreciate what the JevBench folks are doing, but it's true the JevBench (like all benchmarks) isn't representative of every real-world workloads. I particularly don't love how JevBench handles confidence, and rewards being highly confident in certain cases. Benchmarking is hard, and JevBench is probably more representation of its target workloads than TPC-C is :)

As for actual intelligence, there's only so much you can expect from 2B without reasoning. One of the design challenges in training this model was to avoid forgetting too much, and the KL-to-frozen-base step partially exists for that purpose. Even then, Qwen3.5 2B base has fairly limited single-pass reasoning ability (which shows up for us in the performance on the JevBench hard set).

dev_l1x_be | 9 hours ago

General question: what is this model good for? I have mixed results with Laya on a pdf classifier (Jev was doing much better).

mijoharas | 8 hours ago

so, how are people running these locally? is there a llama.cpp/ollama solution for systemone/jev style apis?

gsnedders | 6 hours ago

What I don’t get from this article is why strands-decider-2b v10 and strands-decider-2b v11 are shown as “other systems”, and with v11 notably outperforming v19 — why are these other systems, and why is v11 not the path taken?
Oh, huh, that's an error in the image I didn't notice! Those comparisons are to a model called 'decider-2B', which isn't ours.

These benchmarks cover strands-decider v19. v21 is better calibrated, and we've got some new ones coming this week that move us further towards the bottom right.

muslimpeacepriz | 5 hours ago

Gone are the days when HN was flooded by "-lang.org"-postw; nowadays it is " model". Equally full of BS.
The kind of stuff that people are doing with decision models today is largely the way I've been using Liquid AI's LFM models, as rapid iteration interrogation tools for atomic decisions. They aren't good at reasoning, but you can reason about what decisions you would need in order to come to a reasonably high quality decision about a thing.

Even though they weren't themselves decision models exactly, they were so fast that you could use them in a similar way. They obliterated Qwen and other models for a lot of tasks in that use case.

I'm not sure I would agree with some of the claims made in this article though. You can definitely do a lot of similar tasks to regular LLMs with decision models, you just have to chain the results. You can even technically infuse a broader perspective into each token choice or force certain context to have priority in the decision making of the next token.

AdamIdrissi | 4 hours ago

Wow this is rlly cool.