Rendered at 17:06:33 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
real_faxenoff 10 hours ago [-]
As a regular user of a bunch of specialized micromodels, I'll tell you this: you won't be happy with such a model (and its JEV counterparts) running permanently in the background on your PC's CPU. You need to offload their processing to the NPU. There are many pitfalls along the way, but the result is worth it.
NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).
I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.
ricardobeat 6 hours ago [-]
Most LLMs cannot run efficiently on current NPUs (except for prefill stage), the hardware was built for a different kind of ML workload.
mermerico 4 hours ago [-]
Decision models are prefill only
spwa4 3 hours ago [-]
That is most of the explanation for their speed, even.
nico 5 hours ago [-]
What models are you running on CPU? Any repos or gists you can share? Curious about the models and your use case. Are you doing multi-language or single language?
I think under Lemonade you can do that, especially with AMD systems, using the models that allow for hybrid operation. Prefill happens on the NPU and token generation happens on the GPU. Not sure if this is faster or better than just doing it all on the GPU side.
shifto 5 hours ago [-]
It's slower but more power efficient. I think only some form of onyx models are available and I couldnt get anything to run on the first gen NPU's in the 7940HS cpu.
mrkn1 5 hours ago [-]
[dead]
adenta 14 hours ago [-]
At this point I can't wait for a comedian to release a decision model backed by humans.
Meet Jerry- it's literally a guy named Jerry answering your questions.
What I don’t get from this article is why strands-decider-2b v10 and strands-decider-2b v11 are shown as “other systems”, and with v11 notably outperforming v19 — why are these other systems, and why is v11 not the path taken?
mjb 43 minutes ago [-]
Oh, huh, that's an error in the image I didn't notice! Those comparisons are to a model called 'decider-2B', which isn't ours.
These benchmarks cover strands-decider v19. v21 is better calibrated, and we've got some new ones coming this week that move us further towards the bottom right.
mattvr 11 hours ago [-]
Why is everyone calling binary choices `noul`? Does this have some meaning or is it just copying Jev’s API?
The new OpenAI Decisions API calls it "predicate". Also calling the API "decisions" rather than "system one". Usually I don't like inventing new standards but I hope the OpenAI schema takes over. We don't need this hype terminology.
hiimkeks 9 hours ago [-]
My guess is if you pronouce this "nool" it sounds similar to "bool", and it's a slice of the name "Bernoulli" because the models generate Bernoulli distributions
haarts 11 hours ago [-]
It's from Bernoulli.
jtfrench 11 hours ago [-]
Interesting. Is that a unit he invented or is it just a reference to his last name that stuck?
mijoharas 11 hours ago [-]
From Bernoulli maps apparently. I still don't understand why[0].
A bit of a garden path path sentence... I believe "maps" in "maps to" was being used as a verb.
So "noul"--short for Bernoulli as in the Bernoulli distribution--"maps to if-statements."
mijoharas 7 hours ago [-]
oh god, you're completely right! nice and obvious on a reread. (I also just learnt about "garden path" sentences.)
I'd googled "bernoulli map" after reading that, and came across this[0], which I wasn't aware of and thought was somehow related so I entrenched my misunderstanding (I didn't dig in.)
Side note: I just wrote the sentence "you're completely right!", and almost changed it because it sounds like slop now. I wonder if we're gonna get an increase in these sentences in human written prose over time as people mimic these sentences, or a decrease as people shy away from them to not sound like AI. :)
Huh. And here I was thinking the obvious thing is that it's a way to have "null" without it accidentally being parsed as null.
tchalla 9 hours ago [-]
The same reason why they’re calling this System One thinking. Everyone wants to be seen doing different things and smart ones.
dprkh 9 hours ago [-]
Claude generated it and it stuck.
keyle 14 hours ago [-]
Fantastically well written. It's rare for me to be able to understand what the AI gurus are talking about, and this was written by humans for humans.
It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?
mjb 5 hours ago [-]
One of the authors here. Thanks!
On humans-for-humans, blog posts with my name on aren't written by AI (https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html) either for my personal blog or at work. I'm a heavy AI user, but this is something I think is best left to humans.
There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence.
We're releasing more of that functionality in strands, and I expect an explosion of innovation in that area.
mattstir 9 hours ago [-]
A use case I have in mind is creating a relatively simple home assistant app that would let me control smart home things like lightbulbs and playing music on some kitchen speakers. I can also imagine having it differentiate between commands like "set lights to 30% brightness" and more general questions like "what's the weather today" and piping the latter over to a cheap LLM to handle. It could probably all be handled by an LLM, but I like how fast these decision models end up being.
nryoo 12 hours ago [-]
Picking lunch menu..?
olgava 10 hours ago [-]
[flagged]
woadwarrior01 10 hours ago [-]
The ~2-week-old Intern-Decision family of models (0.8B, 2B and 4B) have the same Qwen3.5 base model family (albeit the instruction-tuned variants) and pointer head architecture.
clef from cloudflare runs on llama.cpp - being locked-in to strands cli would be a bummer and will slow down adoption.
Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:
$ uv run --with transformers==5.19.0 convert_lora_to_gguf.py ~/Downloads/lora --dry-run --verbose
[...]
File "/Users/user/repos/llama.cpp/conversion/base.py", line 630, in map_tensor_name
raise ValueError(f"Can not map tensor {name!r}")
ValueError: Can not map tensor 'layers.0.linear_attn.in_proj_a.weight'
mjb 5 hours ago [-]
It's very doable, but we haven't done it yet (although some great community folks did an ONNX version of v21).
I haven't looked in depth, but it should be doable without modifications to llama.cpp. You can't just convert the LoRA, through - there's a whole pointer head and some custom layers that need to be correctly handled (in fact, the LoRA mostly exists to get the base model to behave the way the pointer head needs).
dev_l1x_be 4 hours ago [-]
General question: what is this model good for? I have mixed results with Laya on a pdf classifier (Jev was doing much better).
ricardobeat 6 hours ago [-]
I wish these would stop using JevBench. It focuses way too much on text classification tasks, and some of the models perform very poorly on tasks that need actual intelligence.
mjb 4 hours ago [-]
One of the authors here.
I appreciate what the JevBench folks are doing, but it's true the JevBench (like all benchmarks) isn't representative of every real-world workloads. I particularly don't love how JevBench handles confidence, and rewards being highly confident in certain cases. Benchmarking is hard, and JevBench is probably more representation of its target workloads than TPC-C is :)
As for actual intelligence, there's only so much you can expect from 2B without reasoning. One of the design challenges in training this model was to avoid forgetting too much, and the KL-to-frozen-base step partially exists for that purpose. Even then, Qwen3.5 2B base has fairly limited single-pass reasoning ability (which shows up for us in the performance on the JevBench hard set).
girvo 9 hours ago [-]
Does anyone know if it is worth fine-tuning one of these decision models on the shape of the questions you want it to work on, vs the more general versions? I'm using Jev pretty successfully at work at the moment, but am curious about what is doable
mjb 5 hours ago [-]
One of the authors here.
Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.
As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).
Everything you need to fine-tune, or even retrain from scratch, is in the github repo.
weinzierl 8 hours ago [-]
I'm interested in this as well and maybe to broaden the scope of the question a little:
If I have a sizeable amount of labeled data and need decisions calibrated to that data should I
1. Ignore the hype and train a traditional classifier
2. Finetune an LLM based decision model
3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow
If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?
mjb 5 hours ago [-]
All are viable options.
(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.
(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.
(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.
The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.
mrkn1 6 hours ago [-]
[dead]
gxcsoccer 8 hours ago [-]
I think jev points to an interesting way to fine-tune open models
The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results
muslimpeacepriz 14 minutes ago [-]
Gone are the days when HN was flooded by "-lang.org"-postw; nowadays it is " model".
Equally full of BS.
SubiculumCode 12 hours ago [-]
Are any of these multimodal yet? I'd love to try asking a model with calibrated probabilities to answer question like, "do these shapes match?". Sure, you can ask a LLM....
Can someone explain this architecture a bit more in depth? They say the pointer head scores the hidden state at each option against the hidden state of the answer. But the LLM produces hidden states per token, so an option can span multiple tokens, no?
On multiple tokens, we put each option (however many tokens it is) onto a line, and then read the hidden state at the end of that line to use at the option state. This is done option-by-option.
The hidden state at the end of the question is the query `q` (again, not sensitive to how many tokens the question is), each options end-of-line state is the key `k`, and the per-option logit is calculated as `q.k / sqrt(256)`.
fxwin 4 hours ago [-]
Appreciate the transparent process! Is there a reason why you chose to process the entire list of options as one input to the LLM torso instead of splitting them and processing options separately? From what i can tell from my (limited) experiments, there is a lot of variation caused by simply reordering options, plus a heavy bias toward the first option listed [0], and my guess is that this is caused by hidden state of each option "leaking" into subsequent options
No particularly principled reason, no. It's one of the (many) design variants we haven't had time to experiment with yet.
fxwin 5 hours ago [-]
as far as gpt-5.6 could help me understand the code in their repository (and as far as their architecture is accurate [0]), the "hidden state at each option" refers to the hidden state at the last token in each option, which is then scored by the (learned) pointer head against the hidden state at the <answer>-position.
To vaguely confirm that this is plausible, i tested the sample query from the post with differently arranged options (This should produce slightly different outputs since each option influences all subsequent hidden states, including those of the other options):
So each of these arrangements favor billing, but with pretty different confidence scores, and it looks like there is a pretty big bias toward the first option listed - which makes me wonder how viable it would be to have the torso process these options independently instead
Their docs say:
> An option's score depends on what it says, not where it sits. There is no per-option parameter anywhere in the model, so nothing can learn that "the first option is usually the right one"
This isn't quite correct. There might not be an explicit "per-option" parameter, but the hidden states themselves "contain" all prior options. My gut feeling tells me they messed up data augmentation by incorrectly permuting options during training.
Edit: the above numbers were obtained with the v19 version/checkpoint, v21 is much less sensitive to ordering, but still shows a first position bias.
There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.
avereveard 12 hours ago [-]
About half a second per decision on six cores
davidwritesbugs 10 hours ago [-]
Isnt that a bit slow for these?
mrkn1 6 hours ago [-]
[flagged]
mijoharas 3 hours ago [-]
so, how are people running these locally? is there a llama.cpp/ollama solution for systemone/jev style apis?
soltanov 12 hours ago [-]
Benchmark calibration does not establish reliability on unfamiliar production inputs.
kimseungyong 8 hours ago [-]
I wait for this open-source model.
I should let the development agent select a model and run it to improve token efficiency.
Is it being used this much these days?
4 hours ago [-]
ContinuityLab 7 hours ago [-]
Swapping out the text-generation head for a dedicated pointer head on a small footprint model is a pragmatic approach for low-latency local decision pipelines.
stephantul 11 hours ago [-]
2B being called small is such a sign of the times
Ujj-001 8 hours ago [-]
is jev commoditized now ?
chelseahermes 3 hours ago [-]
[flagged]
cosprax 1 hours ago [-]
[flagged]
lin7c 11 hours ago [-]
[flagged]
happybox2016 8 hours ago [-]
[flagged]
Fluid_Mechanics 14 hours ago [-]
[flagged]
yieldcrv 12 hours ago [-]
a strand type game
davvie 12 hours ago [-]
Looks really nice, I think I could use it on my Mac mini for some smaller automations
NPU performance will be twice as high, while power consumption will be four times lower. No additional fan noise (if you know what I mean).
I'll wait another month until the first phase of the =battle royale= among models of this kind wraps up, put together a solution for the NPU/iGPU, and post it on HF.
Btw, would love your opinion on this: pre trained classifiers that run and train on CPU https://github.com/nicobrenner/jeffy
https://fastflowlm.com/
Meet Jerry- it's literally a guy named Jerry answering your questions.
[1] - https://chattjb.org/about
Jerry, "The Decider" https://www.youtube.com/watch?v=r8VbzrZ9yHQ
These benchmarks cover strands-decider v19. v21 is better calibrated, and we've got some new ones coming this week that move us further towards the bottom right.
It is from BerNOULli distribution [0]
[0] https://en.wikipedia.org/wiki/Bernoulli_distribution
EDIT: formatting
[0] https://news.ycombinator.com/item?id=49723267 (see parent for reference)
So "noul"--short for Bernoulli as in the Bernoulli distribution--"maps to if-statements."
I'd googled "bernoulli map" after reading that, and came across this[0], which I wasn't aware of and thought was somehow related so I entrenched my misunderstanding (I didn't dig in.)
Side note: I just wrote the sentence "you're completely right!", and almost changed it because it sounds like slop now. I wonder if we're gonna get an increase in these sentences in human written prose over time as people mimic these sentences, or a decrease as people shy away from them to not sound like AI. :)
[0] https://en.wikipedia.org/wiki/Dyadic_transformation
It can technically be used for a lot of use cases, I'd like people to chime in on ideas on this?
On humans-for-humans, blog posts with my name on aren't written by AI (https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html) either for my personal blog or at work. I'm a heavy AI user, but this is something I think is best left to humans.
There are a ton a ways to use models like this. Model routing is a popular emerging one. The one that really interests me is hybrid agentic workflows - filling the gap between deterministic workflows (e.g. AWS StepFunctions) and fully agentic workflows that use frontier intelligence.
We're releasing more of that functionality in strands, and I expect an explosion of innovation in that area.
https://huggingface.co/collections/internlm/intern-decision
It is just the right mix of size, capability and speed to make it generally useful for adhoc bulk classification tasks.
For those wanting to run it in browser: https://huggingface.co/alxnahas/strands-decider-2B-webgpu
Since it's a LoRa on Qwen, I assume this is runnable via llama.cpp. Pity that the PEFT/LoRa->GGUF translation is left to the user. Anyone got past:
I haven't looked in depth, but it should be doable without modifications to llama.cpp. You can't just convert the LoRA, through - there's a whole pointer head and some custom layers that need to be correctly handled (in fact, the LoRA mostly exists to get the base model to behave the way the pointer head needs).
I appreciate what the JevBench folks are doing, but it's true the JevBench (like all benchmarks) isn't representative of every real-world workloads. I particularly don't love how JevBench handles confidence, and rewards being highly confident in certain cases. Benchmarking is hard, and JevBench is probably more representation of its target workloads than TPC-C is :)
As for actual intelligence, there's only so much you can expect from 2B without reasoning. One of the design challenges in training this model was to avoid forgetting too much, and the KL-to-frozen-base step partially exists for that purpose. Even then, Qwen3.5 2B base has fairly limited single-pass reasoning ability (which shows up for us in the performance on the JevBench hard set).
Yes, at this size (2B) you can get better performance and calibration fine tuning on your questions. You can also improve calibration on your problem type just be recalibrating without fine-tuning.
As to what's doable, that depends on what you expect. It's a single pass through a 2B model, no reasoning, so it's never going to be particularly 'intelligent'. On domain problems, though, you can get great in-distribution performance (and possibly better in-distribution performance than you'd get using the same volume of data to train a specialized classifier).
Everything you need to fine-tune, or even retrain from scratch, is in the github repo.
If I have a sizeable amount of labeled data and need decisions calibrated to that data should I
1. Ignore the hype and train a traditional classifier
2. Finetune an LLM based decision model
3. Shoehorn (probably a small subset of) the data into the context of the LLM classifier somehow
If the answer is 3. where does the data belong? In the input content? Request wide state? In the question instructions? In the criteria? How much of my data can and should I use?
(1) would take the most data (probably, depends on your domain) and probably generalize worst, but would be closest to compute-optimal in the end.
(2) doesn't take much data, you can tweak calibration to your needs, and might only take a few minutes on a beefy GPU.
(3) is the way to go if you need a lot of general knowledge, any amount of reasoning, and don't need good calibration. Likely the least inference efficient of the three options for a given accuracy and calibration.
The hard part of (3) is keeping the answers to the questions independent. If you dump them all together with the state into the prompt, the answers to the second question will depend on the first (and the answer to the first if you it step-by-step). You can do it by tweaking the inference process with the right masking or use of batching, or by fiddling with cache control using a provider's API.
The goal isn’t to replace current models, it’s to train a model for a specific domain so it can handle multiple-choice and yes/no questions quickly, helping the overall system run faster and get better results
v21 supports vision: https://huggingface.co/StrandsAgents/strands-decider-2B-hobs...
(Disclaimer, I work at Cloudflare, but not on models)
I believe image classification/analysis by deciders (not just OCR, not everything is about text) is still lacking.
Cloudflare's Clef had fair results on my test, but it's larger and slower. Wondering about Strands.
On multiple tokens, we put each option (however many tokens it is) onto a line, and then read the hidden state at the end of that line to use at the option state. This is done option-by-option.
The hidden state at the end of the question is the query `q` (again, not sensitive to how many tokens the question is), each options end-of-line state is the key `k`, and the per-option logit is calculated as `q.k / sqrt(256)`.
[0] https://news.ycombinator.com/item?id=49987076
To vaguely confirm that this is plausible, i tested the sample query from the post with differently arranged options (This should produce slightly different outputs since each option influences all subsequent hidden states, including those of the other options):
"billing,sales,retail" produces billing -> 0.843, retail -> 0.092, sales -> 0.065
"retail,billing,sales" produces billing -> 0.470, retail -> 0.468, sales -> 0.062
"retail,sales,billing" produces billing -> 0.517, retail -> 0.415, sales -> 0.068
"billing,retail,sales" produces billing -> 0.803, retail -> 0.146, sales -> 0.051
"sales,retail,billing" produces billing -> 0.647, retail -> 0.127, sales -> 0.225
So each of these arrangements favor billing, but with pretty different confidence scores, and it looks like there is a pretty big bias toward the first option listed - which makes me wonder how viable it would be to have the torso process these options independently instead
Their docs say:
> An option's score depends on what it says, not where it sits. There is no per-option parameter anywhere in the model, so nothing can learn that "the first option is usually the right one"
This isn't quite correct. There might not be an explicit "per-option" parameter, but the hidden states themselves "contain" all prior options. My gut feeling tells me they messed up data augmentation by incorrectly permuting options during training.
Edit: the above numbers were obtained with the v19 version/checkpoint, v21 is much less sensitive to ordering, but still shows a first position bias.
[0] https://github.com/strands-labs/strands-decider/blob/main/do...
{ "model": "strands-decider-2B-hobson-v19", "answers": { "is_urgent": { "type": "noul", "noul": 0.8287 } }, "usage": { "input_tokens": 86, "output_tokens": 1 }, "latency_ms": 1732.17 }
This is how I got it running - https://gist.github.com/2891eb0db9ea92c1a4e860d44f556292
There's a lot more to be done if we optimize for MLX & let it run on a Mac mini instead of the docker wrapper.
I should let the development agent select a model and run it to improve token efficiency.
Is it being used this much these days?