RC RANDOM CHAOS

Strands built a 2B model that never writes text

Strands Decider 2B is an open-source 2B decision model that scores choices in ~115ms, cheap enough to run guardrail checks before every agent tool call.

· 4 min read
Strands built a 2B model that never writes text

Strands Decider 2B answers a yes/no or multiple-choice question in a median of about 115 milliseconds on an Nvidia RTX 3090, and around 153ms for small tasks on an M3 MacBook. It does that without producing a single token of text. The model cannot write a sentence, summarize a document, or hold a conversation. What it can do is pick one of the options you hand it and tell you how confident it is.

That constraint is the whole design. Strands labs built the model by taking a pretrained Qwen3.5-2B torso, removing the LM head that an ordinary LLM uses to generate text, and replacing it with a pointer head of just over a million parameters. The pointer head scores the hidden state at each candidate answer against the hidden state at an <answer> position, and the torso is fine-tuned with a rank-16 LoRA adapter. Every output comes out of a single parallel pass, so the model always returns one of the options you defined and never wanders into free text.

Strands calls this a decision model, or a “system one” model, part of a class that has drawn attention since TypeSafe AI shipped Jev earlier this month. Decision models pick between options (“is ‘turn on the lights’ about the coffee machine, yes or no”) and assign numerical scores (“is this sentiment positive, between 0 and 1”). Giving up arbitrary output buys lower latency, better accuracy at a given size, and one property that matters if you are wiring these into a control path: each decision carries a calibration score, an estimate of how likely that answer is to be correct. Frontier LLM inference APIs do not expose anything equivalent. Asking several questions about the same prompt is cheap too.

How it scores

On JevBench’s public set, strands-decider-2b lands 3rd of 33 models in the 2B class on combined accuracy and calibration, with calibration measured by Brier score, and 1st of 30 if you exclude the models that sit just over 2B. It gets 100% of the set’s easy tasks right. The team is open that this is version 19, the second major architecture after an earlier slot-head design that performed worse, and that the repository documents every change across versions.

A check in the before-tool hook

The example worth studying ships in the repo under examples/strands/. Strands agents expose an InterventionHandler with a before_tool_call method that runs before any tool executes. In the demo the agent has a get_weather tool and a deliberately eager system prompt, so when a user asks “what’s the weather?” with no location, the agent guesses a city and tries to call the tool. Before that call runs, the decider reads the conversation and the proposed arguments and answers two yes/no questions: are these argument values grounded in anything the user actually said, and is it too early to call this tool. A few lines of Python turn those predictions into a decision, and the agent goes back to ask which city you meant instead of confidently reporting the weather somewhere nobody mentioned.

The handler returns a typed action: Proceed, Deny, Confirm (stop and ask a person), or Guide (hand the model its turn back with feedback rather than blocking outright). Strands has no opinion about what goes inside the handler, so the same shape holds whether you call this decision model, a Cedar policy, or another agent. A check this cheap can sit in a path where an LLM call never could. A 115ms grounding test in front of every tool call is affordable; a second LLM round trip per tool call usually is not.

Where it fits, and where it doesn’t

Beyond guardrails, the early uses Strands lists are the rote parts of agent plumbing: model routing, tool selection, evaluations, memory, context management, and policy classification. The hybrid pattern is the one to watch for cost, letting an LLM make the hard calls while the decider handles the easy, repetitive ones.

Two cautions before you lean on it. The model is explicitly weak at complex problems and useless for anything that needs generated text, so treat it as a classifier in the loop, not a stand-in for your reasoning model. And the demo’s questions, threshold, and policy were all picked by hand as an illustration, not a recommendation; calibrating those for a real workload is the actual work. Strands says integration libraries are on the way, so for now you wire it up yourself.

The weights, the training data, and the training scripts are all published on GitHub and Hugging Face under an open-source release, which is more than most model drops include. If you want to find out whether a calibrated 2B classifier earns a place in your agent’s hot path, you can retrain it on hardware you already have.

Share

Keep Reading

Latest on the Wire

Full wire →

New signal daily · RSS

Stay in the loop

New writing delivered when it's ready. No schedule, no spam.