The team behind the Strands Agents toolkit published a two-billion-parameter model on 1 October that cannot write a single sentence. Strands Decider 2B takes a question plus a short list of possible answers and returns a probability for each one. That makes it a decision model, and the design is built for speed rather than eloquence.
The authors’ own example feeds in a customer message saying payouts have been failing for three days and asks which team should handle it: billing, sales, or retail. The model returns billing at 0.845, retail at 0.091, and sales at 0.064. It never explains itself. Your own code reads those numbers and decides whether to route the ticket, ask a person, or do nothing.
The authors, Marc Brooker, Mike Chambers, and Fabio Nonato de Paula, writing on the Strands Agents blog, say that trade gives them three things. The model always lands on one of the offered options, it attaches a reliability score to every answer, and it can answer many questions about the same input cheaply. They add that frontier LLM inference APIs do not expose a comparable reliability score. The cost is capability: by the team’s own account, a model like this does far worse than a reasoning model on hard problems, and it cannot do coding, chat, or summarization at all.
Under the hood, the team started from Qwen3.5-2B, an existing open language model, and removed the part that produces text. In its place sits a pointer head, a small add-on of just over a million parameters that looks at each candidate answer and scores how well it fits the question. Training adjusted the rest of the model through a LoRA adapter at rank 16, a lightweight method that changes only a small slice of the weights. The released version is the nineteenth iteration, and the post says an earlier design with a different scoring head performed significantly worse.
The performance numbers all come from the authors’ own testing, not an outside lab. Accuracy was measured on the public set of JevBench, a test named for TypeSafe AI’s Jev model, which launched earlier this month and set off the current interest in this model class. Calibration, meaning whether a stated confidence matches how often the answer is right, was scored with the Brier score, again on that public set. The post reports a third-place finish among 33 models in the 2B size class, and first place among 30 once models slightly over 2B are excluded. It also reports a perfect score on the easy JevBench tasks, which says little about hard ones.
Speed is the pitch. The authors report a median of about 115 milliseconds per decision on a local Nvidia RTX 3090 graphics card, and about 153 milliseconds for small tasks on an M3 MacBook. One caveat sits in the figure caption: the latency chart was measured against version 18, not the released version 19.
The most concrete demo shows why a fast scorer matters. A sample agent is deliberately eager, so when a user asks “What’s the weather?” without naming a place, it guesses a city and calls its weather tool. Before the call runs, the decision model answers two yes/no questions about it: did the user actually supply these argument values, and should the call wait? A short Python snippet converts the answers into an action, and the agent asks which city was meant. The authors call this a demo, not advice, and say they picked the questions, the threshold, and the policy by hand.
The team also describes early uses it is seeing for this class of model: choosing which model should handle a request, selecting tools, running guardrails, and handling memory. It also mentions hybrid agents, where a large model handles the hardest calls and a decision model takes the routine ones to cut cost and delay. Those are the authors’ observations, not measured results.
The release is unusually open. Code is on GitHub, weights are on Hugging Face, and the training data and scripts are included, so anyone can inspect or retrain the model on hardware they already own.
The market angle is cost. A safety check that needs a second call to a hosted language model is a check many teams skip, while one that finishes in roughly a tenth of a second on local hardware can sit in front of every tool call. Teams weighing that idea should build a small test set from their own tickets and tool calls first, because a public ranking says little about how the scores hold up on their own data.
Reported by the Strands Agents blog on 1 October 2026.