artificial intelligence

Strands Decider 2B: A Local Referee for AI Agents

Strands Decider 2B: A Local Referee for AI Agents

Picture an agent receiving a support message about a failed payout. Before it writes a reply, it may need to route the case, decide whether the request sounds urgent, and check whether a proposed tool call is grounded in anything the user actually said.

A full language model can handle those decisions, but it is a little like hiring a novelist to operate a three-position switch. Strands Decider 2B takes a narrower route: it reads a piece of state, chooses from defined options, and returns a confidence signal alongside the decision. The result is a small decision model designed for local inference and the fast, repetitive choices inside agentic workflows.

One version detail matters. The October 1, 2026 launch introduced the v19 model, while the repository identified v21 as its current reference checkpoint by October 7, 2026. The earlier v19 weights remain available, so older examples may name a different checkpoint even though the core design is the same. (strandsagents.com)

A model that does not try to chat

What is a decision model useful for if it cannot write a paragraph? It is useful precisely because many parts of an agent do not need a paragraph. They need a bounded answer that another piece of software can act on.

A large language model, or LLM, generates text one token at a time. A decision model instead scores the alternatives you provide. Strands Decider exposes three forms of question:

  • A noul question is yes or no and returns the probability of true.
  • A choice question selects one option from a set, such as billing, sales, or retail.
  • A score question places the input on an ordered scale, such as calm, frustrated, or very angry.

The option names do not have to be baked into the model during training. You describe them in the request, which means the same checkpoint can classify support tickets, select tools, or judge whether an answer meets a quality bar. All three question types use the same underlying option-scoring mechanism and read the result in different ways. (github.com)

The confidence value is the part that makes this more than ordinary classification. Calibration means that a model's confidence should line up with how often it is correct. The repository reports that, on short classification tasks it has not seen before, answers above 0.9 confidence are correct about 95 percent of the time. That is a useful starting point for routing policy, not a universal safety guarantee; production thresholds still need to be tested against your own data and consequences.

The trick is replacing generation with a pointer

The architecture begins with a pretrained language-model body, or torso. Strands uses Qwen3.5-2B-Base, a model with roughly 1.9 billion learned parameters. Normally, its language-model head predicts the next token so the model can continue a sentence.

Strands removes that generation head and adds a small pointer head instead. The torso reads the state, question, and options. A hidden state, which is the model's numerical representation of the text at a particular position, is taken from the answer position and compared with the hidden state at the end of each option. The pointer head turns those comparisons into scores.

That design produces one score per option in a single forward pass. There is no token-by-token decoding loop and no chance for the model to wander outside the allowed answer set. The pointer head contains only about a million parameters, while the torso is adapted with a rank-16 LoRA adapter. LoRA, short for Low-Rank Adaptation, is a small trainable add-on that changes how the larger model behaves without retraining every original weight.

state + question + options
 ↓
Qwen3.5-2B torso + rank-16 LoRA
 ↓
pointer head → option scores + confidence

Why the small size matters

The original v19 release reported 167 correct answers out of 231 tasks on the public JevBench set, with a median decision time of about 115 milliseconds on an Nvidia RTX 3090. Small tasks on an Apple-silicon Mac were reported at roughly 153 milliseconds. Latency increases as the input grows, but the model is still practical for local experiments on a GPU, a Mac, or a CPU.

The design also makes repeated questions about the same text cheaper. The shared state is encoded once, then the model evaluates the shorter question portions. Asking whether a message is urgent, which team should receive it, and whether a response is adequate does not require rereading the whole message three separate times. (strandsagents.com)

Those numbers need context. JevBench has only 231 public tasks, and the repository warns that small differences between single runs are difficult to interpret. That is one reason the project publishes its data, evaluation scripts, and training recipe rather than presenting a single leaderboard number as the whole story. The documented v19 recipe takes about 11 hours on one RTX 3090, which puts meaningful experimentation within reach of a well-equipped developer workstation. (github.com)

Run a decision locally

The command-line interface keeps the first experiment compact:

pip install strands-decider

strands-decider ask StrandsAgents/strands-decider-2B-hobson-v21 \\
 --state 'Help! My payouts have been failing for 3 days.' \\
 --choice 'Which team should handle this?=billing,sales,retail' \\
 --noul 'Does this convey urgency?' \\
 --score 'How frustrated is the writer?=calm,frustrated,depressed'

The state is the text being judged. Each flag adds a question about that same state. The command returns the selected choice, per-option probabilities, the yes/no probability, or the expected position on the ordered scale. The first run downloads the model components; later requests can use the local copy. The CLI supports CUDA, Apple silicon through MPS or MLX, and CPU execution, with device-specific setup documented by the project. (github.com)

Put the decider before the tool call

A useful pattern appears in the project's Strands agent example. An eager language model sees a weather request with no city and invents one. Before the weather tool executes, the decider checks two questions: are the tool arguments grounded in the user's words, and is it premature to call the tool before clarification?

The surrounding policy can stay ordinary Python:

if not grounded or premature:
 ask_for_clarification
else:
 call_tool

This is a hybrid agent. The language model handles interpretation and conversation; the decision model handles a repeatable gate. The same pattern fits model routing, argument validation, triage, guardrails, evaluation, and context management. A low-confidence result can trigger confirmation or human review instead of forcing the agent to pretend it knows more than it does. (strandsagents.com)

Know what it cannot do

Strands Decider 2B is not a smaller chatbot. It does not replace a reasoning model for coding, document summarization, open-ended conversation, or complex multi-step analysis. Its strength comes from making the answer space explicit, so poorly chosen options or vague question wording can still produce a misleading result.

The local HTTP server is intended for experimentation and binds to 127.0.0.1 without authentication by default. Keep that boundary in mind before placing it on a shared network. Used with clear options, measured thresholds, and a sensible fallback path, the model fills a useful gap between a heavyweight LLM and a hand-built classifier.

That is the interesting idea behind Strands Decider 2B: not a replacement for the agent's voice, but a fast, inspectable referee for the decisions surrounding it. Let the larger model handle ambiguity and language; let the decider handle bounded choices where latency, consistency, and calibrated confidence matter. (github.com)

ahsan

ahsan

Hello! I am Mr Ahsan, the writer of the Website. I am from Netherland. I like to write about technology and the news around it.

Comments (0)

No comments yet. Be the first to respond!

Leave a Comment

Your comment will be visible after review.