Agent·

AWS Strands Decider 2B — Pointer-head decision models — What does it mean?

The news. On October 1, 2026, AWS's Strands team released Strands Decider 2B, a decision model built on a Qwen3.5-2B torso, with weights, training data and training scripts published. AWS reports a median decision latency of about 115 ms on an Nvidia RTX 3090 and about 153 ms on an M3 MacBook for small tasks (its latency graph was measured on v18), and says the released model is iteration v19; an earlier version used a "slot head" that performed significantly worse. On the public set of JevBench, a benchmark for this model class, it ranks 3rd of 33 models in the 2B class on accuracy and Brier-score calibration. Read the release →

Picture two ways to collect a vote. A blank answer sheet lets the voter write anything — a name, a misspelling, an essay. A printed ballot lists the candidates and offers only tick boxes. Strands Decider turns a language model from a blank answer sheet into a printed ballot: it can only mark one of the boxes you printed.

A normal LLM answers with its LM head. After the transformer layers run, the head scores every token in the vocabulary, the model samples one, appends it and runs again (one token at a time, from logits). Ask it "which team should handle this: billing, sales or retail?" and the reply is whatever tokens it writes — "billing", "Billing team", or a sentence of reasoning. Structured outputs can constrain which tokens are allowed, but the model is still writing, one step at a time.

Strands Decider removes that head. The caller sends a state, a question and the options, and the prompt includes an <answer> position. The pointer head scores the hidden state at each option's position against the hidden state at the answer position, so the output is a position in the input, not a word from the vocabulary. The name echoes pointer networks, where a model answers by pointing at part of its input. On the ballot, the voter's pen can only land in a box that was printed — there is no write-in line, so an out-of-list answer is impossible by construction.

Because every option is scored in the same pass, the model returns a score for every option, not just the winner. In AWS's own routing example the three scores are billing 0.845, retail 0.091 and sales 0.064, and the CLI also reports a separate confidence of 0.768 for the choice; the post does not document how that confidence is computed. AWS argues this reliability signal is something frontier LLM inference APIs do not expose, and that is what makes thresholds possible: route automatically above a cut-off, hand off to a person or a bigger model below it.

The tradeoff is stated plainly in the release. Answering in one parallel pass makes a decision model much worse than a reasoning model at hard problems, and it cannot write text at all — no code, no chat, no summaries. AWS's suggested pattern is a hybrid: let an LLM make the hard decisions and let the decider make the easy, repetitive ones in the routing step or in a policy check before a tool runs.

PropertyGenerative LLM (LM head)Decision model (pointer head)
OutputFree text, one token per stepOne of the supplied options, in one pass
Answer outside the list?Possible unless decoding is constrainedImpossible by construction
Per-option scoreNot exposed by frontier inference APIs, per AWSReturned for every option (source)
Hard, multi-step problemsStrong, especially reasoning modelsSignificantly worse, per AWS
Latency (Strands Decider 2B)—~115 ms median, RTX 3090 (v18 graph); ~153 ms for small tasks, M3 MacBook (source)

Where the change pays off — a worked example. Hold three things fixed: one support message ("my payouts have been failing for 3 days"), three options (billing, sales, retail) and the same 2B torso. With an LM head, the output layer is a matrix of hidden size × vocabulary — for a hidden size of 2,048 and a 150,000-token vocabulary that is about 307 million parameters (illustrative), and it ranks all 150,000 tokens at every step. With the pointer head, the output layer is about 1 million parameters — roughly 0.05% of the 2B model — and it ranks exactly 3 positions, once. The three scores AWS shows add up to 0.845 + 0.091 + 0.064 = 1.000, so they behave like a probability distribution over the options alone, and the gap between first and second place (0.845 vs 0.091) is one number a router could threshold on. The torso itself is adapted with a rank-16 LoRA, so most of the 2B weights stay frozen during training.

The same cheapness changes where a check can go. AWS's demo runs the decider before every tool call and asks it two yes/no questions about the proposed call: are the argument values grounded in something the user said, and is it too early to call the tool at all. When a deliberately eager agent guesses a city for a weather lookup, the decider answers "not grounded" and the agent goes back to ask which city was meant. The point is not that a 2B model is smart; it is that a decision costing about a tenth of a second can sit on every step of the agent's cost profile, where an extra LLM call would not.

Goes deeper in: AI Agents → Workflow Patterns → Chaining + Routing

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based