AWS Strands Decider 2B Open-Sourced: 2B-Param Decision Model for Agent Routing, Tool Gating at 115ms
AWS Strands Labs released Strands Decider 2B, a 2-billion-parameter open-source decision model that replaces the LM head of Qwen3.5-2B with a pointer scoring head. It achieves 72.3% accuracy on JevBench public set with 115ms median latency on a local RTX 3090. Instead of generating freeform text, it selects between predefined options and outputs calibrated confidence scores — ideal for agent routing, tool gating, and argument grounding checks that don't need a full LLM call.
Public-source compilation
Synthesized from public posts/docs. Prefer the original source for primary claims.
How we review content · Primary source · Strands Agents Blog
TL;DR
AWS Strands Labs released Strands Decider 2B on October 1, 2026 — a 2-billion-parameter open-source decision model built on the Qwen3.5-2B torso with the LM head removed and replaced by a ~1M-parameter pointer scoring head. It achieves 72.3% accuracy (167/231) on the JevBench public set with a median latency of ~115ms on a local RTX 3090. Fully open-sourced under Apache-2.0 with training data and scripts, installable via pip install strands-decider.
This marks the rapid mainstreaming of "decision models" (also called System One models) — a category pioneered by TypeSafe's Jev just two weeks earlier. Decision models don't generate freeform text; they select between predefined options and output calibrated confidence scores, making them ideal for high-frequency, low-complexity decisions in AI agent workflows: model routing, tool selection, argument grounding checks, and policy classification — all at a fraction of the cost and latency of a frontier LLM call.
What Is a Decision Model? Why Do Agent Developers Need It?
In AI agent workflows, a large fraction of decisions are fundamentally "pick one from a finite set": the user says "turn on the lights" — which tool should the agent call? Are the arguments grounded in what the user actually said? Is it too early to call this tool? Answering these with GPT-6 or Claude Opus 5.5 is slow, expensive, and lacks reliable confidence scores.
After TypeSafe AI launched Jev in mid-September 2026, the "decision model" concept (also called System One model) gained rapid industry traction. The core idea: take a pre-trained LLM torso, remove its text-generation ability, and replace it with a head that scores between options. You sacrifice flexibility for speed, accuracy, and controllability — the model always outputs one of the predefined options, never freeform text.
AWS Distinguished Engineer Marc Brooker saw Jev and tried building his own version. The project was successful enough to briefly reach #1 on the JevBench ranking for its size class, leading AWS to clean it up and release it as a Strands Labs project.
Technical Architecture of Strands Decider 2B
Base Model and Modification
- Base model: Qwen3.5-2B (Alibaba open-source)
- Modification: LM Head removed, replaced with Pointer Head
- Pointer Head parameters: ~1 million (1M+), much smaller than the base
- Fine-tuning: Rank-16 LoRA adapter
- Current release: v19 (19 architecture iterations)
The Pointer Head works by scoring the hidden state at each option position against the hidden state at the <answer> position. A single forward pass produces the probability distribution across all options — orders of magnitude faster than having an LLM generate an answer token by token.
Supported Question Types
| Type | Description | Example |
|---|---|---|
| noul (yes/no) | Returns Yes/No + confidence | "Are the tool's argument values grounded in facts the user provided?" |
| choice (multi-select) | Selects from N options + confidence | "Which team should handle this? billing/sales/retail" |
| score (ordinal) | Returns a score between 0 and 1 | "Is this phrase positive sentiment?" |
Each question type comes with calibrated confidence scores — something frontier LLM inference APIs cannot directly provide.
Performance: JevBench and Latency
Accuracy and Calibration
JevBench is TypeSafe's public evaluation set for decision models. Strands Decider 2B on the public set:
- Accuracy: 167/231 ≈ 72.3%
- Brier score (calibration): ~0.348 (lower is better)
- ECE (Expected Calibration Error): ~0.050
- Ranking: 3rd of 33 in the 2B class; 1st of 30 excluding models just over 2B
Internal benchmarks: held-out short 0.641 (n=6000), ContractNLI 0.872, MuSiQue 0.884.
Latency
| Hardware | Median Latency | Notes |
|---|---|---|
| NVIDIA RTX 3090 | ~115ms | Scales approximately linearly with task token count |
| Apple M3 MacBook | ~153ms | Small tasks, not much worse than RTX 3090 |
Key advantage: decision latency scales approximately linearly with task size, without the KV cache bloat problem of LLMs. 115ms means it can sit in paths where an LLM call could never go — like checking arguments before every single tool call.
Practical Application: Gating in Agent Workflows
Installation
pip install strands-deciderBasic Usage
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 --state "Help! My payouts have been failing for 3 days!" --choice "Which team should handle this?=billing,sales,retail"Output:
choice_0 -> billing (confidence 0.768)
billing 0.845
retail 0.091
sales 0.064before_tool_call Gating
The most canonical deployment is hooking Strands Decider into an Agent's InterventionHandler.before_tool_call to run two checks before any tool executes:
- Argument grounding: "Are the tool's argument values grounded in facts the user actually provided?" (yes/no)
- Prematurity check: "Is it premature to call this tool now, before clarifying with the user?" (yes/no)
If the arguments were guessed by the agent (not stated by the user), or the timing is wrong, the decision model returns Deny or Guide, blocking the tool call and sending the agent back to ask the user. The entire check completes in ~115ms — imperceptible to the user.
The handler returns typed actions: Proceed (let it through), Deny (block), Confirm (ask a human), or Guide (hand the model back its turn with feedback). The same InterventionHandler shape works whether you're calling Strands Decider, a Cedar policy, or another agent.
Hybrid Agent Architecture
The mature usage pattern is a hybrid agent: decision model handles simple, high-frequency decisions (the "easy tasks" it gets 100% correct), frontier LLM handles complex reasoning. This division can significantly reduce cost and latency while maintaining decision quality.
Open Source and Reproducibility
Strands Decider 2B is thoroughly open-sourced:
- Code: github.com/strands-labs/strands-decider (Apache-2.0)
- Weights: huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19
- Training data: Fully public
- Training scripts:
training/recipe.sh, ~1 hour on H100 cluster, ~11 hours on single RTX 3090 - Version history: Every change from v1 to v19 is documented, traceable
Marc Brooker told TechCrunch that building models in this class costs "hundreds to thousands of dollars" — frontier lab resources aren't required. TypeSafe CEO Diogo Almeida cautioned that most current imitators are "ML people wanting to implement a cool architecture" rather than "a team deeply dedicated to making intelligence useful."
Known Limitations
AWS transparently listed the model's limitations:
- Weak on long-document multi-step reasoning: Performs poorly on tasks requiring multi-step reasoning over long documents
- Question phrasing sensitivity: Rephrasing the question may yield the same answer (lacks true reasoning)
- Poor rubric migration: Swapping scoring rubrics for
score/noultypes can degrade performance significantly - Not a general-purpose LLM: Cannot be used for coding, chat, document summarization, or any task requiring text generation
Always calibrate thresholds with your own traffic before production deployment — this is AWS's official recommendation. Don't use default parameters in production.
Industry Trend: Decision Model Segment Heating Up
Strands Decider 2B's release is part of a rapid acceleration in the decision model space:
- 2026-09-18: TypeSafe AI launches Jev, creating the decision model category
- 2026-09-28: Marc Brooker's personal implementation hits #1 on JevBench for its size class
- 2026-09-30: OpenAI announces a Jev-like decision model offering
- 2026-10-01: AWS Strands Labs releases Strands Decider 2B
- 2026-10-02: Cloudflare releases Clef decision models (covered in our previous report)
TechCrunch noted that "dozens of similar models" have emerged in just two weeks since TypeSafe debuted Jev. Decision models may become a standard component of AI agent infrastructure — like routers for networks, decision models for agent workflows.
Developer Action Items
- Try it now:
pip install strands-decider, run the CLI examples to feel the output format - Identify use cases: Audit your agent workflow for decisions that are "pick one from a finite set" — these are candidates for decision model replacement
- Calibrate thresholds: Test confidence thresholds with your actual traffic to find the Proceed/Deny/Guide boundary
- Build hybrid: Don't try to replace all LLM calls — build a "decision model for simple decisions + LLM for complex reasoning" hybrid pipeline
- Watch calibration: Brier score and ECE matter more than raw accuracy — a 72% accurate but well-calibrated model is more useful than an 85% accurate but overconfident one
Sources: Strands Agents Blog, TechCrunch, GitHub Repository, Hugging Face Weights, JevBench Leaderboard
Topic hub
AI Agent Tutorials & Workflow Guides
Evergreen how-tos for coding agents, content pipelines, and n8n automation—linked to news context and real earn cases.
Explore AI Agent Tutorials & Workflow Guides →Monetization angle
How can you make money from this trend?
WayToClawEarn focuses on verified earn playbooks—not just news. Start from these cases.
DeepSeek + Claude Code Micro SaaS
Run multiple small products on cheap inference
Claude Code bug bounty
Productize agent skills into security services
Related tutorials
Related news
- Cloudflare Open-Sources Clef Decision Models: Returns Probabilities Not Text, How to Cut AI Agent Costs
- DeepSeek Harness v0.2 Desktop Launch: Zero-Config AI Coding Agent, How to Choose?
- ElevenLabs Doubles to $22B Valuation with Eleven v4 Turbo: How to Choose Voice AI Agents?
- OpenAI DevDay 2026: GPT-6.1 Sol at 1/5 Astra Price, Dots Autonomous Agent, Codex Cloud — What Developers Need to Know