WayToClawEarn
Medium impactStrands Agents Blog

AWS Strands Decider 2B Open-Sourced: 2B-Param Decision Model for Agent Routing, Tool Gating at 115ms

AWS Strands Labs released Strands Decider 2B, a 2-billion-parameter open-source decision model that replaces the LM head of Qwen3.5-2B with a pointer scoring head. It achieves 72.3% accuracy on JevBench public set with 115ms median latency on a local RTX 3090. Instead of generating freeform text, it selects between predefined options and outputs calibrated confidence scores — ideal for agent routing, tool gating, and argument grounding checks that don't need a full LLM call.

Edisen Lu · WayToClawEarnVia Strands Agents BlogPublished Oct 2, 2026

Reviewed from public sources · AI-assisted drafting under editorial oversight. How we work · Original source

Public-source compilation

Synthesized from public posts/docs. Prefer the original source for primary claims.

How we review content · Primary source · Strands Agents Blog

TL;DR

AWS Strands Labs released Strands Decider 2B on October 1, 2026 — a 2-billion-parameter open-source decision model built on the Qwen3.5-2B torso with the LM head removed and replaced by a ~1M-parameter pointer scoring head. It achieves 72.3% accuracy (167/231) on the JevBench public set with a median latency of ~115ms on a local RTX 3090. Fully open-sourced under Apache-2.0 with training data and scripts, installable via pip install strands-decider.

This marks the rapid mainstreaming of "decision models" (also called System One models) — a category pioneered by TypeSafe's Jev just two weeks earlier. Decision models don't generate freeform text; they select between predefined options and output calibrated confidence scores, making them ideal for high-frequency, low-complexity decisions in AI agent workflows: model routing, tool selection, argument grounding checks, and policy classification — all at a fraction of the cost and latency of a frontier LLM call.

What Is a Decision Model? Why Do Agent Developers Need It?

In AI agent workflows, a large fraction of decisions are fundamentally "pick one from a finite set": the user says "turn on the lights" — which tool should the agent call? Are the arguments grounded in what the user actually said? Is it too early to call this tool? Answering these with GPT-6 or Claude Opus 5.5 is slow, expensive, and lacks reliable confidence scores.

After TypeSafe AI launched Jev in mid-September 2026, the "decision model" concept (also called System One model) gained rapid industry traction. The core idea: take a pre-trained LLM torso, remove its text-generation ability, and replace it with a head that scores between options. You sacrifice flexibility for speed, accuracy, and controllability — the model always outputs one of the predefined options, never freeform text.

AWS Distinguished Engineer Marc Brooker saw Jev and tried building his own version. The project was successful enough to briefly reach #1 on the JevBench ranking for its size class, leading AWS to clean it up and release it as a Strands Labs project.

Technical Architecture of Strands Decider 2B

Base Model and Modification

  • Base model: Qwen3.5-2B (Alibaba open-source)
  • Modification: LM Head removed, replaced with Pointer Head
  • Pointer Head parameters: ~1 million (1M+), much smaller than the base
  • Fine-tuning: Rank-16 LoRA adapter
  • Current release: v19 (19 architecture iterations)

The Pointer Head works by scoring the hidden state at each option position against the hidden state at the <answer> position. A single forward pass produces the probability distribution across all options — orders of magnitude faster than having an LLM generate an answer token by token.

Supported Question Types

TypeDescriptionExample
noul (yes/no)Returns Yes/No + confidence"Are the tool's argument values grounded in facts the user provided?"
choice (multi-select)Selects from N options + confidence"Which team should handle this? billing/sales/retail"
score (ordinal)Returns a score between 0 and 1"Is this phrase positive sentiment?"

Each question type comes with calibrated confidence scores — something frontier LLM inference APIs cannot directly provide.

Performance: JevBench and Latency

Accuracy and Calibration

JevBench is TypeSafe's public evaluation set for decision models. Strands Decider 2B on the public set:

  • Accuracy: 167/231 ≈ 72.3%
  • Brier score (calibration): ~0.348 (lower is better)
  • ECE (Expected Calibration Error): ~0.050
  • Ranking: 3rd of 33 in the 2B class; 1st of 30 excluding models just over 2B

Internal benchmarks: held-out short 0.641 (n=6000), ContractNLI 0.872, MuSiQue 0.884.

Latency

HardwareMedian LatencyNotes
NVIDIA RTX 3090~115msScales approximately linearly with task token count
Apple M3 MacBook~153msSmall tasks, not much worse than RTX 3090

Key advantage: decision latency scales approximately linearly with task size, without the KV cache bloat problem of LLMs. 115ms means it can sit in paths where an LLM call could never go — like checking arguments before every single tool call.

Practical Application: Gating in Agent Workflows

Installation

terminal
pip install strands-decider

Basic Usage

terminal
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19   --state "Help! My payouts have been failing for 3 days!"   --choice "Which team should handle this?=billing,sales,retail"

Output:

code
choice_0 -> billing (confidence 0.768)
  billing                  0.845
  retail                   0.091
  sales                    0.064

before_tool_call Gating

The most canonical deployment is hooking Strands Decider into an Agent's InterventionHandler.before_tool_call to run two checks before any tool executes:

  1. Argument grounding: "Are the tool's argument values grounded in facts the user actually provided?" (yes/no)
  2. Prematurity check: "Is it premature to call this tool now, before clarifying with the user?" (yes/no)

If the arguments were guessed by the agent (not stated by the user), or the timing is wrong, the decision model returns Deny or Guide, blocking the tool call and sending the agent back to ask the user. The entire check completes in ~115ms — imperceptible to the user.

The handler returns typed actions: Proceed (let it through), Deny (block), Confirm (ask a human), or Guide (hand the model back its turn with feedback). The same InterventionHandler shape works whether you're calling Strands Decider, a Cedar policy, or another agent.

Hybrid Agent Architecture

The mature usage pattern is a hybrid agent: decision model handles simple, high-frequency decisions (the "easy tasks" it gets 100% correct), frontier LLM handles complex reasoning. This division can significantly reduce cost and latency while maintaining decision quality.

Open Source and Reproducibility

Strands Decider 2B is thoroughly open-sourced:

Marc Brooker told TechCrunch that building models in this class costs "hundreds to thousands of dollars" — frontier lab resources aren't required. TypeSafe CEO Diogo Almeida cautioned that most current imitators are "ML people wanting to implement a cool architecture" rather than "a team deeply dedicated to making intelligence useful."

Known Limitations

AWS transparently listed the model's limitations:

  1. Weak on long-document multi-step reasoning: Performs poorly on tasks requiring multi-step reasoning over long documents
  2. Question phrasing sensitivity: Rephrasing the question may yield the same answer (lacks true reasoning)
  3. Poor rubric migration: Swapping scoring rubrics for score/noul types can degrade performance significantly
  4. Not a general-purpose LLM: Cannot be used for coding, chat, document summarization, or any task requiring text generation

Always calibrate thresholds with your own traffic before production deployment — this is AWS's official recommendation. Don't use default parameters in production.

Industry Trend: Decision Model Segment Heating Up

Strands Decider 2B's release is part of a rapid acceleration in the decision model space:

  • 2026-09-18: TypeSafe AI launches Jev, creating the decision model category
  • 2026-09-28: Marc Brooker's personal implementation hits #1 on JevBench for its size class
  • 2026-09-30: OpenAI announces a Jev-like decision model offering
  • 2026-10-01: AWS Strands Labs releases Strands Decider 2B
  • 2026-10-02: Cloudflare releases Clef decision models (covered in our previous report)

TechCrunch noted that "dozens of similar models" have emerged in just two weeks since TypeSafe debuted Jev. Decision models may become a standard component of AI agent infrastructure — like routers for networks, decision models for agent workflows.

Developer Action Items

  1. Try it now: pip install strands-decider, run the CLI examples to feel the output format
  2. Identify use cases: Audit your agent workflow for decisions that are "pick one from a finite set" — these are candidates for decision model replacement
  3. Calibrate thresholds: Test confidence thresholds with your actual traffic to find the Proceed/Deny/Guide boundary
  4. Build hybrid: Don't try to replace all LLM calls — build a "decision model for simple decisions + LLM for complex reasoning" hybrid pipeline
  5. Watch calibration: Brier score and ECE matter more than raw accuracy — a 72% accurate but well-calibrated model is more useful than an 85% accurate but overconfident one

Sources: Strands Agents Blog, TechCrunch, GitHub Repository, Hugging Face Weights, JevBench Leaderboard

AWSStrands Deciderdecision modelopen source AIagent routingJevQwen3.5AI cost optimization

View source →

Educational reference only: cases summarize public sources and may use AI-assisted drafting under editorial review. Not financial advice; outcomes are not guaranteed.