Neural network decision node glowing blue with branching paths

AWS Strands Decider 2B: Open 2B Decision Model Released for Local Agents

Amazon Web Services quietly shipped a new class of model on October 1, 2026. Strands Decider 2B is not a chat model. It is a 2-billion-parameter decision model built on Alibaba’s Qwen3.5-2B-Base, released under Apache 2.0 with weights, training data, and scripts. The goal is fast, local, structured choices instead of token generation.

What Strands Decider 2B actually does

Traditional language models generate text. Decision models score predefined options and return a choice with a calibrated confidence in a single forward pass. AWS Strands Labs removed the language-modelling head from Qwen3.5-2B-Base and replaced it with a ~1 million parameter pointer head. The head compares the hidden state at the <answer> position against the hidden state at each option’s last token.

The torso is fine-tuned with a rank-16 LoRA adapter. The released checkpoint is v19. It does not output text, it outputs a decision. The model reads a state and typed questions, then returns a choice, a yes/no probability, or a score with calibrated confidence. Because there is no generation, latency stays low.

On JevBench v1 public accuracy, AWS reports 0.723 accuracy, 167 of 231 tasks correct. That places Decider 2B third of 33 in the 2B class and first of 30 when excluding models just over 2B. The calibration is competitive with other models in this class.

Local first, no hosted API

Strands Decider 2B is intended for local experimentation and production guardrails. It runs on a consumer GPU or CPU. AWS measured a median latency of about 115 ms on an RTX 3090, with p95 at 299 ms. The model also runs on Apple silicon Macs. There is no hosted inference service from AWS; developers self-host.

The repo ships with a CLI and an HTTP server. The bundled server binds to 127.0.0.1 with no authentication, so production deployments need an auth layer. The weights are on Hugging Face as StrandsAgents/strands-decider-2B-hobson-v19. Installation is pip install strands-decider.

Training recipe is public. About 11 hours on one RTX 3090 or 1 hour 10 minutes on eight H100s. The data sources and a preregistered record of every training run are included, which is unusual for a commercial lab.

Architecture and why it matters

The core idea is “system one” reasoning: fast pattern matching instead of slow chain-of-thought. The model keeps Qwen3.5-2B-Base’s multimodal vision tower. Requests can carry base64 images as part of the state, using the same v19 checkpoint.

Use cases Strands Labs highlights: model routing, tool selection, evaluations, guardrails, memory management, context pruning, and policy classification. A hybrid agent can use a large LLM for hard decisions and Decider 2B for rote decisions, cutting cost and latency.

The intervention system in the Strands Harness SDK shows a concrete pattern. Before a tool call, Decider 2B reads the conversation and the proposed tool call and answers yes/no questions: are argument values grounded in the user’s actual statement, and is it too early to call the tool. The agent can then proceed, deny, confirm, or guide back to the user.

How it compares to other decision models

October 1, 2026 was a busy day for open decision models. Cloudflare released Clef and Clef-flash, 27B and 9B open-weight decision models hosted on Workers AI at 209 ms and 38 ms median latency. Amazon’s Strands Decider 2B is smaller and CPU friendly.

Compared to Jev, the model that opened this class in September, Decider 2B is Apache 2.0, fully open weights, and includes training data. It is weaker on hard problems than larger Jev variants, but it is free to run locally without a hosted bill.

The decision model trend reflects a broader shift toward structured outputs. Instead of parsing free text, agents can ask a model to pick among options with a probability. That reduces hallucination in guardrails and tool use.

What this means for local AI developers

Strands Decider 2B lowers the barrier to running decision logic on edge hardware. A 2B model with LoRA fits on a laptop GPU and runs in tens of milliseconds. That makes it practical for real-time agent interventions without cloud latency.

The open release also means reproducible research. With weights, data, and scripts, researchers can fine-tune on domain-specific choices, e.g., medical triage options, code review approvals, or fraud detection thresholds.

AWS is positioning Strands Labs as an experimental arm. Decider 2B is the first contribution, and the company expects this class to drive innovation in agentic AI over the next year. If decision models become a standard component, we may see more small, specialized checkpoints instead of one giant generalist.

Hardware for running local decision models and LLM inference

Decider 2B runs on modest hardware, but if you want headroom for larger MoE models and vision workloads, a modern GPU helps. The ASUS ROG Astral RTX 5090 White OC Edition is a 32GB GDDR7 card built for extreme workflows and local LLM inference. ASUS ROG Astral GeForce RTX 5090 32GB GDDR7 White OC Edition Graphics Card

For a turnkey build, the Corsair Vengeance i5200 pairs an Intel Core Ultra 9 285K with an RTX 5090, 64GB DDR5, and 4TB total storage. Corsair Vengeance i5200 Gaming PC – Liquid Cooled Intel Core Ultra 9 285K CPU, NVIDIA GeForce RTX 5090 GPU, 64GB Dominator Titanium RGB DDR5 Memory, 2+2TB M.2 SSD

HP’s Omen Max 45L with Ryzen 9 9900X3D and RTX 5090 is another high-end prebuilt for local AI experimentation and 4K gaming. HP Omen MAX 45L 5090 Gaming Desktop – AMD Ryzen 9 9900X3D, GeForce RTX 5090, 32GB DDR5 RAM, 2TB SSD

What this means

Strands Decider 2B validates a new pattern: small, open-weight decision models that run locally and return structured choices with confidence. Released October 1, 2026 under Apache 2.0, it is built on Qwen3.5-2B-Base with a pointer head, runs in ~115 ms median on an RTX 3090, and ships with full training artifacts. For developers building agents, it offers a cheap, fast guardrail layer that can be self-hosted without per-token costs.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *