Most model releases try to do more. Strands Decider 2B does less, intentionally. It cannot write code, summarize documents, or hold a conversation. What it can do is pick between options and assign reliability scores to those picks, in around 115 milliseconds on local hardware. That narrow focus is exactly the point.
Strands announced Strands Decider 2B as part of its strands-labs initiative, a project the team set up earlier this year for hands-on experimentation with agentic AI approaches. The model sits in a category that has been attracting real attention lately: decision models, sometimes called system one models. TypeSafe AI’s Jev, launched just weeks before this release, is the most prominent comparable. The category is still new enough that no one has fully defined its ceiling.
What decision models actually do
Unlike a standard LLM, a decision model does not generate arbitrary text. It picks from a set of options you provide. Give it a string and ask whether it relates to the coffee machine or the lighting system, and it returns an answer from those choices plus a confidence score. Ask it to classify sentiment on a 0-to-1 scale, and it does that. Ask it to write you a Python function, and it cannot.
The tradeoff is real. Decision models are worse than reasoning models at complex problems. But the things they give up in flexibility, they gain in speed and consistency. They always return one of the options you specify. They never hallucinate an answer that wasn’t on the list. And they attach a calibration score to every decision, something frontier LLM inference APIs do not offer.
For agentic workflows, that combination matters. Routing a query to the right tool, selecting memory, applying guardrails, classifying a policy, these are all discrete decision problems. They don’t need paragraph-length reasoning. They need fast, reliable answers with a confidence signal attached.
Model specs and architecture
Strands Decider 2B is built on a Qwen3.5-2B base. The team removed the standard language model head and replaced it with a pointer head that scores each option against hidden states in the model. The pointer head is small, just over a million parameters. The base model is fine-tuned with a rank-16 LoRA adapter. The result is a 2 billion parameter model that runs on a local CPU or GPU without specialized infrastructure.
Performance on JevBench’s public evaluation set puts the model third out of 33 models in the 2B class, and first among those at or under 2B parameters. Brier score calibration tracks alongside accuracy, which matters because a model that is confident and wrong is worse than a model that knows what it doesn’t know. Latency sits at a median of around 115ms on an Nvidia RTX 3090. On an M3 MacBook, that rises to roughly 153ms for small tasks, which is still well within the range required for real-time agent decision loops.
Where developers are already using it
The Strands team has shared early use cases from developers building with the model. The range is wider than you might expect for such a constrained tool:
- Model routing and tool selection in agentic pipelines
- Guardrails and policy classification
- Memory and context management
- Hybrid agent architectures that pair decision models with LLMs, using the faster model for routine calls and the LLM only for harder decisions
- Game-playing, maze navigation, and task automation
The hybrid agent pattern is the one worth watching. Using a decision model for the easy, repetitive choices in an agent loop and reserving LLM calls for genuinely hard decisions cuts both cost and latency. As agentic systems get more complex and run longer chains of actions, that kind of efficiency starts to compound.
Access and availability
The model is fully open source. Weights are on Hugging Face under the identifier StrandsAgents/strands-decider-2B-hobson-v19. The full repo is on GitHub and includes training data, training scripts, architecture notes, and version history across all 19 iterations. The team documented every architectural change, which makes it a useful reference for anyone who wants to build their own decision model rather than just use this one.
Getting started takes a single pip install. The CLI supports both choice questions and numerical scoring tasks, and the repo includes examples that connect the model to a Strands agent running locally alongside Amazon Bedrock for LLM calls.
The decision model category is early. But the direction is clear: agentic systems need faster, cheaper, more reliable components for the parts of the workflow that don’t require full language generation. Strands Decider 2B is a concrete, open contribution to that problem, and the benchmark numbers suggest it’s a competitive one.




