Strands Decider 2B works differently from a conventional generative model. AWS built it from the core of Qwen3.5-2B but removed the component responsible for text generation. In its place, the company added a smaller mechanism that scores the choices supplied to the model.
That means the output is a decision rather than a written explanation. The model can, for example, classify a request, pick which tool an agent should use, route a task to the appropriate destination or determine whether an action should proceed.
AWS also designed the model to report how confident it is in each choice. That feature is intended to give developers another signal when deciding whether an automated workflow should continue, ask for clarification or hand a decision to a different system.
The project originated with AWS distinguished engineer Marc Brooker after he experimented with building a model inspired by TypeSafe AI's Jev. His prototype briefly reached the top of the Jevbench ranking for models of its size, according to TechCrunch, before Amazon engineers developed it into the Strands Labs release.
“What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step — ‘what is the next thing for me to do here, based on where I am?’” Brooker told TechCrunch.
He said conversations with AWS customers showed that many agent workflows did not need the full capabilities of a large language model at every stage. Instead, some tasks could be reduced to choosing among known alternatives.
Brooker described the approach as offering “a workflow step that can be structured in a way that is more reliable, thanks to the confidence scores, thanks to the closed domain of answers, lower latency, potentially lower cost.”
AWS says Strands Decider 2B can run on local CPUs or GPUs and return answers in tens of milliseconds. In its own testing, the company reported median latency of roughly 115 milliseconds on an Nvidia RTX 3090 and about 153 milliseconds for small tasks on an M3 MacBook.
The company is evaluating the model across accuracy, confidence calibration and speed. On JevBench's public benchmark, AWS said Strands Decider 2B ranked third among 33 models in the 2B class and first among 30 models when slightly larger models were excluded.
AWS is positioning the model for uses including model routing, tool selection, evaluations, guardrails, memory management, context management and policy classification. It also sees a role for hybrid agent architectures in which larger language models handle difficult reasoning while smaller decision models take care of simpler choices.
The tradeoff is that Strands Decider 2B is far less flexible than a general-purpose language model. Because it cannot generate arbitrary text, AWS says it is not suited to tasks such as coding, chatbots or document summarization. Its parallel decision architecture also performs worse on complicated reasoning problems.
Brooker said that improving these models will require balancing speed and decision accuracy against the broader language capabilities inherited from their underlying models.
“There is a very careful balance to be found where you want to push its performance on accuracy and calibration on these kinds of tasks, without degrading its performance on understanding different languages, on having the kind of knowledge it has, which is what makes it general purpose and interesting and useful,” he told TechCrunch.
TypeSafe founder and CEO Diogo Almeida was more cautious about the wave of decision models following Jev.
“I get that people think it’s a gold rush, but they might be underestimating the difficulty of making the models actually smart,” Almeida told TechCrunch.
AWS has released Strands Decider 2B as open source through Strands Labs, along with its model weights, training data and training scripts. The company says developers can download the model, run it locally and build their own systems on top of it.
This analysis is based on reporting from TechCrunch & Strands.
Image courtesy of SQ Magazine.
This article was generated with AI assistance and reviewed for accuracy and quality.