AI agents

AWS releases Strands Decider 2B for local agent validation

AWS launched Strands Decider 2B, a downloadable decision model that validates AI agent actions locally with sub-100ms latency.

Amazon Web Services has introduced Strands Decider 2B, a compact decision model designed to run locally and validate AI agent actions before execution. Released on October 1, 2026, this tool provides developers with a fast, reliable mechanism to check agent outputs without relying on external APIs or generating invalid text.

What happened

The launch positions AWS directly against emerging decision models like TypeSafe’s Jev, which recently sparked a wave of similar tools from major AI vendors. While OpenAI recently previewed its hosted Decisions API using the Luna model, AWS took a different approach by releasing a fully downloadable model. This package includes not just the weights, but also the training data and scripts, allowing engineers to inspect and adapt the underlying recipe.

Strands Decider 2B is part of the broader Strands ecosystem, incubated at Strands Labs, AWS’s experimental hub for agentic AI. It follows the recent release of Strands Harness, which provides the infrastructure for running longer-lived agents. By offering an open, local alternative to hosted services, AWS aims to give developers more control over the critical validation steps in their agent workflows.

How it works

Decision models differ from standard large language models by restricting their output space. Instead of generating free-form text, they select from developer-supplied options or return numerical scores. Strands Decider uses Qwen3.5-2B as its base, referred to as the “torso.” The team removed the standard language-model head responsible for text generation and replaced it with a pointer head containing just over a million parameters. This head scores the provided answer options.

The backbone utilizes a rank-16 low-rank adaptation (LoRA) adapter. By limiting the possible outputs to only those supplied by the developer, the model cannot invent options that do not exist. This architecture ensures that every output is valid within the defined context, though it does not guarantee correctness in every instance. The trade-off yields faster decisions and calibrated confidence scores, which applications can use to determine next steps.

In practice, this allows for pre-execution checks. For example, if an agent proposes calling a weather tool without a specified city, the Decider evaluates whether the arguments are grounded in the conversation. If information is missing, the system can route the agent back to ask the user for clarification rather than executing a flawed tool call. This process runs through Strands’ intervention system, enabling developers to define policies for proceeding, denying, or requesting human confirmation.

Key details

  • The model is built on Qwen3.5-2B with a custom pointer head replacing the text generation layer.
  • It achieves decision times under 100 milliseconds on an Nvidia RTX 3090 GPU.
  • On an M3 MacBook, median response times for small tasks are approximately 150 milliseconds.
  • Strands Decider 2B ranks second among public models with roughly 2 billion parameters on JevBench.
  • It ranks first among public models that provide a full training recipe and data.
  • The release includes all previous architectural iterations in the repository for transparency.

Why it matters

For software engineers building autonomous agents, reliability is often the biggest hurdle. Generative models can hallucinate tool arguments or miss critical constraints, leading to errors that are difficult to debug. Strands Decider 2B addresses this by inserting a fast, deterministic check between the agent’s reasoning and its actions. This separation of concerns allows generative models to handle complex conversation and creativity, while the decision model handles routing and validation.

The local nature of the model is significant for latency-sensitive applications and data privacy. Running validation on-device eliminates the network round-trip time associated with hosted APIs. Furthermore, the inclusion of training data and scripts supports fine-tuning for specific domains. Developers are no longer black-boxed into a vendor’s proprietary logic; they can adapt the model to their unique operational constraints and verify its behavior through the provided benchmarks.

What you can do

  • Download the Strands Decider 2B model, training data, and scripts from the AWS repository.
  • Integrate the model into your agent workflow to validate tool arguments before execution.
  • Use the provided LoRA adapter structure to fine-tune the model on your specific decision tasks.
  • Benchmark the model’s latency on your local hardware to ensure it meets your real-time requirements.
  • Review the historical iterations in the repository to understand the architectural evolution and performance trade-offs.
  • Experiment with the intervention system to define custom policies for handling low-confidence decisions.

Tools from the Bytechap store

Keep reading

All stories