Fine-tuning search agents with multi-turn reinforcement learning on SageMaker
Amazon demonstrates how multi-turn RL fine-tunes a Qwen3.6-27B search agent, cutting failure rates from 22% to under 1% while improving retrieval quality.
Dieser Artikel ist nur auf Englisch verfügbar.
Amazon Web Services has published a technical deep-dive on using Amazon SageMaker AI to fine-tune large language model search agents with multi-turn reinforcement learning. Released in October 2026, the guide details how engineers can optimize smaller models for complex retrieval tasks without the high latency and cost of frontier models. The approach focuses on training agents to make better decisions across entire conversation sequences rather than isolated steps.
What happened
Search agents powered by LLMs are increasingly common in enterprise settings, but they face a reliability gap. While large frontier models can handle multi-step reasoning, they are expensive and slow. Smaller models are faster and cheaper but often fail to navigate complex tool environments effectively when prompted directly. Traditional supervised fine-tuning requires costly expert demonstrations that rarely exist for specific internal tools, while single-turn reinforcement learning misses the dependencies between sequential actions.
To address this, AWS introduced a workflow using Amazon SageMaker AI multi-turn reinforcement learning (MTRL). This method treats agentic tasks as a sequence of decisions, optimizing the model based on the final outcome of a multi-turn interaction. The team fine-tuned a Qwen3.6-27B model, demonstrating that this approach allows smaller models to achieve the reliability of larger ones by learning environment-specific behaviors directly. The process uses serverless infrastructure, removing the need to manage GPU clusters manually.
The results showed significant improvements in both retrieval quality and operational stability. By training the agent to maximize a trajectory-level reward signal, the system learned to avoid common failure modes like exceeding token budgets or getting stuck in loops. This enables organizations to deploy efficient, specialized search agents that understand their specific data landscapes and toolsets without relying on massive general-purpose models.
How it works
Amazon SageMaker AI MTRL frames the agent’s task as a series of decisions within an environment. Instead of scoring individual responses, it evaluates the entire multi-turn trajectory. The system generates training data through multi-turn rollouts, where the agent interacts with tools like lexical search (BM25) and vector search. It then optimizes the model using policy gradient algorithms such as Proximal Policy Optimization (PPO) or Clipped Importance Sampling Policy Optimization (CISPO).
A key component is the reward function. In this implementation, the team used nDCG@10 (Normalized Discounted Cumulative Gain at rank 10) as the primary metric. This standard information retrieval score measures how well the top ten retrieved documents match the ideal ranking. If the agent fails to complete the task within the allowed turns or token limits, it receives a penalty reward of -1. This negative feedback explicitly teaches the model to be efficient and avoid error states without requiring complex intermediate reward shaping.
The infrastructure handles the heavy lifting asynchronously. Rollout generation and gradient updates run in parallel, keeping training fast while maintaining bounded off-policy staleness. Engineers configure the job with minimal hyperparameters, such as batch size and concurrency limits, while the service manages algorithm selection and advantage estimation defaults. This low-code interface allows teams to focus on defining their tools and rewards rather than managing distributed training logistics.
Key details
- Model used: The experiment fine-tuned the Qwen3.6-27B model, supported in the US West (Oregon) region.
- Performance gain: On the BrowseComp-Plus benchmark, nDCG@10 improved by 23.7 percent, rising from 0.5136 to 0.6354.
- Reliability boost: The failure rate on BrowseComp-Plus dropped dramatically from 22.89 percent to just 0.68 percent after fine-tuning.
- Reward metric: The system optimized directly for nDCG@10, penalizing timeouts or token overruns with a -1 reward.
- Infrastructure: Training runs are serverless with per-token pricing, supporting resumable jobs if time limits are reached.
- Datasets: Training included FRAMES, BRIGHT, and Enterprise RAG datasets, while testing used FreshStack, WixQA, and Wands.
Why it matters
For software engineers building retrieval-augmented generation (RAG) systems, this approach offers a path to reduce costs without sacrificing quality. Many teams currently rely on large, expensive models because smaller open-weight models struggle with multi-step tool use. By fine-tuning a 27-billion parameter model with MTRL, developers can achieve comparable reliability at a fraction of the inference cost. This is particularly valuable for enterprise search applications where query volume is high and latency requirements are strict.
The reduction in failure rates is equally critical for production systems. An agent that frequently exceeds its turn limit or token budget creates poor user experiences and increases computational waste. The penalty-based reward design proved effective at teaching the model to recognize and avoid these boundaries. This means deployed agents are more predictable and easier to monitor, reducing the operational burden on engineering teams who must otherwise build complex guardrails around unstable base models.
Furthermore, the ability to train against direct task metrics like nDCG@10 aligns model optimization with business goals. Instead of optimizing for generic language likelihood, the model learns to retrieve relevant documents specifically for the company’s data structure. This specificity allows for better performance on domain-specific queries, such as internal technical documentation or product catalogs, where generic semantic similarity might miss crucial keyword matches or structural nuances.
What you can do
- Start by defining your agent’s available tools, such as BM25 and vector search endpoints, and ensure they are accessible via an API.
- Prepare your training and validation datasets in the format required by Amazon SageMaker AI MTRL, reserving five percent for validation.
- Define a trajectory-level reward function that reflects your specific success metric, such as nDCG@10 or exact match accuracy.
- Configure the MultiTurnRLTrainer SDK with default hyperparameters initially, adjusting batch size and concurrency only if needed.
- Monitor training progress using MLflow integration to inspect trajectories and ensure the agent is learning efficient search patterns.
- Evaluate the fine-tuned model on held-out benchmarks before deploying to production to verify improvements in both quality and failure rates.

