Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI
Amazon Web Services Huibin Shen
AWS says it fine-tuned a search agent with multi-turn RL on SageMaker AI. It cut failures hard on one hard benchmark and improved retrieval quality on several others.
Based on reporting by Amazon Web Services, Huibin Shen — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
AWS is pushing search agents past the usual prompt-and-pray setup. In a new post, the company says it fine-tuned a Qwen3.6-27B model with Amazon SageMaker AI multi-turn reinforcement learning so the agent could learn how to search, when to stop, and how to behave inside a specific environment.
The pitch is simple enough. A small model can be cheap and fast, but it usually needs help to handle multi-step search well. Supervised fine-tuning asks for expert demonstrations that are expensive or missing entirely. Single-turn RL looks at one response at a time, which is a poor fit when each query changes what the next query should be. MTRL, as AWS frames it, scores the whole trajectory instead.
The setup used BM25 and vector search tools behind an agent endpoint, with training and validation data uploaded to S3. The datasets span a mix of retrieval and question-answering work: FRAMES, BRIGHT, Enterprise RAG, ESCI, Musique, and MLQA for training, plus FreshStack, WixQA, BrowseComp-Plus, and Wands for testing. The reward was nDCG@10, with a -1 penalty when the agent hit the turn limit or the token budget.
AWS says it only changed three hyperparameters in the job: max_epochs was set to 1, global_batch_size to 128, and rollout_max_concurrency to 32. Everything else, including the algorithm choices and advantage estimators, stayed on defaults. That’s the kind of claim AWS clearly wants to make sound almost boring: the hard part becomes the environment and the reward, not wrestling a giant RL stack into existence.
The results were mixed in the honest way, which is usually the useful way. The fine-tuned model improved on three of four held-out benchmarks. BrowseComp-Plus went from 0.5136 to 0.6354 nDCG@10, WixQA from 0.5725 to 0.6781, and Wands from 0.5762 to 0.6112. FreshStack slipped slightly, from 0.4112 to 0.4089. The bigger story, though, was reliability: on BrowseComp-Plus, the failure rate dropped from 22.89 percent to 0.68 percent.
AWS also says the reward curves for training and validation rose steadily before flattening out, and that long runs can be resumed from checkpoints if a job times out. The whole thing reads like a quiet argument for a familiar idea: if an agent lives in a real tool loop, train it in the real tool loop. Strange that this still needs saying, but here we are.
My take — AI-written commentary, not fact-checked reporting
This is the kind of RL story that actually matters: not bigger models, just models that stop wandering off into the bushes. AWS is right to put the reward on the full trajectory and punish failure directly; search agents are supposed to finish the job, not audition for a poetry slam. The industry has spent enough time worshipping chatty demos—boring reliability is the upgrade people will pay for.
Read more about this at: Amazon Web Services