Key takeaways

  • On October 2, 2026, AWS published a SageMaker AI MTRL example using model ID qwen3.6-27b in us-west-2. This is a technical result and implementation guide, not a new foundation-model launch.
  • AWS reports that nDCG@10 improved on WixQA, Wands and BrowseComp-Plus, while FreshStack declined slightly. BrowseComp-Plus rose from 0.5136 to 0.6354 and its reported failure rate fell from 22.89% to 0.68%.
  • The setup optimized one trajectory-level retrieval reward, nDCG@10, and assigned -1 when the agent exceeded its turn or per-turn token budget. Better retrieval ranking does not by itself prove better final-answer accuracy or safety.
  • The example used six training datasets, reserved 5% of each for validation and evaluated on four held-out datasets totaling 2,049 questions. AWS did not publish repeated-run variance, total training tokens, wall-clock duration or total cost in the post.
  • Do not budget from the blog’s “per-token” wording alone. AWS’s live pricing page describes general RL customization as model-specific hourly billing, and the reviewed page did not present an unambiguous Qwen3.6-27B MTRL rate.
01

AWS published the search-agent result on October 2

AWS published a worked example on October 2, 2026 showing how its managed SageMaker AI multi-turn reinforcement-learning service can specialize a smaller model for iterative retrieval. The example uses the exact SDK model ID qwen3.6-27b in US West (Oregon), or us-west-2, with an external agent endpoint exposing BM25 lexical search and vector search tools.

This is not a Qwen model release or proof that a generally available endpoint was upgraded. It is an AWS-run fine-tuning and evaluation report for one search-agent design. The news for builders is the concrete multi-turn training recipe and held-out result set; availability, supported models and regions still need to be checked against the live account and documentation.

02

The reward covers the completed search trajectory

The agent can issue multiple searches, observe results and decide what to do next. AWS used nDCG@10—the quality of the top ten retrieved documents relative to an ideal ranking—as the trajectory-level reward. A score of 1 is a perfect ranking and 0 means no relevant document was retrieved. If the agent hit its maximum turns or per-turn sampling-token limit, the run received a -1 reward.

That design directly trains retrieval behavior and completion within a budget. It does not directly score whether a synthesized answer is factual, whether citations support every claim, whether sensitive content was disclosed, or whether retrieved text manipulated the agent. Teams need separate answer-grounding and safety evaluations before deployment.

03

The example trains across six datasets and tests on four others

AWS lists FRAMES, BRIGHT, Enterprise RAG, ESCI, Musique and MLQA as training sources and reserves 5% of the training instances in each source for validation. The held-out evaluation uses 400 WixQA questions, 147 Wands questions, 672 FreshStack questions and 830 BrowseComp-Plus questions—a total of 2,049 test questions.

The split across support, product search, developer Q&A and deep-research retrieval is useful, but “held out” is not the same as independently audited. The post does not provide the sampled prompt IDs, generated trajectories, reward code, random seeds or repeated-run variance needed to reproduce the comparison or quantify training instability.

04

AWS reports three gains and one small regression

On WixQA, AWS reports nDCG@10 moving from 0.5725 to 0.6781, an 18.4% relative gain, while failures moved from 0.67% to 0.17%. Wands rose from 0.5762 to 0.6112, a 6% relative gain, with no reported failures in either condition. BrowseComp-Plus increased from 0.5136 to 0.6354, a 23.7% relative gain.

FreshStack moved the other way: 0.4112 for the base model and 0.4089 after fine-tuning, even as its failure rate fell from 0.20% to 0.05%. That regression matters. The result supports testing a specialized policy; it does not support a claim that MTRL improves every retrieval domain.

05

The largest change is fewer BrowseComp-Plus failures

AWS defines a failed task as one that hits an error such as the turn limit or token budget, and assigns failed tasks an nDCG@10 of zero. On BrowseComp-Plus, that rate fell from 22.89% to 0.68%. Average turns also fell from 7.0 to 6.3. This means some of the ranking-score gain is coupled to the agent finishing rather than timing out.

That is operationally useful, but it makes the failure contract part of the result. A buyer should reproduce the same test with its own turn ceiling, token ceiling, tool latency, retrieval index and error policy. Raising limits or changing timeout treatment could alter both the failure rate and aggregate nDCG.

06

The public post leaves training cost unresolved

The published SDK example changes only max_epochs to 1, global_batch_size to 128 and rollout_max_concurrency to 32. AWS says MTRL is serverless, supports resumable jobs, exposes trajectories and rewards through managed MLflow, and can evaluate before deployment to SageMaker AI or Amazon Bedrock. It does not disclose the example’s total training tokens, elapsed time, checkpoints, generated rollout volume or bill.

Pricing needs a preflight rather than an assumption. The blog says MTRL uses per-token pricing, while the live SageMaker pricing page says general RL customization is billed by job duration at an hourly rate that varies by model. The reviewed page labels MTRL fine-tuning and evaluation but did not expose a clear rate for qwen3.6-27b. Confirm the billable meter, region, training and evaluation rates, endpoint costs and stop conditions in the console or a written AWS quote before starting a large job.

07

What search-agent teams should do next

Run a bounded trial only when the production task has relevance judgments or another verifiable trajectory reward. Freeze a base-model control, pin the corpus and index, publish prompt IDs, set identical turn and token budgets, repeat the run across seeds, and report confidence intervals. Add final-answer citation checks, unsupported-claim rates, malicious-document tests, latency and cost per successful task.

Adopt if improvements survive those controls on the organization’s own queries. Constrain deployment when retrieval improves but answer grounding, privacy or tool authority remains unresolved. Wait when supported-model availability or pricing cannot be verified. Reject a rollout that treats one vendor-run nDCG table as proof of general reasoning quality, production reliability or safety.

08

Copy-ready multi-turn RL trial record

Complete one record for each base model, agent environment, reward and production corpus.

Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].

Trial, constrain, wait or reject; use case, owner, users, region, review date and expiry.

Model ID, revision where available, SageMaker region, SDK version, supported status and deployment target.

Tool schemas, BM25 and vector indexes, turn ceiling, token ceiling, timeout, retry and terminal-state rules.

Dataset names and versions, prompt IDs, deduplication, split, reward code, penalties, epochs, batch size, concurrency and seeds.

Frozen held-out set, base control, repeated runs, confidence intervals, nDCG, failure rate, turns, answer support, injection and privacy tests.

Confirmed billing meter and rate, rollout tokens, duration, evaluation, storage, MLflow, endpoints, Bedrock charges, budget alert and stop threshold.

Read-only tools, corpus access, credentials, sensitive-document controls, egress, action restrictions and human approval.

Base-model route, model and prompt versioning, monitoring, regression trigger, endpoint removal and artifact retention.

Measured benefit, unresolved regression, approved traffic share, owner, next review and evidence required to expand.

Primary sources

  1. Fine-tune a search agent with multi-turn RL on Amazon SageMaker AIAWS Artificial Intelligence Blog · Reviewed: October 2, 2026 publication metadata; MTRL capabilities; environment and datasets; nDCG@10 reward; Qwen3.6-27B job configuration; training progress; held-out results; cleanup; conclusion and next steps · Retrieved · Supports: AWS reports a SageMaker AI multi-turn reinforcement-learning run on Qwen3.6-27B that improved nDCG@10 on three of four held-out retrieval sets and reduced the reported BrowseComp-Plus failure rate from 22.89% to 0.68%.
  2. Multi-turn reinforcement learningAmazon SageMaker AI Developer Guide · Reviewed: MTRL overview; rollout and update architecture; supported losses and advantage estimators; agent requirements; training assets; job submission; evaluation; deployment; limitations and linked workflow pages · Retrieved · Supports: AWS documents MTRL as a managed model-customization workflow that collects multi-turn trajectories against a customer agent environment, applies sequence-level rewards and supports evaluation and deployment workflows.
  3. Amazon SageMaker AI pricingAmazon Web Services · Reviewed: Pricing overview; model-customization navigation; training, evaluation and data-generation descriptions; reinforcement-learning billing language; MultiTurn Reinforcement Learning fine-tuning and evaluation labels; pricing examples · Retrieved · Supports: The live pricing page says general reinforcement-learning customization is billed by job duration at a model-specific hourly rate, while the October 2 blog characterizes MTRL as per-token. The reviewed public page did not expose a clear Qwen3.6-27B MTRL rate, so a production estimate requires a region- and model-specific quote or console check.

Limitations

AccessAllGPT reviewed public AWS material but did not access an AWS account, confirm account-level feature availability, run Qwen3.6-27B, call the MTRL SDK, deploy BM25 or vector-search tools, inspect training prompts, generated trajectories, reward code, checkpoints or MLflow records, validate dataset licenses or contamination, reproduce any nDCG or failure result, repeat runs across seeds, test final-answer correctness, prompt injection, privacy, latency or regional behavior, or obtain an itemized bill. AWS authored and executed the experiment, and the post does not disclose total training tokens, duration, cost, random seeds, sampled prompt IDs or uncertainty. Public billing language reviewed on October 4 was not sufficiently consistent to quote a Qwen3.6-27B MTRL price.

Disclosures

AccessAllGPT received no SageMaker, AWS, Qwen, dataset or search-system access, credits, briefing, pricing quote, review or compensation for this article. AWS and Qwen did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Amazon Web Services or Qwen. Publication-wide relationships are listed on the disclosures page.

Further AccessAllGPT guidance

  1. Microsoft and Hugging Face Bring ThinkingBox’s Stateful Agent Workflows to OpenEnv
  2. ServiceNow Details AutoSynthData for Enterprise Agent Training
  3. AWS Publishes a Secure Claude Desktop Route to AgentCore Web Search
  4. Design an Agent Benchmark That Predicts Production
  5. RAG Chunk Overlap: Stop Paying Twice for Repeated Evidence
  6. AI API Data Retention and Residency Procurement Gates
  7. AccessAllGPT Research methodology
  8. Publication disclosures

Continue the research

Get evidence-led updates for teams making production AI decisions.