Key takeaways

  • ServiceNow CoreAI published the AutoSynthData method on October 2, 2026. It uses target-model failures and stronger-teacher successes to generate tasks near the target model’s current capability boundary.
  • In the reported Hybrid experiment, AutoSynthData generated 2,000 tasks in about 18 hours using Gemma-4-26B-A4B-it as target and Qwen3.8-27B as teacher. ServiceNow reports a 7.2 percentage-point Pass@1 gain and calls that a 35% relative improvement.
  • A second ITSM run generated 1,994 tasks in 66 hours using DeepSeek-V4.1-Flash as teacher; reported mean Pass@1 rose from 18.77% to 27.18%. Both findings are vendor-run results in the benchmark environment.
  • The pipeline replays intended solutions, checks positive and mutated negative outcomes, limits repair attempts and reviews batch diversity. These controls are promising design details, not proof that every accepted verifier or sample is correct.
  • What is public now is the method description and the earlier Apache-2.0 EnterpriseOps-Gym dataset and evaluation framework. ServiceNow did not publish AutoSynthData code, the generated training corpora, fine-tuned checkpoints, an API, pricing or a release schedule in the reviewed announcement.
01

ServiceNow published the method and two training results on October 2

ServiceNow CoreAI described AutoSynthData on October 2, 2026 as a pipeline for turning an enterprise agent’s measured failures into new training tasks. In the headline Hybrid run, the target was google/gemma-4-26B-A4B-it and the teacher was Qwen/Qwen3.8-27B. ServiceNow says the system generated 2,000 synthetic samples in about 18 hours, then used them for supervised fine-tuning.

The reported best checkpoint, at epoch 5, improved mean Pass@1 by 7.2 percentage points, which the authors describe as a 35% relative gain. They also report verifier success moving from 63.01% to 68.55% and say the checkpoint closed 59% of the original Pass@1 gap to the reference model. These are ServiceNow-run results in EnterpriseOps-Gym Hybrid; no independent reproduction accompanies the post.

02

A second domain improved from 18.77% to 27.18%

ServiceNow also reports an ITSM experiment with the same Gemma target and DeepSeek-V4.1-Flash as teacher. AutoSynthData generated 1,994 training samples in 66 hours, and mean Pass@1 reportedly increased from 18.77% to 27.18%. The authors attribute the longer generation time partly to a larger teacher and to running ITSM before later pipeline optimizations.

Two positive domain results establish an internal replication across Hybrid and ITSM, not general performance across every enterprise workflow, model, teacher or training method. The announcement does not provide confidence intervals, per-task paired outcomes, failed-generation counts, compute specifications, token use, model charges or total fine-tuning cost.

03

The curriculum starts with failures rather than generic prompts

AutoSynthData first runs a target model and stronger teacher on diagnostic tasks. It records the capability being tested, tools and workflow structure, where the target fails, how the teacher succeeds, required final-state properties and dimensions that can vary. Those observations become sanitized capability specification cards.

ServiceNow says the generator does not receive the original evaluation prompts, entities, trajectories or verifier details. It instead creates new prompts, initial states, entity combinations, solution paths and verifiers from the cards. That separation is intended to reduce direct task copying, but AccessAllGPT could not inspect the private cards, generation prompts or resulting corpus to test semantic leakage.

04

Target, multiply and verifier gates form the production loop

The target phase creates independent core samples around identified gaps. The multiply phase makes variants only from accepted target samples; a multiplied example cannot seed another multiplied example, which limits generation drift. An environment adapter handles state setup, execution, reference replay, deterministic verification, solver runs and profiling while a shared controller coordinates coverage and dataset construction.

For the configuration described, ServiceNow favors candidates solved by the target on no more than one of three trials and by the stronger solver on at least two of three. Positive verification executes the reference trajectory and checks the resulting state. Negative verification mutates expected outcomes and requires those incorrect states to fail. A critic may diagnose and repair a failed candidate, but retries are bounded and repaired tasks must pass the gates again.

05

Batch review addresses repetition, not just sample validity

Individually executable tasks can still make a narrow training set. AutoSynthData therefore reviews accepted and rejected samples for overrepresented task families, missing capability dimensions, repeated examples and low-yield targets. Its controller reduces generation in crowded regions and redirects effort toward gaps.

That is a useful operational distinction: a passing verifier tests one candidate’s final state, while batch review asks whether the collection offers diverse learning signal. Neither check automatically establishes realism, unbiased coverage or transferable gains. Those properties require review against actual deployment traffic and a held-out evaluation that generation never used.

06

The public benchmark is not the new training release

EnterpriseOps-Gym is publicly available under Apache-2.0, with task prompts, selected tools, MCP server configuration and SQL-verifier fields across Calendar, CSM, Drive, Email, HR, Hybrid, ITSM and Teams. Its paper describes 1,150 expert-curated tasks, 512 tools and 164 database tables. At the pinned revision, the dataset card’s prose table totals 1,115 tasks while its machine-generated split metadata enumerates 649 examples in the oracle configuration and 637 in each distractor-tool configuration. Teams should therefore bind any reproduction to a revision and reconcile the paper, card and hosted files before comparing scores.

The AutoSynthData announcement uses that environment to generate new training examples, but does not link public AutoSynthData source code, the 2,000- or 1,994-sample corpora, trained checkpoints, generation manifests or a managed service. There is no announced API identifier, product tier, price or availability date. Readers can inspect the benchmark today; they cannot reproduce the full new pipeline from the announced artifacts alone.

07

The benchmark’s earlier findings explain the focus on planning

The EnterpriseOps-Gym authors previously reported a 37.4% best overall task-success rate among 14 evaluated models and a 53.9% best clean-refusal rate on infeasible tasks. They report that human-authored plans improved execution by 14–35 percentage points, while added distractor tools had little effect. Those results motivated treating strategic planning as the principal training gap.

They remain author-reported benchmark measurements from March, not proof that an AutoSynthData-trained model is safe for deployment. Higher Pass@1 can coexist with policy violations, unintended side effects or poor refusal behavior. Buyers need separate acceptance thresholds for task completion, state integrity, access control, refusal and recovery.

08

What agent teams should do next

Use the release as a design pattern, not a downloadable product. Start with one stateful environment where final outcomes can be checked independently. Freeze target and teacher versions; isolate training from held-out evaluation; preserve diagnostic tasks; log every generated candidate, repair and rejection; and measure tokens, compute, expert review and elapsed time per accepted sample.

Adopt the pattern when you own a resettable environment and can audit verifiers and final state. Constrain generated data to supervised experiments when policy or side-effect checks are incomplete. Wait for code, corpora and checkpoints if reproducibility is required. Reject a production rollout when improvement is demonstrated only on generation-adjacent evaluation or when the agent can mutate consequential state without deterministic authorization and rollback.

09

Copy-ready synthetic agent-data release gate

Complete this record before using failure-led synthetic tasks to train or approve an enterprise agent.

Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].

Domain, tools, state schema, reset behavior, policies, unavailable actions and production differences.

Target, teacher and critic IDs, revisions, inference settings, prompts, orchestration and fallback behavior.

Diagnostic inputs, capability cards, generation context, held-out set, overlap tests and leakage reviewer.

Feasibility, realism and difficulty thresholds; target and teacher trial counts; retry ceiling and rejection reasons.

Positive replay, negative mutations, alternative valid paths, policy checks, side effects and independent reviewer.

Accepted, repaired and rejected counts; families, duplicates, coverage gaps, provenance, license and sensitive-data checks.

Dataset revision, hyperparameters, epochs, checkpoints, compute, tokens, elapsed time, cost and failure logs.

Pass@1, verifier components, refusal, policy compliance, side effects, latency, cost, uncertainty and baseline.

Adopt, constrain, wait or reject; approved authority, human approvals, canary scope, rollback and revalidation trigger.

Primary sources

  1. AutoSynthData: Generating Training Data for Enterprise AgentsServiceNow CoreAI on Hugging Face · Reviewed: October 2, 2026 publication metadata; task definition; capability-card construction; target and multiply phases; sample and batch quality controls; Hybrid and ITSM experiment setups; reported training volume, elapsed time and evaluation changes; closing scope statement · Retrieved · Supports: ServiceNow describes AutoSynthData as an internal pipeline that converts target-model failures into generated, executable and verifier-backed agent-training tasks. The post reports two supervised fine-tuning experiments, but does not announce public AutoSynthData code, generated training sets, checkpoints, pricing or a managed API.
  2. EnterpriseOps-Gym dataset at revision c8e538eServiceNow-AI on Hugging Face · Reviewed: Pinned repository tree; dataset card metadata; Apache-2.0 license; oracle and distractor-tool configurations; eight domain splits; task schema; SQL-verifier description; dataset loading example; linked evaluation framework and paper · Retrieved · Supports: The pinned public dataset revision contains four configurations across eight domain splits and exposes task prompts, selected tools, server configuration and verifier fields under Apache-2.0. It is the pre-existing evaluation dataset used to illustrate AutoSynthData, not the newly generated AutoSynthData training corpus.
  3. EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise SettingsServiceNow CoreAI authors on arXiv · Reviewed: March 13, 2026 submission metadata; abstract; environment and task construction; SQL verification; model evaluation; human-plan ablation; infeasible-task results; limitations; dataset and framework references · Retrieved · Supports: The authors report a 1,150-task, eight-domain benchmark with 512 tools and 164 database tables. Their model study reports a 37.4% best overall success rate, 53.9% best clean-refusal rate on infeasible tasks and 14–35 percentage-point gains from human-authored plans. These are author-reported benchmark results, not independent measurements of AutoSynthData.

Limitations

AccessAllGPT reviewed public, author-controlled materials and did not independently reproduce AutoSynthData or EnterpriseOps-Gym results. We did not obtain AutoSynthData code, capability cards, generation prompts, generated training corpora, fine-tuned checkpoints, private evaluation traces, compute logs or cost records. We did not run the target or teacher models, replay a task, inspect SQL verifiers, test task novelty, measure contamination, validate the 18- or 66-hour durations, confirm the 7.2-point or 18.77%-to-27.18% gains, calculate statistical uncertainty, audit safety outcomes, or interview the authors. The March paper, dataset-card prose and generated split metadata expose different task counts—1,150, 1,115 and 649 in the oracle configuration at the pinned revision—which may reflect filtering, release scope or later revisions; we did not resolve them. All experiment results are vendor-reported and may not transfer beyond the tested configurations.

Disclosures

AccessAllGPT did not receive ServiceNow products, model access, generated data, checkpoints, code, compute, a briefing, review or compensation for this article. ServiceNow, Hugging Face, Google, Qwen and DeepSeek did not sponsor, review or endorse it. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with ServiceNow or organizations cited. Publication-wide relationships are listed on the disclosures page.

Further AccessAllGPT guidance

  1. Ai2 Releases AstaBrief 8B, an Open-Weight Model for Cited Scientific Reports
  2. Anthropic Publishes “Claude-Shaped Science” as BootLoops 1.0 Goes Open Source
  3. Design an Agent Benchmark That Predicts Production
  4. Where Human Approval Belongs in AI Automation
  5. AI Agents vs. Workflows: Choose the Right Automation Pattern
  6. Build or Buy an LLM Evaluation Platform?
  7. AccessAllGPT Research methodology
  8. Publication disclosures

Continue the research

Get evidence-led updates for teams making production AI decisions.