ServiceNow has shipped a pipeline that converts a target model's own failures into validated, environment-grounded training tasks, then uses them to close capability gaps.
AutoSynthData, from ServiceNow CoreAI, generates agentic training data by evaluating a target model in a target environment, identifying where it fails and where a stronger teacher succeeds, and distilling those gaps into sanitized capability specification cards. The generator never sees the original evaluation prompts, entities, trajectories, or verifier details. It produces new tasks with different states, entities, and solution paths, then runs each candidate through a quality-control loop: solver evaluation, positive verification (does the intended solution pass the verifier?), negative verification (do mutated incorrect outcomes fail?), and a bounded critique-and-repair cycle. A batch-level meta-review then checks coverage, diversity, and redundancy, redirecting generation away from overrepresented task families.
The pipeline runs in two phases. The target phase generates core samples from capability cards in parallel; the multiply phase creates novel variants of accepted samples, each with its own user request, environment state, entity configuration, reference trajectory, and verifier. Multiplied samples cannot seed further multiplication, anchoring expansion to the vetted set. Generation control is separated from environment execution via a shared controller and an environment-specific adapter.
In EnterpriseOps Gym experiments, AutoSynthData generated 2,000 synthetic training samples in about 18 hours for the Hybrid domain using Gemma-4-26B-A4B-it as target and Qwen3.8-27B as teacher. Fine-tuning on those samples improved mean Pass@1 by 7.2 percentage points — a 35% relative gain — and raised verifier success from 63.01% to 68.55%, closing 59% of the original Pass@1 gap between Gemma and the reference model. A second run on the ITSM domain, using DeepSeek-V4.1-Flash as teacher, generated 1,994 samples in 66 hours and lifted mean Pass@1 from 18.77% to 27.18%.
The mechanism is explicitly designed to move with the model: after post-training, the updated model is re-evaluated, tasks it now solves reliably are deprioritized, and persistent failures guide the next generation round. The authors note the same difficulty-calibrated frontier could support reinforcement learning, though current experiments focus on supervised fine-tuning.