AutoSynthData: Generating Training Data for Enterprise Agents

ServiceNow CoreAI built AutoSynthData to turn capability gaps in enterprise agents into training data. The system evaluates a target model in its environment, uses failures and a stronger teacher's successes to identify what the model needs to learn, then generates and validates new tasks. As the model improves, the curriculum shifts toward remaining difficulties.
Tasks are defined as a system specification, user prompt, and verifier. A useful task must be feasible in the environment, realistic, and difficult enough that the current model does not solve it reliably. Verifiers must be consistent, sound, and complete.
AutoSynthData evaluates the target model on diagnostic tasks and distills findings into sanitized capability specification cards. The generator receives only these cards, not the original evaluation prompts, entities, trajectories, or verifier details. Generation has two phases: a target phase that creates core training samples and a multiply phase that expands accepted samples into novel variants. Each candidate undergoes solver evaluation, positive and negative verification, and bounded repair before acceptance. Batch-level meta-review tracks coverage, diversity, and redundancy.
In experiments with EnterpriseOps Gym, AutoSynthData generated 2,000 synthetic training samples in about 18 hours for the Hybrid domain, using Gemma-4-26B-A4B-it as the target and Qwen3.8-27B as the teacher. Fine-tuning on this dataset produced a best checkpoint at epoch 5 that improved mean Pass@1 by 7.2 percentage points, a 35% relative improvement, and raised verifier success from 63.01% to 68.55%, closing 59% of the original Pass@1 gap between Gemma and the reference model.
In the ITSM domain, with DeepSeek-V4.1-Flash as the teacher, AutoSynthData generated 1,994 samples in 66 hours. The run took longer primarily because it used a larger teacher model and preceded pipeline optimizations. Synthetic SFT raised mean Pass@1 from 18.77% to 27.18%.
The experiments focus on supervised fine-tuning, but the same mechanism could support reinforcement learning: generate tasks that challenge the current policy, train, then move the generation target with the updated policy. The authors plan to test this difficulty-calibrated frontier beyond SFT.
Based on reporting from the original publisher. Visit the source for full context and later updates.