ServiceNow Releases AutoSynthData for Generating Enterprise Agent Training Data
Key point
The framework improves Pass@1 scores by up to 8.4 percentage points in ITSM and Hybrid enterprise environments by targeting model capability gaps.
Details
ServiceNow CoreAI has released AutoSynthData, a framework designed to generate synthetic training data for enterprise agents by identifying and addressing specific model weaknesses. The system analyzes failures of a Target model against successes from a stronger Teacher model to identify Capability Gaps, which are then converted into new, verified tasks for post-training.
Mechanism and Pipeline
The process defines an agentic task as a tuple of system specification, user prompt, and verifier. The pipeline operates in a loop:
- Diagnosis: Evaluate the Target model to identify failure patterns.
- Characterization: Use a Teacher model to define successful behaviors and feasibility.
- Generation: Create new tasks based on sanitized capability specification cards, ensuring they are feasible, realistic, and difficult enough to expose weaknesses.
- Verification: Tasks undergo positive verification (valid solutions accepted) and negative verification (invalid outcomes rejected) to ensure soundness.
Quality Control and Curriculum
AutoSynthData employs rigorous quality control to prevent data drift and ensure difficulty calibration:
- Sample-level: Tasks are preferred if the Target model fails them (≤1/3 success rate) while the Teacher model succeeds (≥2/3 success rate).
- Batch-level: A meta-review process monitors for overrepresented task families and adjusts resource allocation to cover missing capabilities.
- Curriculum: As the model improves, the training frontier moves to harder tasks, excluding those the model now consistently solves.
Experimental Results
Experiments were conducted in the EnterpriseOps Gym environment using Gemma-4-26B-A4B-it as the target model.
Hybrid Environment
- Teacher Model: Qwen3.8-27B
- Data: 2,000 synthetic samples generated in ~18 hours.
- Outcome: Mean Pass@1 improved by 7.2 percentage points (a 35% relative improvement). The verifier success rate rose from 63.01% to 68.55%, closing 59% of the gap between the target and reference models.
ITSM Environment
- Teacher Model: DeepSeek-V4.1-Flash
- Data: 1,994 synthetic samples generated in 66 hours.
- Outcome: Mean Pass@1 increased from 18.77% to 27.18%, demonstrating consistent improvements across different enterprise domains.
The framework currently focuses on Supervised Fine-Tuning (SFT) but is designed to support Reinforcement Learning (RL) by providing difficulty-calibrated learning signals.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.