ServiceNow CoreAI has posted AutoSynthData on Hugging Face to build training data for agents in an enterprise environment. The post says a broadly capable model can still fail in a particular setting: a workflow it handles poorly, a combination of tools it misuses, or a constraint it does not respect. One failure does not say how to train. Training needs many new tasks that test the same capability in different situations. Those tasks have to be possible in that environment, look like requests a person would actually make, and come with a reliable check for success.

[1]
Eight empty boxes on pale paper, with three dark circular holes below the left ones and torn scraps beside them.
Eight boxes along the top are empty. Only three round holes below the left side are punched through to black, with scraps beside them. That is keeping a few tasks worth training on, and leaving the rest blank. An illustration, not an experiment screenshot., AI-generated illustration, not a news photograph

A task has three parts: a system specification, a user prompt, and a verifier. The specification covers instructions, environment policies, and any initial state that is needed. It should not add arbitrary constraints just to manufacture difficulty. A user prompt has to be feasible, realistic, and still too hard for the current model to solve consistently. The verifier should agree with the prompt, the specification, and the state. It should reject trajectories that fail the task or break a constraint, and it should accept other valid solutions rather than one reference path. A lax verifier can reward the wrong behavior. An overly tight one can penalize a valid solution.

In the EnterpriseOps Gym experiment, both the target model and a stronger teacher run the evaluation tasks. The team records the capability being tested, the tools and workflow, where the target fails, how the teacher succeeds, and what a correct final state must satisfy. Those findings become capability cards. The evaluation tasks decide what to learn. The generator does not receive the original prompts, entities, trajectories, or verifier details. It receives the cards and builds new tasks with different prompts, states, and solution paths.

[1]

Dataset construction has two phases. The target phase generates core samples from the cards in parallel. Each candidate goes through validation, execution, solver evaluation, and repair. The multiply phase expands accepted samples into new variants, each with its own request, state, entities, reference trajectory, and verifier. A multiplied sample cannot seed another multiplied sample.

In the configuration used here, they favor tasks the target model solves on no more than one of three trials and the stronger solver solves on at least two of three. Positive verification runs the reference trajectory in the environment and checks the result against the verifier. Negative verification changes part of the expected outcome and confirms those states no longer pass. Failed candidates go to a critic and can be repaired a limited number of times, then the gates run again. A batch review also asks which task families are overrepresented and which capabilities are missing, and it changes the next batch accordingly.

The post says the useful training distribution moves as the model improves. After fine-tuning, tasks the model now solves reliably are less useful, and persistent failures still say what to generate next. The experiments are supervised fine-tuning. The same mechanism could support reinforcement learning. They say they plan to test that. This release did not.

[1]

On the Hybrid domain, the target is Gemma-4-26B-A4B-it and the teacher is Qwen3.8-27B. About 18 hours produced 2,000 synthetic samples. The best checkpoint was epoch 5. Mean Pass@1 rose by 7.2 percentage points, a 35 percent relative improvement, and verifier success rose from 63.01 percent to 68.55 percent. The checkpoint closes 59 percent of the original Pass@1 gap between Gemma and the reference model. The post does not name that reference model, and it does not give Hybrid Pass@1 as a before-and-after pair. The training tasks were newly generated from capability cards. The generator did not receive the original evaluation tasks. The result is for EnterpriseOps Gym Hybrid, not for every enterprise environment.

On ITSM, the same Gemma is the target and the teacher is DeepSeek-V4.1-Flash. Generating 1,994 samples took 66 hours. The post says this run was slower mainly because the teacher was larger and because it came before later pipeline changes that improved throughput. Synthetic supervised fine-tuning raised mean Pass@1 from 18.77 percent to 27.18 percent. That is a gain in a second domain, not the Hybrid 7.2-point figure moved across.

[1]

要点

  • The generator receives capability cards, not the original prompts, entities, trajectories, or verifiers.
  • In the configuration used here, the target solves at most one of three trials, and the stronger solver at least two of three.
  • Hybrid: about 18 hours produced 2,000 samples. Gemma’s mean Pass@1 rose 7.2 percentage points, and verifier success rose from 63.01% to 68.55%.
  • ITSM: 1,994 samples took 66 hours, and Pass@1 rose from 18.77% to 27.18%. The experiments are supervised fine-tuning. Reinforcement learning was not tested.