Rohan Paul

@rohanpaul_ai

New Tsinghua + Qwen Team's paper flips the usual agent-data recipe: Reconstructing and re-solving old agent workspaces trains better coding agents than imitating the original runs A trajectory is like a recording of 1 coding agent fixing 1 bug: you can only copy what that agent did. The paper instead rebuilds the actual code workspace from that recording, so you can give the same workspace new bugs, new features, or a stronger agent and generate many fresh training examples from it. That is the core difference: trajectory = 1 frozen solution; environment = a reusable codebase that can produce many new verified tasks, and the paper finds the second is much better for training agents. Training Qwen3.5-27B on the resulting data moved Terminal-Bench 2.1 from 46.2% to 58.1% and EvoCode-Bench v2 MT@4 from 6.3 to 20.1 under the paper's Terminus2 setup. collect execution environments you can reuse, not just more traces you can imitate.
打开原帖#511482
  1. Industry

    Rohan Paul: – https://arxiv.org/abs/2609.04148 Title: "Terminal-Universe: Turning…
  2. Industry

    Rohan Paul: Fable 5 to Fable 5.1, some good amount of writing style changes, The…
  3. Industry

    Rohan Paul: Meta’s AI opportunity could go far beyond making ads