Rohan Paul

@rohanpaul_ai

New Tencent paper shows, if an agent keeps improving, its training tasks cannot stay still: continuously harder environments produced better Terminal-Bench 2.1 results than co-evolution. Synthetic terminal tasks eventually become too easy. Once the agent solves them reliably, they stop giving much useful RL signal. Instead of waiting for the model to fail and then building new tasks around those failures, this paper evolves the tasks themselves. It gradually makes each environment less familiar, adds rarer required skills, or makes the job take more steps. Each harder version is checked and introduced as the agent improves. That produced progressively harder environments when tested with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol. On Terminal-Bench 2.1, Qwen3.6-27B reached 71.5% versus 62.9% with co-evolution, while Qwen3.6-35B-A3B reached 64.9% versus 55.1%.
打开原帖#511482
  1. Industry

    Rohan Paul: Mark Zuckerberg just revealed that Meta is already training post-Watermelon models.
  2. Industry

    Rohan Paul: seeing a strange problem with ChatGPT on browser.
  3. Industry

    Garry Tan: Harness wars are full on now and Muse is very impressive