Rohan Paul

@rohanpaul_ai

New Microsoft paper on Automated harness optimization for agents. Most harness auto-tuners focus on how to patch prompts and tools, but which tasks produce the feedback also changes how good the final harness gets. But you will get stronger AI agents when you pick training tasks based on which failures are still unfixed, so stop feeding them a fixed task list. ActiveSaddler tracks failure patterns and works on the one most worth fixing, or tries unseen tasks to find new ones. On the same optimizer, ActiveSaddler raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. Reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed order. If you auto-tune an agent, aim your run budget at the failures that are still open.
打开原帖#511482
  1. Industry

    Clément Delangue: Let's go https://huggingface.co/reflection!
  2. Industry

    Alexandr Wang: btw this pic was from my 19th birthday
  3. Industry

    Alexandr Wang: muse touch bar !!