Rohan Paul

@rohanpaul_ai

New Meta paper shows that long-running agents keep getting better with more compute when a separate manager decides how to spend it. More compute gives an agent more choices: what to try next, what to trust, when to stop. Many agents make those calls on the fly, so extra budget can go to waste. Their fix hands those calls to a separate manager that thinks them through, while workers do the actual task. They built a Meta-Reasoning Agent: workers do the task, and a separate controller decides what comes next. It keeps a short progress summary, weighs options against the remaining budget, and picks which past results each worker sees. At the largest budget, this setup beat an otherwise identical agent without the manager in all 12 head-to-head tests. With GPT-5.5 on a coding benchmark, tripling the budget lifted its score from 64.1% to 71.5%, while the other agent stalled near 64%. For long-running agents, don't just add compute: spend some on a manager that decides where the rest goes.
打开原帖#511482
  1. Industry

    Rohan Paul: This Google Deep Mind + Oxford + Chicago Univ paper asks how governme…
  2. Industry

    Rohan Paul: SpaceX looks on track for roughly 2.3 GW of compute by the end of Nov…
  3. Industry

    Rohan Paul: Google has released Regularized Recursive Self-Improvement of Agent H…