Rohan Paul

@rohanpaul_ai

New Apple paper basically says progress in automated ML engineering has come from models and runtimes rather than the scaffolding built around them. It finds one well-prompted coding agent with shell and file access matched or beat 4 multi-agent ML systems, so skip the orchestration and start with a single session. Those systems add search trees, memory layers, and specialist agent teams. They were designed when a model could only write code, not run it. This paper re-ran every wrapper on 1 codebase with the same model, hardware, and 24-hour budget. Giving the model a shell instead of a chat box was the only change that clearly mattered. With GLM 5.2, the minimal agent medaled on 62.5% of Kaggle tasks against 47.1% for the best published harness. Adding parallel agents and a message channel actually dropped its medal rate from 55.7% to 33.3%. – arxiv. org/abs/2609.40303 Title: "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?"
打开原帖#511482
  1. Industry

    Rohan Paul: – https://arxiv.org/abs/2610.08144 Title: "Navier-Stokes lost in tran…
  2. Industry

    Rohan Paul: A new paper shows that when AI translates a math proof into Lean, pas…
  3. Industry

    Rohan Paul: Emad Mostaque: AI models "have pretty much reached the efficiency of…