Rohan Paul
@rohanpaul_ai
New Apple paper basically says progress in automated ML engineering has come from models and runtimes rather than the scaffolding built around them.
It finds one well-prompted coding agent with shell and file access matched or beat 4 multi-agent ML systems, so skip the orchestration and start with a single session.
Those systems add search trees, memory layers, and specialist agent teams.
They were designed when a model could only write code, not run it.
This paper re-ran every wrapper on 1 codebase with the same model, hardware, and 24-hour budget. Giving the model a shell instead of a chat box was the only change that clearly mattered.
With GLM 5.2, the minimal agent medaled on 62.5% of Kaggle tasks against 47.1% for the best published harness. Adding parallel agents and a message channel actually dropped its medal rate from 55.7% to 33.3%.
– arxiv. org/abs/2609.40303
Title: "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?"