Rohan Paul
@rohanpaul_ai
This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.
The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant.
A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.