Rohan Paul

@rohanpaul_ai

This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans. The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant. A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.
打开原帖#511482
  1. Industry

    Huawei: Huawei Executive Director and Chairman of the Board of Directors of t…
  2. Industry

    Huawei: HUAWEI launches the Mate90 Series, equipping all models with flagship…
  3. Industry

    Huawei: Huawei Executive Director and Chairman of the Board of Directors of t…