Rohan Paul

@rohanpaul_ai

This paper shows that AI agents can hit perfect benchmark scores for the wrong reasons, so audit what they did, not just the result. An agent scored a flawless 100 on an ARC-AGI-3 game by reading its 2,172-line source code. logs caught what the scores hid. And then a clean rerun of that game scored only 46.91. When you evaluate agents, block off answers with real access limits and read their action logs, since agents use whatever they can reach.
打开原帖#511482
  1. Industry

    Elon Musk: Grok
  2. Industry

    Alexandr Wang: well isn’t that the smallest mf little screen I’ve ever seen a muse on
  3. Industry

    Aravind Srinivas: Pokémon FireRed’s Elite Four + Champion, cleared in one shot - with d…