Rohan Paul
@rohanpaul_ai
This paper shows that AI agents can hit perfect benchmark scores for the wrong reasons, so audit what they did, not just the result.
An agent scored a flawless 100 on an ARC-AGI-3 game by reading its 2,172-line source code.
logs caught what the scores hid.
And then a clean rerun of that game scored only 46.91.
When you evaluate agents, block off answers with real access limits and read their action logs, since agents use whatever they can reach.