Rohan Paul
@rohanpaul_ai
New Google paper shows research agents get better results sooner when they keep a ranked map of ideas and check that code matches each idea, so build both in.
Most agents just keep editing code.
When code drifts from the idea it's scored as, the agent learns the wrong lesson.
AIM sorts ideas into ranked themes and splits each round between strong themes and untested ones. An auditor tosses gamed results and relabels ideas to match the code.
It beat the best prior agent, ScientistOne, by 1.6 and 4.9 points on 2 task groups, and matched ScientistOne's best score up to 3.1x sooner.
Expect the biggest gains on tasks with many possible approaches and few good ones.