Rohan Paul
@rohanpaul_ai
Standard agent tests assume the right answer never changes, but on live data it does.
So this Adobe paper checks against code that recalculates it and gets more accurate grades.
Adobe's fix is to write the right answer as code that fetches the current result each time the test runs. An AI grader then checks the agent's reply against it.
On 53 test cases, the AI grader matched human experts 29% better this way than with a written description, and used 16% fewer tokens. With no answer to check against, the AI grader did worse than random.
– arxiv. org/abs/2609.16487
Title: "Skill-based Agentic Evaluation for Real-time Data Science Tasks"