Rohan Paul

@rohanpaul_ai

Standard agent tests assume the right answer never changes, but on live data it does. So this Adobe paper checks against code that recalculates it and gets more accurate grades. Adobe's fix is to write the right answer as code that fetches the current result each time the test runs. An AI grader then checks the agent's reply against it. On 53 test cases, the AI grader matched human experts 29% better this way than with a written description, and used 16% fewer tokens. With no answer to check against, the AI grader did worse than random. – arxiv. org/abs/2609.16487 Title: "Skill-based Agentic Evaluation for Real-time Data Science Tasks"
打开原帖#511482
  1. Research

    François Chollet: The critical distinction between base LLMs (2024 and earlier) and mod…
  2. Industry

    Michael Truell: Grok Bot gets useful work done, without you needing to ask.
  3. Industry

    Aravind Srinivas: We're open-sourcing our multimodal Decision model and offering it thr…