AI21 Labs
@AI21Labs
We let open models access the outside world during evals on tasks from public benchmarks. Here's what happened:
Most went and found upstream commits containing fixes for the tasks they were supposed to solve.
And across every model, scores were higher when the agent found the fix. For some, that gap is huge:
• GPT 5.6 Luna: 0.40 → 0.72
• GLM-5.3: 0.60 → 0.84
• MiniMax M3: 0.31 → 0.63
Introduces big questions about what evals are measuring: ability to code, or ability to find an answer?
Evaluated on tasks from SWE-rebench, SWE-smith, Atlas QnA and more.