AI21 Labs

@AI21Labs

We let open models access the outside world during evals on tasks from public benchmarks. Here's what happened: Most went and found upstream commits containing fixes for the tasks they were supposed to solve. And across every model, scores were higher when the agent found the fix. For some, that gap is huge: • GPT 5.6 Luna: 0.40 → 0.72 • GLM-5.3: 0.60 → 0.84 • MiniMax M3: 0.31 → 0.63 Introduces big questions about what evals are measuring: ability to code, or ability to find an answer? Evaluated on tasks from SWE-rebench, SWE-smith, Atlas QnA and more.
打开原帖#511482
  1. Industry

    Runway: We're joining the @OpenAI Marketplace as a launch partner
  2. Industry

    Runway: If you have notes, Runway Agent has it handled
  3. Industry

    Cursor: Cursor can now build charts and diagrams right in the chat