Rohan Paul

@rohanpaul_ai

New Meta paper finds that having 2 different coding agents review each other's patches catches far more silent bugs than giving 1 agent a bigger budget. Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend. Agents editing real training code can leak test data, break a gradient, or miswire a flag. The code still runs, so you burn GPU hours and get numbers that look valid. The reason is uncorrelated mistakes: agents from the same product fail the same way, so there is little for review to catch. The effect repeated on an unrelated training codebase. – arxiv. org/abs/2609.39551 Title: "RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models"
打开原帖#511482
  1. Industry

    Alexandr Wang: clearing my schedule today to give muse media training 😑
  2. Industry

    Clément Delangue: Super excited about this and I know weights are coming (https://huggi…
  3. Industry

    Aravind Srinivas: Perplexity wins on Hugging Face Decision Index (benchmark for decisio…