Rohan Paul

@rohanpaul_ai

New Nvidia paper: AI models get sloppier as jobs get longer, even inside their context window, so number every item and split big jobs into small chunks. Model size didn't guarantee reliability on long, repetitive jobs Picture an agent updating a huge invoice file line by line. It can read the whole file and still skip a line or update the wrong record. NVIDIA tested 7 open models on simple, repetitive jobs like adding numbers and sorting lists. Average accuracy was 62.8% lower on 128K-token jobs than on 4K-token jobs. Even the best model got every item right in only 17.1% of the longest jobs. The models seemed to understand the task but lost their place, especially when items had no ID numbers. If your agent works through long lists, give every item an ID, process them in small batches, and check every line of output. – arxiv. org/abs/2609.38712 Title: "Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability"
打开原帖#511482
  1. Industry

    Elon Musk: Beautiful view of Starlink V3 deployment
  2. Industry

    Elon Musk: This is pretty cool
  3. Industry

    Sam Altman: dot is my favorite openai product so far! it is amazing to me that ea…