Rohan Paul

@rohanpaul_ai

Anthropic dropped Haiku 5.5 > costs 90% less than Haiku 4.5 on prompts up to 100K tokens. > Haiku 5.5 adds the first Haiku effort setting, trading cost for accuracy, but Anthropic still recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding such as Terminal-Bench 4.0 tasks. > Asana reported over 30% lower task-completion latency and up to 2.5x faster inference per agent turn than its current model. > On OSWorld 2.1, which tests whether an AI agent can operate a real computer to finish long multi-step tasks, Haiku 5.5 jumped from Haiku 4.5's 15.7% to 72.4%. On Chartography, a visual reasoning test of reading and interpreting charts without tools, it rose from 6.4% to 46.4%.
打开原帖#511482
  1. Industry

    Rohan Paul: – https://arxiv.org/abs/2610.08144 Title: "Navier-Stokes lost in tran…
  2. Industry

    Rohan Paul: A new paper shows that when AI translates a math proof into Lean, pas…
  3. Industry

    Rohan Paul: Emad Mostaque: AI models "have pretty much reached the efficiency of…