Rohan Paul

@rohanpaul_ai

A new benchmark called JevBench just dropped. for models whose output is a bounded software decision rather than open-ended prose and follows TypeSafe’s 15 Sept release of Jev, which takes application state plus fixed choices and returns a typed answer with probabilities instead of prose. This benchmark's score deliberately combines Intelligence, Calibration, Speed and Cost because deployment can fail even when raw accuracy is high. e.g. GPT-5.6 Luna records substantially higher hard-case accuracy than Jev 1.13.0, yet Jev leads the composite because the benchmark also prices latency, calibration and cost. The geometric mean prevents exceptional performance on 1 axis from fully compensating for a weak one. The result is evidence about a narrow typed-decision workload, not evidence that Jev is generally more capable than GPT-5.6 Luna.
打开原帖#511482
  1. Industry

    Alexandr Wang: muse is here to win over all your hearts 🌹
  2. Frontier

    Nathan Lambert: from my latest on RSI: https://www.interconnects.ai/p/where-i-stand-o…
  3. Frontier

    Nathan Lambert: A big problem with the AI forecasting discourse is that people ask “w…