跳到正文
ByteWoops | AI 观察员
首页
报道原帖
新锐明星
关于
中文EN
搜索登录
ByteWoops | AI 观察员
首页

最新

报道原帖

GitHub 榜

新锐明星
关于
登录English

Rohan Paul

@rohanpaul_ai

Sep 11, 2026, 14:43

Absolutely beautiful new paper from Stanford Univ + Together AI "Intelligence per Watt: Measuring Intelligence Efficiency of Local AI" - It found that hybrid local-cloud routing reduced energy, compute, and cost by 60% to 80% against its batched-cloud baseline. - From 2023 to 2025, local AI’s intelligence-per-watt improved 5.3×, while locally serviceable query coverage jumped from 23.2% to 71.3%. - An iPhone 16 Pro achieved roughly 7× higher intelligence-per-watt than workstation GPUs on the same model and precision, showing how power-efficient mobile AI can be for lightweight queries. - A diverse pool of 20+ local models actually beat the paper’s 3 frontier cloud models on 3 of 4 benchmarks when each query was routed to the best model, showing how much model diversity can matter. - Dropping precision from FP16 to FP4 cut inference energy by 3X–3.5X, while costing roughly 2.5 percentage points of accuracy per precision step; in one test, a larger FP4 model even beat a smaller FP16 model. - The biggest remaining weakness is concentrated at the hard end: on the paper’s hardest reasoning slice, about 95% of problems were still unsolved by local models, even as easier and medium-difficulty tasks improved rapidly.
打开原帖#511482

相关阅读

  1. Industry

    Rohan Paul: Jacob Coxon's next interview on CBS News (ex Anthropic+Open AI resear…Sep 11, 2026
  2. Frontier

    Kirk Borne: Read online the short book "From Python to NumPy" by Nicolas Rougier…Sep 11, 2026
  3. Frontier

    Boris Cherny: The latest Threat Intelligence report is an absolutely terrifying and…Sep 11, 2026
ByteWoops | AI 观察员

独立、严谨、可追溯的 AI 前沿资讯。

了解更多报道原帖GitHub 榜关于投稿技能
用户中心登录用户中心Agent API
每周研究简报

人工智能研究、系统与社会

RSS
© 2026 ByteWoops隐私