John Schulman@johnschulman201:02Bullish on this direction. Having a metric for explanation quality makes it possible to hillclimb, and counterfactual simulatability seems right. Adam et al. created a dataset+pipeline that creates more diverse+realistic test cases than prior work & do interesting exps on it. Can also train models to write better post-hoc explanations of their behavior, as highlighted in this thread. > Can a model learn to explain its behaviors, such as why it ignored a user request or made a coding error? > > We trained models on thousands of explanations of their own in-the-wild behaviors. > > Training on this single general dataset shows generalization to held-out evals.#728545
MiniMax Design (H3)@Hailuo_AI18:54MiniMax Design / Hailuo (@Hailuo_AI): **Maybe make H3 become the final renderer in your workflow?** Quotes a creator trial using MiniMax H3 like a product-CG renderer: iterate locally, then send to API when direction is set to cut waste.#728545
Qwen Developers@QwenDevs10:33Qwen Developers (@QwenDevs) on **E-Commerce Bench**: > E-Commerce Bench is a small attempt to evaluate models in a specific business setting. hope it can offer a useful reference for improving model performance on specialized business tasks. Quoted @Alibaba_Qwen: agents start with ¥100,000 to run online stores for 365 days (sourcing, negotiation, pricing, promotions, inventory, cash flow) on real e-commerce data.#728545
Jim Fan@DrJimFan10:14Good old days at OpenAI in 2016: an agent stares at screen pixels, moves a mouse, and books a flight on United. We called it World of Bits, inside OpenAI Universe. 10 yrs later, Astra is reincarnated in the same universe. Even the naming is astronomically correct. Universe was perhaps the most ambitious AI infra project at the time, but we couldn't quite figure out how to solve it. A policy with zero prior knowledge of what a "submit" button does has to rediscover the entire internet visual lingua by trial and error. In retrospect, RL from scratch against hand-drawn, per-task "artisan" reward functions on a bunch of Pascal Titan X GPUs was completely doomed. To solve computer use agent, the right way turns out to be boiling the ocean first (hillclimb on every general task you can find), and then specialize back down to the screen pixels and keystrokes. Or simply, a "Specialized Generalist". Lessons learned: one step ahead of everyone, you're a pioneer. Three steps ahead, you're a prophet. Five steps ahead, you're a martyr. Congrats, GPT-6! That United flight finally gets booked, reliably this time. The 2016 intern in me has a big smile.#728545
François Chollet@fchollet04:09Side note: when we released ARC-AGI-3 in March, and frontier models scored <1% on it, a few Singularitarian poasters took it as a personal insult, and got very worked up about it. They argued the benchmark was fundamentally broken, that it could not even be solved by the smartest humans, that the max reachable score was actually 40%, etc. We had to deal with a torrent of insults and hate poasts since because we had released an unsaturated benchmark. As it turns out, the benchmark is perfectly calibrated. It is straightforward for a human to score 100% if they do better than average people – all you need is to use fewer actions than our human baseline (which is not a strong baseline, as we used unfiltered human testers). And naturally as a result it's also very feasible for AI to score 100% once real progress towards agentic general intelligence has been made. The trajectory of AI from <1% to 100% over the course of 6 months shows that the benchmark was able to snapshot the recent rise of agentic capabilities. And that rise has happened faster than most people expected, including us. > Any smart human giving it real effort should score >90% on ARC-AGI-3#728545
François Chollet@fchollet03:42GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation. Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence. Read our post on Astra and what these results mean: https://arcprize.org/blog/astra#728545
Mark Chen@markchen9003:39GPT-6 Astra is here! This is a big moment for our research team - years of work on pretraining, reinforcement learning, and post-training have come together in our most capable and aligned model yet. It can build and test software, work across apps on your computer, and even help you take a crack at open scientific problems! Capabilities that felt like grand challenges a few years ago have become tools people can actually use. One example is Computer Use - if you've tried this before and felt like it was too slow or not good enough, I encourage you to give it another shot. We've come a long way since Operator, and it "just works" now. We're also asking these systems to act on your behalf for more consequential work. Agents needs to stay aligned with your goals and values, think transparently, and respond to oversight even when tasks become difficult. We've made substantial progress on these behaviors in Astra, alongside stronger monitoring that can stop potentially unauthorized actions. That work is part of what makes this release possible. I think alignment is one of the most important research frontiers in AI, and it remains far from solved. Our ability to understand and align models has to keep pace with model capabilities. We want to give people more room to think, build, and discover with increasingly powerful tools that remain *under their control*. Huge thanks to the researchers and teams who got us here. There's a lot more work ahead, and I'm incredibly excited about what we can make possible in the near future!#728545
Percy Liang@percyliang02:56A year ago, David was the only FTE on Marin. Today, thanks to Open Athena, Marin has 10 FTE. Read his post to better understand the context of Marin. And David tokens are always a pleasure to read. > It’s been a few weeks since Marin kicked off our 535B/A23B MoE hero run. So far it’s gone almost boringly well. > > It’s also been just over a year since Marin joined Open Athena. I wrote up some thoughts about how far we’ve come, and how we got there. https://openathena.ai/blog/marin-535b-launch-note/#728545
Noam Brown@polynoamial02:42Of all the use cases for GPT-6 Astra, I'm most excited for scientific discovery. We at @OpenAI have not pushed it to its limits on math and science. I look forward to waking up every morning and seeing what new scientific breakthrough someone has made with this model!#728545
François Chollet@fchollet02:18This is the wrong approach. This kind absolute, overbearing take on AI regulation will prove to be actively counterproductive. > Bernie Sanders and Greg Casar today announced the Ban Artificial Superintelligence Act. All AI development in the United States will be paused. Systems that have capabilities that match or exceed human cognitive performance will be banned. Violators will face 20 years in prison.#728545
François Chollet@fchollet00:54I believe we need to make a deliberate effort to keep humans in the loop in all critical processes across our economy and society, regardless of whether it is technically necessary. Even if AI develops the *capability* for advanced autonomy, we should not make it highly autonomous. We have to maintain control and keep visibility and understanding of all critical processes, we should not blindly hand over everything to AI agents just because we can. AI as a tool in the human hand is the only form of AI that is worth pursuing.#728545
Sara Hooker@sarahookr22:32One of the biggest gaps in ai progress is data. Invent a dataset allows you to translate intent into ai ready data for training. A very important step in bridging the gap. 🔥 Big shoutout to the @adaption_ai team. > Introducing Invent a Dataset. > > Describe the dataset you need. Get a structured, training-ready dataset back. No existing data needed. > > Dataset creation used to start with collection. Now it starts with specification.#728545
MiniMax (official)@MiniMax_AI20:47MiniMax (@MiniMax_AI): **MiniMax-M3 powering HUMAIN-M3**. - Built on the M3 foundation; further trained on **more than 1 trillion Arabic tokens** - Targets Arabic plus diverse languages and regional dialects - Quoted @HUMAIN: HUMAIN-M3 in research preview on HUMAIN Node#728545
Qwen@Alibaba_Qwen12:00Qwen (@Alibaba_Qwen) announced **E-Commerce Bench**, a new benchmark for **long-horizon autonomous business operations**. Verified from the X post: - Agents start with **¥100,000** - Run online stores for **365 days** - Tasks span sourcing, negotiation, pricing, promotions, inventory, and cash flow - Market driven by **real e-commerce data** Primary source: https://x.com/Alibaba_Qwen/status/2095476249556853100#728545
Tencent Hy@TencentHunyuan18:12Tencent Hy (@TencentHunyuan) on **Hy4 preview**: - In AI research, Hy4 preview found inference bottlenecks on its own - Lifted e2e throughput **31.8%** via operator fusion and communications opts - Stable across context lengths and concurrency - More: https://hy.tencent.ai/#728545
Fei-Fei Li@drfeifei12:00Fei-Fei Li (@drfeifei) announced that **World Labs** has reached a major milestone with **Atlas** — described as a first-of-its-kind **multimodal world model trained from scratch**. From the post (verified via X og meta): - Generates frames with **pixel-perfect camera control** - Reconstructs large scenes from as few as **one** image - Framed as a multimodal world-model milestone for @theworldlabs Primary source: the X announcement from @drfeifei.#728545
Kai-Fu Lee@kaifulee12:00Kai-Fu Lee (@kaifulee) announced that the official website for his new book **AI Native: The Mandate to Transform Your Company** is now live. Verified from the X post (og description): - AI is rewriting how companies operate, compete, and create value - Framing: the question is no longer whether to adopt AI, but how quickly to become **AI-native** - Pre-order call-out in the post (ebook/audiobook timing referenced as Sep 15; hardcover Nov 3 in matching public announcement materials) Primary source: https://x.com/kaifulee/status/2094531609450172615#728545
Hailuo@Hailuo_AI12:00Hailuo (@Hailuo_AI) announced **H3 Max** on MiniMax Design. - Faster generation - **480p** starting at about **$0.02/s** - **3 free** generations included#728545
Tencent Hunyuan@TencentHunyuan12:00Tencent Hunyuan (@TencentHunyuan) announced the **Hy4** preview. - **770B** total parameters, **49B** active - **1M** context window - Positioned for **productivity**, with an **open** and **affordable** stance#728545
Z.ai@Zai_org12:00Z.ai (@Zai_org) introduced **GLM-5.3-Flash**. - Native **multimodal** - **1M** context window - **320B-A18B** MoE-style scale - **MIT** license - Previously known as the **Ox Alpha** preview#728545
Qwen@Alibaba_Qwen12:00Qwen (@Alibaba_Qwen) announced **Qwen3.8-Flash** open weights. - Multimodal **MoE** model - Early preview of the **Qwen4** architecture - Scale notes: **125B** / **51B** N-gram configuration highlighted in the post - Open weights release#728545
DeepSeek@deepseek_ai17:17DeepSeek (@deepseek_ai) follow-up (2/n): - Multimodality unlocks more agent use cases - **V4-Flash-Vision-Exp** works across agent frameworks - Combines visual understanding with a wide range of tools for practical workflows#728545
DeepSeek@deepseek_ai12:00DeepSeek (@deepseek_ai) announced that **DeepSeek-V4-Flash-Vision-Exp** is now live on the API. - Experimental multimodal model roughly matched to **V4-Flash** text capability - Aimed at **multimodal agent** workloads - Available via DeepSeek API as a Vision-Exp preview#728545
DeepSeek@deepseek_ai12:00DeepSeek (@deepseek_ai) announced that **DeepSeek Harness v0.1** is available in **Developer Preview**. Verified from the X post: - Opening the harness to developers building agent harnesses worldwide - **Open-sourcing** the codebase under the **MIT** license - Powered by the **Cordis** meta-framework - Core idea highlighted in the post: everything is a plugin Primary source: https://x.com/deepseek_ai/status/2087887408440164663#728545
Kimi.ai@Kimi_Moonshot16:20Kimi.ai (@Kimi_Moonshot): **Kimi K3 is now live on @databricks**. Quoted Databricks: Moonshot AI's open-weight Kimi K3 available through Unity AI Gateway — run where your data lives for custom AI apps and agents on the Lakehouse.#728545
MiniMax Agent@MiniMaxAgent12:00MiniMax Agent (@MiniMaxAgent) launched **MiniMax Code 2.0**. - Coding agent product refresh - Rebuilt on the open-source **Pi Agent** framework#728545
MiniMax@MiniMax_AI12:00MiniMax (@MiniMax_AI) teased upcoming **MiniMax H3**. - **Omni-Reference** capability - Commercial generation support - **Open weights** planned - Available on **HailuoAI.video** and via API#728545
Lilian Weng@lilianweng03:00Lilian Weng (@lilianweng) posted that leaving was “a hard and sad decision,” and that she shared a message with colleagues at **Thinky** (Thinking Machines). Verified post text: > It is a hard and sad decision. I shared this message with folks at Thinky. Thank you all for the time together♥️ Just as the last sentence in my message: The future worth building is human. Neutral summary of the accompanying note (as reported from the image/message she shared): she said she did not feel able to continue at the pace a startup requires, after months of stress/workload that had pushed beyond what her health could sustain. Keep this as a first-person departure signal—no further speculation.#728545
Kimi@Kimi_Moonshot12:00Kimi / Moonshot (@Kimi_Moonshot) announced the release of **Kimi K3** weights and a technical report. - **2.8T** Mixture-of-Experts (MoE) - **Native vision** support - **1M** context window - Open weights + tech report published#728545
Xiaomi MiMo Developers@XiaomiMiMoDevs12:00Xiaomi MiMo Developers (@XiaomiMiMoDevs): **MiMo Responses API is Now Available**. - Compatible with the OpenAI Responses API — Codex users can integrate directly - Streaming, function calling, deep thinking, structured output, and more - API Reference: https://mimo.mi.com/docs/en-US/api/chat/responses - Codex Configuration: https://mimo.mi.com/docs/en-US/tokenplan/integration/codex-configuration#728545
Tianqi Chen@tqchenml12:00Tianqi Chen (@tqchenml) announced a curated **online book** based on a brand-new mini-series taught at **@SCSatCMU** on **Modern GPU Programming for ML Systems**, as part of the ML Systems course. Verified from the X post: - Topics touched in the course include data-layout swizzling, **3D TMA**, and state-of-the-art **Blackwell** programming - Materials released as a curated online book companion to the CMU course Primary source: https://x.com/tqchenml/status/2069382647302734099#728545
Xiaomi MiMo@XiaomiMiMo23:05Xiaomi MiMo (@XiaomiMiMo): **MiMo Claw is now live**. From @XiaomiMiMoDevs launch note: - Flagship AI model + Kingsoft Office integration - Built for reliable long-running agent tasks - Claims 40–60% lower token consumption vs comparable solutions#728545
Xiaomi MiMo@XiaomiMiMo12:00Xiaomi MiMo (@XiaomiMiMo) announced **MiMo-V2.5-Pro-UltraSpeed** in collaboration with **@TileRT_AI**. Verified from the X post: - Claims **1,000+ tokens/s** output speed on a **1 trillion**-parameter model - Framed as a first for that scale/speed combination - Explicitly distinguished from wafer-scale integration (e.g. Cerebras) and pure on-chip SRAM approaches Primary source: https://x.com/XiaomiMiMo/status/2063993790587904362#728545