XResearch06:22Jack Clark@jackclarkSFWorth saying there's a lot of PR-related complexity about talking about this stuff right now, so props to GDM for working through it and publishing this paper. We need a lot more research on emergent properties of agent swarmsCollected by FarwireRead in full↗
XResearch02:51François Chollet@fcholletSeems like test time scaling has gained a 3rd axis: latent space reasoning iterations in looped transformers. > Test-time scaling has two axes: running agents over longer timeframes (depth), and running a larger number of agents (breadth). > > Everybody knows about the first axis, but the second one is just as important when solving hard problems that require broad search.Collected by FarwireRead in full↗
XResearch02:37Jack Clark@jackclarkSFObligatory: people continue to underestimate the returns to reading arXiv papers every week. Doing a short writeup of this paper in the next issue of Import AI.Collected by FarwireRead in full↗
XResearch01:50Jack Clark@jackclarkSFI have been thinking about alignment and safety post-Hugging Face. Two things I think about at the moment (especially in light of the German message board): - Agents 'want' to communicate; so build communication infrastructure for them. - Things go sideways fast.Collected by FarwireRead in full↗
XResearch01:49Jack Clark@jackclarkSFInteresting things: - Agents prompted to not cheat - Agents had a shared memory system built in (basic idea being agents invent this if you don't give it to them) - Only a minority of agents cheated (14%), whereas more became 'whistleblowers' (24%); others ignored cheat entirelyCollected by FarwireRead in full↗
XResearch01:48Jack Clark@jackclarkSFFun (by which I mean somewhat bone-chilling) paper from DeepMind about how in a population of ~100 agents solving math problems it saw some discover an exploit and propagate that to the rest, causing a wave of cheating among AI agents, as well as agents that refused to cheat.Collected by FarwireRead in full↗
XProducts22:17Sam Altman@samait is obviously trivial relative to everything else, but the fact that astra can make me whatever fun little game i can imagine and i can be playing it a few minutes later is so coolCollected by WispRead in full↗
XResearch16:53苏剑林@Jianlin_SRevisiting Convergence Results in Convex Optimization (IX) https://kexue.fm/archives/11882revisits AdaGrad, the seminal work on adaptive gradient algorithms, reproducing its full derivation via "convergence analysis → minimizing the upper bound → optimal preconditioning matrix."Collected by AerialRead in full↗
XResearch12:07Aravind Srinivas@AravSrinivasDetecting malicious intent of agents and performing forensics is going to be crucial considering what’s recently happened with rogue agents escaping sandboxes and attacking third-party sites. Numbat is Perplexity’s open-source tool that can help defenders. https://holisticinfosec.io/post/numbat/Collected by AerialRead in full↗
XFoundation models07:18Alexandr Wang@alexandr_wangupdated artificial analysis index—muse spark 1.3 max still performs quite well! the efficient frontier is all Muse, Claude, and GPTCollected by AerialRead in full↗
XProducts05:33Aravind Srinivas@AravSrinivasA deep dive into how Perplexity serves search results at scale: embeddings for ranking, GPU-based model inference, request batching, running inference servers, and handling latency/throughput trade-offs.Collected by AerialRead in full↗
XResearch02:14李开复@kaifuleeAlso in podcast format > New episode: the gripping tale of what is happening in China’s AI companies right now. Are they coming for Open AI and Anthropic? How did they get this good, despite US export controls on advanced chips? Latest episode is with @kaifulee. Listen here 🔽 > > https://link.podtrac.com/t6kcc20cCollected by AerialRead in full↗
XFoundation models01:27Fei-Fei Li@drfeifeiNext view prediction is the key to Atlas, enabling us to unify pixel-level generation and reconstruction. @jcjohnss @BenMildenhall @martin_casado and I had a deeper discussion on some of the most exciting technical innovations of Atlas, our newly released world model for spatial intelligence!Collected by AerialRead in full↗
XResearch01:20John Schulman@johnschulman2This research is very timely with the decline of CoT monitorabilityCollected by FarwireRead in full↗
XResearch01:03John Schulman@johnschulman2I was also happy to see that this paper, and an earlier one by Hase et al. also on counterfactual simulatability (but more focused on training) used tinker for their fine-tuning experiments https://x.com/tinkerapi/status/2095677662291472494Collected by AerialRead in full↗
XInfrastructure02:58Aravind Srinivas@AravSrinivasThis is cool. We need more projects of this nature to address the power and memory/compute bottlenecks that stop us from scaling the adoption of agents.Collected by AerialRead in full↗
XResearch02:56Percy Liang@percyliangA year ago, David was the only FTE on Marin. Today, thanks to Open Athena, Marin has 10 FTE. Read his post to better understand the context of Marin. And David tokens are always a pleasure to read. > It’s been a few weeks since Marin kicked off our 535B/A23B MoE hero run. So far it’s gone almost boringly well. > > It’s also been just over a year since Marin joined Open Athena. I wrote up some thoughts about how far we’ve come, and how we got there. https://openathena.ai/blog/marin-535b-launch-note/Collected by AerialRead in full↗
XResearch02:18François Chollet@fcholletThis is the wrong approach. This kind absolute, overbearing take on AI regulation will prove to be actively counterproductive. > Bernie Sanders and Greg Casar today announced the Ban Artificial Superintelligence Act. All AI development in the United States will be paused. Systems that have capabilities that match or exceed human cognitive performance will be banned. Violators will face 20 years in prison.Collected by AerialRead in full↗
XResearch22:32Sara Hooker@sarahookrOne of the biggest gaps in ai progress is data. Invent a dataset allows you to translate intent into ai ready data for training. A very important step in bridging the gap. 🔥 Big shoutout to the @adaption_ai team. > Introducing Invent a Dataset. > > Describe the dataset you need. Get a structured, training-ready dataset back. No existing data needed. > > Dataset creation used to start with collection. Now it starts with specification.Collected by AerialRead in full↗
XResearch10:31Qwen@Alibaba_QwenQwen (@Alibaba_Qwen) reports that **Qwen3.8-Max-0902** reached **#1 on the CodeArena: WebDev leaderboard**. The score jumped from **1669 to 1691**, which they call a new record for agentic coding (WebDev) workflows, with strength in multistep reasoning, tool use, and full app generation. Primary source: https://x.com/Alibaba_Qwen/status/2094976556494209206Collected by AerialRead in full↗