On September 24, Zhidongxi reported that Simate, a physical-intelligence company founded only three months ago, disclosed that a general-purpose physics-fast-system model it developed with the help of its AI auto-research system AutoResearch has topped the RoboDojo leaderboard, a physical-AI evaluation benchmark, and demonstrated multi-step long-horizon tasks such as brewing tea on real robots: the robot must remember completed steps, track progress, and decide the next action from changing conditions. The company says it has raised hundreds of millions of yuan cumulatively, with a core team from end-to-end autonomous driving, world models, and humanoid-robot control.

[1]

The nut graf: this company is worth a standalone piece not because it is another robotics startup, but because it moves the RSI (recursive self-improvement) narrative into the physical world — AutoResearch lets AI agents carry out the execution work of model modification, training, evaluation, and experiment management, while human researchers keep hypothesis judgment and key decisions. Days after Anthropic disclosed that Claude can already lead about 26 percent of AI research work under human supervision, Simate is trying to replicate that division of labor for robot intelligence. In industry context, another physical-AI company (Xirang Kaiwu) reported financing of hundreds of millions of yuan on the same day, and Simate's own cumulative funding is also in the hundreds of millions — capital is now placing parallel bets on the physical-world version of AI-building-AI. 71 72 The evaluation claim needs to be unpacked. RoboDojo is a third-party benchmark for general robot manipulation, covering generalization, memory, fine manipulation, long-horizon task execution, and open tasks; Simate says its model ranks first on the leaderboard. But this is a self-reported result, with no independent re-verification of the ranking, and the model's exact scale and architecture are held back for a future technical report. The disclosed technical content itself carries information: the physics-fast system introduces a 4D perception-and-memory mechanism — it understands the physical environment by processing depth, geometry, motion, and contact, while organizing historical information around task progress and continuous state tracking, seeking a balance between long-horizon consistency and real-time reaction — and explores efficient visual encoding, spatiotemporal modeling, and feature compression so that larger models can serve real-time on-device operation. That is the fast-system side; by plan it will later coordinate with a general-purpose slow-reasoning system to push toward zero-shot and few-shot general manipulation. 73 74 AutoResearch deserves attention as a separate product line. The disclosed interface is no longer a collection of experiment scripts: there is a From Experiments to One Policy pipeline board and cross-scenario evaluation result pages, used to manage experiment branches, track evaluation performance, and gradually converge multiple rounds of experiments into a more stable policy. Agents can read model code, training configs, and historical experiment records, modify components and configs through the SiPAI pluggable model framework, call infrastructure to run training and evaluation, then decide to continue, adjust, or stop, and fold new findings into the next research round. Simate says AutoResearch is already open for trials, with researchers from Tsinghua, MIT, and HKUST using it. The team composition supports the route: founder Zhang Ying comes from a leading autonomous-driving company, Zhan Fangneng is an HKUST assistant professor working on world models, and 00s-born Ji Mazeyu has participated in real-robot verification of general humanoid controllers, previously co-founding a physical-AI startup later acquired by Meta. 75 76 Caveats and judgment: Simate divides RSI research into weak, medium, and strong tiers — weak RSI targets well-bounded problems, medium RSI is human-AI multi-stage planning, and strong RSI requires the system to autonomously find problems and hold direction, which still needs senior researchers in charge. That taxonomy is itself an honest boundary statement: what can be deployed today is the weak and medium tiers; strong RSI remains a goal. For readers, the metrics worth tracking are not leaderboard positions but three things: whether AutoResearch truly closes the modify-train-evaluate-feedback loop on its own models and improves scores (public evidence beyond self-report); whether the fast-slow system coordination materializes; and whether, when AI does physical research, the gap of distorted feedback between simulation validation and real-robot validation can be closed.

[1]
A robot debug workshop at dusk: in the foreground a workbench with a flowchart-covered blueprint (pure graphics, no text), a pen and a glass of water; midground, a small six-axis robotic arm performing a fine pick-and-place on the table, with a few wooden blocks of different sizes scattered nearby; background, a wall screen showing rows of bar charts (pure graphics) and a row of equipment racks with indicator lights, dusk light slanting in from a window. No people.
The robot learns new tasks, the AI rewrites the next model version, AI-generated illustration, not a news photo