On September 27, The Decoder reported that a team including researchers from Fudan University looked back at how people and agents divided work while the team built a model. The material is more than 700 task logs, 56 participants, and the logs of the agents they used. The model is Atria Dawn Preview, a mixture-of-experts system with 744 billion parameters, aimed at research and engineering tasks.

The team says it leads on five of 16 benchmarks, including web search and cybersecurity, and that it does not lead overall. The report names the split: ahead on AutomationBench, CyberGym, and MLE-Bench Lite, behind on GDPval and SWE-Bench Pro. Training ties each task to a real execution environment and checks it against tests, metrics, or source material.

[1]
Embroidery of a hand choosing one red thread beside a long finished line of stitches.
The stitches already run a long way, and the hand is still choosing the next one. It stands for more execution with the decision left to a person, and it is not a lab photograph., AI-generated illustration, not a news photograph

AI was used in 96.5 percent of the tasks reviewed. Over four weeks, the median ratio of agent actions to human inputs rose from 11 to 28.5. The team says not to read that as more autonomy. Each human decision led to more agent steps. That is not the same as the agents making more decisions.

Of 455 completed tasks that used AI, 151 were rated infeasible without AI at the same scope and quality, about a third. Those 151 were spread across 27 of the 56 participants, not a few heavy users. On methods and parameters, the most common pattern was AI proposes and a human selects, at 55.4 percent. Overall, humans made 85.5 percent of the decisions about methods and parameters, and AI made 9.2 percent. Humans made the final decision on goals and scope in 93.4 percent of cases. Even among the 151 tasks rated infeasible without AI, humans chose the goal 95.4 percent of the time.

[1]

Of 588 tasks with a recorded difficulty, 76 percent moved forward because a person intervened, and in 23 percent the agent solved the problem alone. Human help was almost always information: adding context or clarifying requirements, 35.2 percent, or diagnosing the issue and switching methods, 34.7 percent. People rarely did the work. Partial edits were 3.2 percent, and full takeovers 0.7 percent. When an output needed revision, the AI made the changes itself 75.4 percent of the time after human feedback.

The team describes a risk: if every decision sits on a chain of agent work no person can review, people may be left rubber-stamping. Many participants also ran agents in autonomous modes so long runs would not stop for approval. That boundary was drawn for convenience, not as a formal choice about authority. This is the team's study of its own project, not an outside audit. The article also cites other labs. This piece does not treat those figures as measurements from this study.

[1]

要点

  • Atria Dawn Preview is a 744-billion-parameter mixture-of-experts model. It leads 5 of 16 benchmarks and not overall.
  • The median of agent actions per human input rose from 11 to 28.5 in four weeks. The team says that is not more autonomy.
  • Humans made 85.5 percent of method and parameter decisions, and 93.4 percent of final goal decisions.
  • Full takeovers were 0.7 percent of difficult tasks. The study is of the team's own project.