Fei-Fei Li, Stanford professor and co-founder of World Labs, said in a Bloomberg television interview on Tuesday (September 22) that safety evaluation of AI cannot be left only to the companies building the systems: internal researchers can measure a single model's capability and safety properties, but that cannot replace broader standards set jointly by government, industry, and academia. "I still believe in the importance of benchmarks set by independent institutions and public-sector actors like academia, and of shared responsibility between industry and government."

[1][2]

The nut graf: Li's remarks land squarely on top of the "AI slowdown" debate — over the past two weeks, an Anthropic CEO has publicly called for pacing the frontier, and two prominent safety researchers resigned to join the evaluation organization METR. Li's position is subtly different: she is not asking for a slowdown; she is asking for third-party grading. Translating the "should we slow down" argument into a "who grades the exam" governance question gives the debate a more actionable focal point.

Li's core argument is about who evaluates. Companies hold the most direct technical detail and operational data, but internal evaluation systems are tightly bound to commercial objectives and can hardly claim neutrality; as model capability keeps rising, leaving safety evaluation entirely to developers both erodes public trust and prevents the formation of cross-company, cross-industry baselines. Her prescription is multi-stakeholder: independent institutions and the public sector set benchmarks, industry and government share responsibility. This does not contradict her role as CEO of World Labs — she is not opposed to companies doing safety research; she is opposed to companies monopolizing the power to set their own exam and grade their own paper. Set against the past few weeks, the position aligns with OpenAI's establishment of an independent mathematics advisory group (hosted by IAS, open to public challenge) and reported stress-testing talks with Anthropic: whether or not the field slows down, the industry is moving from self-evaluation toward external evaluation.

Sourcing caveats: Li's remarks in the interview extend a judgment that AI's systemic risks ultimately depend on human decisions; direct quotes should be taken from media reports. She called for "more oversight" and "independent benchmarks" without proposing specific legislation or regulation, and without naming any company or government for criticism. This piece covers her statement as a public position from an influential researcher on governance frameworks, and makes no evaluation of any country's regulatory policy. What makes it worth attention is the timing: while labs argue over slowing down or speeding up, safety researchers are moving toward independent evaluation bodies, and Li's statement gives the "independent evaluation" middle path its most prominent endorsement.

One more layer worth unpacking is what independent evaluation actually buys in practice. Benchmarks built outside the vendor set a baseline that regulators, enterprise procurement, and insurance underwriters can cite without touching proprietary systems; they also survive the churn of model releases — a third-party suite remains comparable across generations, whereas an internal eval dies with the model that was tuned on it. That is why the debate is not merely academic: every large enterprise buying an agent today is already doing a rough version of this, running its own acceptance tests against vendor claims. Li's argument generalizes that instinct from a single buyer to society as a whole. The open question is feasibility — who funds the institutions, who maintains the suites as models shift from text to long-running agents, and how quickly benchmarks stay ahead of the systems they measure. None of that is settled by a television interview, but the direction is now explicit: expect more of the evaluation infrastructure to be treated as public goods rather than vendor marketing.

[1][2]
Early-morning public hearing hall, an academic figure's back at a lectern, one hand raised pointing to a large screen behind showing an abstract scoring-scale graphic (no text); three rows of silhouetted representatives from different institutions sit below, some writing, uniform morning light from ceiling lamps, fine dust in the air. No text anywhere.
Who grades the exam, AI-generated illustration, not a news photo