相关阅读
Frontier
48 Hours, Three Institutional Moves: OpenAI Is Turning Safety From Rhetoric Into InstitutionsFrontier
Measuring a 1pp Error With 38.5% of the Effort: The Economics of Cheaper Agent Evaluation in ProductionFrontier
Let the Rival Be the Auditor: OpenAI and Anthropic's Mutual-Test Deal Writes “Finding Each Other's Flaws” Into a Contract