According to The Information, OpenAI and Anthropic began discussing with their lawyers earlier this year a legally binding agreement: each would gain API access to the other's commercialized AI models and run a series of stress tests hunting for flaws or potential dangers. Both sides committed not to retain each other's data during testing. Unreleased models are outside the scope.

[1][2]

The nut graf: this is not an ordinary red-teaming contract. It is two frontier labs locked in fierce competition, preparing for the first time to write "find each other's flaws" into a legal document. If the deal lands, OpenAI gets to test Anthropic's commercial models, and vice versa — each company effectively becomes an independent auditor of the other's shipped systems. The symbolic weight exceeds the operational details: after the GPT-6 Astra cybersecurity incident, the self-modification allegations against Anthropic's models, and a week in which both CEOs argued publicly about whether to slow down, competition and oversight have picked the same track.

The design details reward close reading. First, scope is limited to commercialized models, excluding unreleased ones — a floor both sides can stand on (neither wants to hand the opponent a look at the next-generation capability in training) and a common denominator (a broken shipping model damages the whole industry's credibility). Second, "no retaining each other's data" is what makes the deal legally survivable: granting API access to a competitor means exposing the most sensitive technical asset, and without a no-retention clause, neither side could sign. Third, the deal is "being discussed," not done — The Information's reporting says negotiations are near a conclusion, but it remains unclear whether either side has finalized anything.

Place this against the broader background and it completes a missing piece of recent safety narratives. Over the past two weeks, safety mechanisms have been either internal (OpenAI's math advisory group, each lab's own evaluations) or external (Anthropic's $2 billion embedded evaluation deal with Accenture, regulators demanding audits). The mutual-testing agreement is a third kind: horizontal, peer-to-peer oversight — not a regulator watching you, not you watching yourself, but your direct competitor watching you. Elon Musk has argued that competing labs should peer-review each other's security flaws before release; this negotiation turns that principle into contract text. Whether it lands, and whether it becomes more than a symbolic ritual, will depend on whether either side actually discloses findings that embarrass the other when the protocol is executed. That is the test to watch: disclosure is the only part of this deal that costs reputational capital, and the only part that proves the arrangement is real. A mutual-testing regime that buries inconvenient findings would be worse than none, because it would convert a genuine adversarial check into a performative one.

[1][2]
Dusk, two neighboring glass office towers; two people gaze across at each other's office floors through a shared glass curtain wall, a lit computer on each desk, city evening glow outside
Rivals as mutual auditors, AI-generated illustration, not a news photo