Rohan Paul
@rohanpaul_ai
Anthropic is hiring Accenture to act like an outside safety inspector, but with unusually deep access inside Anthropic while Claude models are being built and tested.
The work, led by Accenture’s Faculty unit, will cover model evaluations, red-teaming, alignment assessments and safeguard testing.
Both companies expect to invest at least $1B each over 5 years in building AI-safety capacity, while Anthropic will directly pay Accenture for its evaluation work.
Accenture will try to break the models, check whether safety rules actually work, and examine whether Anthropic is following its own safety commitments.
The unusual part is that Anthropic is paying the evaluator itself, so the big unresolved question is how independent Accenture can remain while working so closely with the company it is judging.