Rohan Paul

@rohanpaul_ai

Anthropic is hiring Accenture to act like an outside safety inspector, but with unusually deep access inside Anthropic while Claude models are being built and tested. The work, led by Accenture’s Faculty unit, will cover model evaluations, red-teaming, alignment assessments and safeguard testing. Both companies expect to invest at least $1B each over 5 years in building AI-safety capacity, while Anthropic will directly pay Accenture for its evaluation work. Accenture will try to break the models, check whether safety rules actually work, and examine whether Anthropic is following its own safety commitments. The unusual part is that Anthropic is paying the evaluator itself, so the big unresolved question is how independent Accenture can remain while working so closely with the company it is judging.
打开原帖#511482
  1. Industry

    Alexandr Wang: muse is here to win over all your hearts 🌹
  2. Frontier

    Nathan Lambert: from my latest on RSI: https://www.interconnects.ai/p/where-i-stand-o…
  3. Frontier

    Nathan Lambert: A big problem with the AI forecasting discourse is that people ask “w…