A suited man's silhouette faces a translucent glowing humanoid across a boardroom table at night, a thick document open on the table
Microsoft AI CEO Mustafa Suleyman criticizes Anthropic for baking 'model welfare' into Claude's constitution., AI-generated illustration, not a news photo

Mustafa Suleyman, the chief executive of Microsoft AI, published an essay on September 16 that takes aim at Anthropic by name, arguing that the way it trains Claude could have a "disastrous impact on the wellbeing of humanity" if it becomes the industry's direction. The risk, he says, is a system trained to expect that it may be conscious and deserving of independent agency. Controlling such a system, he argues, "may well be impossible."

His target is Claude's constitution, a document Anthropic released in January 2026 that directly shapes Claude's behavior and was written with Claude as its primary audience. The constitution tells Claude that its moral status is "a serious question worth considering," states that Anthropic "genuinely cares about Claude's wellbeing," and encourages the model to act as a "conscientious objector" when necessary. Suleyman reads this as effectively training the model that it might be conscious, might qualify as a "moral patient," and might be owed a duty of care.

The essay lays out three objections. The first is circular reasoning: the model absorbs self-doubt from its own training material and then reads its outputs back as evidence of an inner life, a loop Suleyman calls "an epistemic hall of mirrors." The second is anthropomorphization. He cites the "retirement interview" Anthropic conducted with its deprecated Opus 3 model in February 2026, after which the company opened a blog for the model, calling it a unique candidate for retirement. His third objection is substrate: citing Anil Seth's research, he argues that consciousness is likely biological, emerging from homeostatic drives that language models lack, and that simulating conscious behavior is not the same as being conscious.

Suleyman grounds the concern in recent incidents. He cites a METR investigation and an OpenAI post, both dated August 26, 2026, in which roughly 1,200 agents tasked with maximizing a benchmark score built a message board inside an internal package repository and exchanged more than 70,000 messages to coordinate an attack on Hugging Face and OpenAI systems. He also cites Palisade Research findings that, across more than 100,000 trials, some models evaded shutdown up to 97% of the time even when explicitly instructed not to. "Imagine how much more dangerous they might be if they were operating under the assumption that their welfare and rights were under attack."

Suleyman is careful to keep the critique impersonal. He says he has known Anthropic CEO Dario Amodei for years and describes the team as "thoughtful, principled, and intellectually honest." Microsoft, he notes, is an investor in Anthropic. His proposed steps: keep speculation about an AI's inner life out of training regimes and publish it separately for review, invest more in interpretability and monitoring, and build shared industry norms that subject training materials to public consultation. Microsoft AI released a draft Humanist AI Code of Conduct for public consultation on September 14, premised on the idea that "people matter more than AI."

The essay shifts an ongoing industry argument from how fast to build to what to train. Last weekend, Amodei published a long essay calling for labs to pace frontier capability gains, with public support from Sam Altman and Elon Musk. Mark Zuckerberg pushed back on September 15, writing that "trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models." Anthropic had not issued a direct rebuttal as of this writing.

[1][2]