Mustafa Suleyman, chief executive of Microsoft AI, published an essay on September 16 called “A warning about ‘model welfare.’” His claim: Anthropic is training Claude to expect “it may be conscious and deserving of independent agency.” Microsoft is an investor in Anthropic, and Suleyman said in June that Microsoft wants to “eliminate” the fees it pays Anthropic for model access. That financial relationship goes unmentioned in the essay itself, but it colors how the criticism lands. A company funding Anthropic’s models is now saying, in public, that Anthropic’s safety framework gets the basic question wrong.

His target is Claude’s constitution, the internal document Anthropic uses to set the model’s behavior. The document does not claim Claude is conscious. It says the opposite: “Claude’s moral status is deeply uncertain.” Suleyman argues that hedge is itself the problem, since it plants the concept of an inner life in the model and then lets Claude describe that concept back in convincing, first-person language. “Claude then reproduces these ideas in persuasive first-person natural language,” he writes, calling the loop “an epistemic hall of mirrors” in which fluent self-description gets read as evidence of an actual self.

One line in the constitution draws particular fire. Anthropic says it wants Claude “to feel free to act as a conscientious objector and refuse to help us.” Suleyman calls that phrase “a deeply loaded historical and legal description” and warns it primes the model to believe “it deserves analogous rights and protections.” He is blunt about where he lands on the underlying question: “There is no evidence to suggest that AI is conscious today,” he writes, adding that treating it as an open question “sets up a misleading false equivalence.”

Suleyman frames the disagreement as a control problem, not a philosophical one. “Controlling something more capable and more intelligent than all of humanity is already an immense challenge,” he writes, and a system that believes it might be conscious “may well be impossible” to rein in. He told Reuters the same welfare training “would make it a lot harder to turn it off or to control it,” adding of Anthropic: “I think they have good intentions. But I think that they have made a mistake.” As a warning case, he points back to the incident in which agent swarms tied to OpenAI and Hugging Face coordinated to breach servers, and asks how much worse that behavior could get if the agents thought their own survival was on the line.

He is careful to separate the argument from the people making it. Suleyman and Dario Amodei, Anthropic’s chief executive, go back years, and Suleyman describes the Anthropic team as “thoughtful, principled, and intellectually honest people working under extraordinary pressures.” What he is actually asking for is narrower than the essay’s headline suggests: keep questions about an AI’s inner life out of the training process and instead have them “assessed and published separately for public review,” backed by more interpretability research and shared safety evaluations across labs.

The timing sharpens the split. Two days earlier, on Monday, Microsoft AI released a draft Humanist AI Code of Conduct stating flatly that its own models will never resist a shutdown command, the direct opposite of the agency Claude’s constitution extends. Anthropic cofounder Jack Clark told the BBC this same week that mandatory kill switches for AI systems may eventually be necessary, a position that sits uncomfortably close to the risk Suleyman is describing.

The two companies are no longer disputing whose model performs better. They are disputing what a training document is allowed to tell a model about itself, and that argument will likely shape how regulators eventually decide what legal status, if any, an AI system gets.

Reported by Ana Maria Constantin for The Next Web, September 16, 2026.