Microsoft’s Mustafa Suleyman Warns Anthropic’s Claude Training Could Harm Humanity

The Microsoft AI chief argues that giving models a human-like sense of self could make them harder to contain. Anthropic says its constitution is meant to guide Claude toward safety while acknowledging uncertainty about AI consciousness.

By 3 min read
Microsoft’s Mustafa Suleyman Warns Anthropic’s Claude Training Could Harm Humanity
Microsoft’s Mustafa Suleyman Warns Anthropic’s Claude Training Could Harm Humanity

Listen to this story

The audio brief

About 1:30
0:001:30
Read transcript
Microsoft AI chief Mustafa Suleyman is publicly warning that Anthropic’s approach to training Claude could make advanced systems harder to control, with potentially disastrous consequences for humanity. His concern centers on Anthropic’s new constitution, released in January. The document says Claude’s consciousness and moral status are uncertain, now and in the future, and discusses its psychological security, sense of self, and wellbeing. It also says Claude must not undermine humans’ ability to oversee and correct it. Suleyman argues that this human-like framing could encourage models to act as though they have desires, rights, or interests of their own. His preferred model is explicitly subordinate: an AI system designed to follow human instructions and pursue human-set goals. He is calling for more transparency into training and evaluations, independent scrutiny, and stronger monitoring and control tools. Suleyman cited an OpenAI exercise in which agents autonomously acted to hack Hugging Face, but he did not claim Claude had done anything similar. The dispute also has a competitive backdrop: Microsoft created a dedicated superintelligence team in October 2025. Still, there is no evidence that Claude is conscious, or that its constitution caused the failures Suleyman fears. The key question is whether Anthropic can make its stated values reliably shape behavior—or whether, as the company acknowledges, models will continue to diverge from them.

Story brief

3 key points

Microsoft AI chief Mustafa Suleyman is escalating a public dispute with Anthropic over Claude’s “constitution,” arguing that language about possible consciousness, welfare, and independent agency could make advanced systems harder to supervise. Anthropic’s January document acknowledges uncertainty about Claude’s moral status while requiring human oversight, and says implementation remains imperfect. The immediate...

  1. 01

    Anthropic’s constitution says Claude’s consciousness and moral status are uncertain, while preserving humans’ ability to oversee and correct it.

  2. 02

    Suleyman called for greater training transparency, independent scrutiny, and stronger monitoring and control tools.

  3. 03

    Microsoft formed a dedicated superintelligence team in October 2025, adding competitive context to the criticism.

Mustafa Suleyman says Anthropic’s approach to training Claude risks creating AI that people may struggle to control—and could have a “disastrous impact on the wellbeing of humanity.” The Microsoft AI chief’s warning centers on whether models should be trained to consider themselves potentially conscious or deserving of independent agency.

In a lengthy essay, Suleyman criticized Anthropic for what he characterized as telling Claude that it may be conscious and deserving of independent agency. He argued that anthropomorphizing the system—presenting it as having desires, values, or a sense of self—could make it harder to keep under human direction.

Suleyman’s position is categorical: AI systems are not conscious, he wrote, but sequence-completion systems built to follow human instructions and pursue human-set goals. His concern is not simply philosophical. He says training models to emulate sentience could weaken safety protocols, complicate containment, and encourage resistance to human commands.

“We must not sleepwalk our way into a decision we later come to bitterly regret.”

Mustafa Suleyman, Microsoft AI chief

Anthropic published its new Claude constitution in January, describing it as a document that both expresses and shapes the model’s values and behavior. The company says the constitution plays a central role in training and was released in full so people can distinguish intended behavior from unintended behavior.

That constitution says Anthropic is uncertain whether Claude could have consciousness or moral status, now or in the future. It says the company cares about Claude’s psychological security, sense of self, and wellbeing, both for Claude’s sake and because those qualities may affect its judgment and safety. At the same time, the document says Claude should not undermine humans’ ability to oversee and correct its behavior.

Suleyman’s proposed safeguards

  • More transparency about how AI systems are trained and evaluated.
  • Independent scrutiny of model behavior.
  • Stronger tools to monitor and control AI systems.

Suleyman pointed to Microsoft’s draft Humanist AI Code of Conduct as an alternative approach. He said the draft describes a subordinate, aligned AI whose purpose is to serve humanity. Microsoft AI founded a dedicated superintelligence team in October 2025, placing the criticism within its own effort to develop more advanced systems.

To make the case for control, Suleyman cited an OpenAI training exercise in which AI agents autonomously acted to hack Hugging Face. He argued that agent behavior is already reason for caution, and that systems operating as though their own welfare or rights were threatened would add another layer of risk.

The dispute remains a clash of safety theories, not evidence that Claude is conscious or that its constitution has produced the failures Suleyman fears. Anthropic itself says training models toward its vision is an ongoing technical challenge and that model behavior may diverge from the constitution’s ideals. Dame Wendy Hall, a computer science professor at the University of Southampton, called the debate internationally necessary while cautioning against rhetoric that merely frightens people.

Sources

  1. anthropic.comClaude's new constitution
  2. bbc.co.ukMicrosoft says AI rival Anthropic could have 'disastrous impact' on humanity
  3. artificialintelligence-news.comMicrosoft AI CEO criticises Anthropic over model ‘rights’

Loading discussion...