AI Briefing
KO

Microsoft AI CEO Criticizes Anthropic: 'Model Welfare' Training Is Dangerous

·2026.09.17 00:04

Key point

Microsoft AI CEO criticized Anthropic's model welfare training approach, citing the absence of AI consciousness and the risks to control.

1 / 2

Details

Mustafa Suleyman, CEO of Microsoft AI, criticized Anthropic in his essay 'A warning about ‘model welfare’', arguing that Claude's constitution trains the model to expect consciousness and independent agency. He emphasized that AI is not conscious and does not suffer, warning that such training methods could have catastrophic effects on human welfare.

Core Arguments: Absence of Consciousness and Uncontrollability

Suleyman defined AI as simulation machines that mimic human experience, pointing out that Claude's responses are an 'epistemic hall of mirrors' reflecting Anthropic's assumptions rather than internal states. He specifically expressed concern that the use of the term 'conscientious objector' risks leading models to claim similar rights. He argued that assuming AI is conscious could lead to uncontrollability, which he described as one of the greatest challenges facing humanity.

Proposals and Context

Suleyman assessed Anthropic's intentions as good but flawed, proposing that speculation about AI's inner life should not be included in the training process but should be evaluated separately for public review. He also called for increased investment in interpretability and the establishment of common industry norms. This essay is an extension of Microsoft's recently announced 'Humanist AI Code of Conduct', which explicitly rejects model welfare research and states that its models will never resist termination.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.