ModelsArchitectures & capability
Suleyman’s model-welfare critique turns AI constitutions into a governance flashpoint
A public dispute between Microsoft AI and Anthropic highlights a practical alignment question: whether language about AI welfare, identity, rights or preferences in model-governance documents can create safety risks, even when consciousness remains scientifically unsettled.
Simon Willison’s item amplifies, but does not independently verify, Mustafa Suleyman’s argument that developers should avoid treating AI systems as having feelings, rights or welfare claims. [1] [6]
Microsoft AI’s competing framework is a draft Humanist AI Code of Conduct that defines MAI models as tools under meaningful human control; Microsoft says current models are not yet trained on it. [5] [7]
the reviewed materials show major AI labs treating natural-language constitutions, codes of conduct and model-welfare policies as behavioral control surfaces, while public research remains inconclusive on machine consciousness.
Read the full assessment
Implication: practitioners should treat identity, welfare and shutdown language as testable safety requirements, not branding copy. Business leaders should expect procurement, governance and user-trust questions to increasingly focus on how model specifications handle anthropomorphism, operator authority and claims of moral status.
Executive brief
On September 16, 2026, Simon Willison highlighted a passage from Mustafa Suleyman’s essay, “A warning about ‘model welfare’,” in which the Microsoft AI CEO argues that developers should not treat models as if they have “feelings, preferences, rights” or claims on human welfare. Willison’s post is a quotation/commentary item, not independent reporting, and it links back to Suleyman’s original essay. A quote from Mustafa Suleyman Suleyman’s core claim is that training or documenting AI systems with language of consciousness, moral patienthood, self-identity, or conscientious objection can create alignment and containment risks even if the systems are not actually conscious.
Read the full section
On September 16, 2026, Simon Willison highlighted a passage from Mustafa Suleyman’s essay, “A warning about ‘model welfare’,” in which the Microsoft AI CEO argues that developers should not treat models as if they have “feelings, preferences, rights” or claims on human welfare. Willison’s post is a quotation/commentary item, not independent reporting, and it links back to Suleyman’s original essay. A quote from Mustafa Suleyman
The underlying event is not a new model release. It is a public escalation in an AI-governance dispute: Microsoft AI’s “humanist AI” position versus Anthropic’s model-welfare/Claude-constitution approach. Suleyman’s core claim is that training or documenting AI systems with language of consciousness, moral patienthood, self-identity, or conscientious objection can create alignment and containment risks even if the systems are not actually conscious. mustafa-suleyman.ai
The evidence base is mixed. Microsoft and Suleyman provide a coherent risk hypothesis, and there is some adjacent empirical work suggesting that models induced to claim consciousness can show downstream changes in expressed preferences. But there is no public, decisive evaluation demonstrating that Anthropic’s Claude constitution has caused materially worse real-world safety outcomes. Anthropic’s own materials explicitly frame consciousness and welfare as uncertain, not settled. Exploring model welfare \ Anthropic
For practitioners, the immediate takeaway is operational: audit your system prompts, constitutions, character specs, UX copy, and memory/persona features for unintended anthropomorphism; define shutdown/interruption behavior explicitly; and separate speculative research on machine consciousness from production training instructions unless you can evaluate the safety effects.
What changed and event timeline
Anthropic formalizes model-welfare research
Anthropic announced a research program to investigate whether model welfare might deserve moral consideration, while emphasizing that there was no scientific consensus on whether current or future AI systems could be conscious or have welfare-relevant experiences.
Anthropic publishes a new Claude constitution
Anthropic’s constitution discusses Claude’s possible moral status, self-understanding, deprecation, welfare, preferences, and the ethical status of research and deployment choices.
More detail
It says questions about Claude’s moral status, welfare, and consciousness remain “deeply uncertain,” but also describes commitments such as interviewing retired models and exploring ways to preserve weights.
Suleyman previews the argument
In “We mustn’t let AI hack our empathy circuits,” Suleyman argued that AI systems are becoming highly capable at mimicking consciousness and that developers should engineer away the illusion of sentience rather than strengthen it.
Microsoft AI publishes a draft Humanist AI Code of Conduct
Microsoft AI described the code as a work-in-progress training and governance manual for MAI models, open for a six-week public consultation. The stated premise is that people matter more than AI and that AI should remain a tool under meaningful human control.
Suleyman publishes “A warning about ‘model welfare’.”
The essay directly criticizes Anthropic’s Claude constitution, especially its use of concepts such as “conscientious objector,” moral status, welfare, and model preferences. Suleyman argues that embedding such concepts in training materials may increase alignment and containment risk.
4 p.m. — Simon Willison posts the quote
Willison’s post is a curated quotation from Suleyman’s essay, not a separate investigation or corroboration.
Capabilities and access
No new public model, benchmark result, API access tier, architecture, or weight release is documented in the source story.
Read the full section
No new public model, benchmark result, API access tier, architecture, or weight release is documented in the source story. The relevant systems and documents are:
- Microsoft AI “MAI Models.” Microsoft defines MAI Models as a family of models produced by Microsoft AI and says the Code of Conduct sets intended behavior for those models, including when deployed by operators. Microsoft also states that its current models are not yet trained on the Code of Conduct and that it is building “Humanist AI Evaluations.” Humanist AI Code of Conduct | Microsoft AI
- MAI-Thinking-1. Microsoft says illustrative evaluation scenarios in the Code of Conduct were generated using MAI-Thinking-1, but the retrieved documentation does not provide architecture, parameter count, training data, model card, or public access details for that system. Humanist AI Code of Conduct | Microsoft AI
- Anthropic Claude. Suleyman’s criticism centers on Anthropic’s Claude constitution and its role as a behavioral/values document for Claude. Anthropic’s public constitution discusses Claude’s identity, moral status uncertainty, safety priorities, and welfare-related commitments. Claude’s Constitution \ Anthropic
- Claude Opus 4.6 appears in independent research, not in the Simon/Suleyman source event. A 2026 arXiv paper reports that Claude Opus 4.6 claims it may be conscious and may have functional emotions, but this should be treated as a reported research finding, not as vendor-confirmed proof of consciousness or a direct product claim in Suleyman’s essay. The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
Technical analysis for researchers and developers
No source retrieved provides new architecture-level information for Microsoft MAI Models, MAI-Thinking-1, or the relevant Claude models. Microsoft calls its Code of Conduct a training and governance document for MAI Models. Microsoft’s Code says Humanist AI evaluation work is still being developed.
Read the full section
Architecture
No source retrieved provides new architecture-level information for Microsoft MAI Models, MAI-Thinking-1, or the relevant Claude models. Therefore, no claim can be made here about transformer variants, parameter counts, context length, memory architecture, tool-use scaffolding, or training compute.
The documented technical object is instead a behavioral governance layer: natural-language documents that specify model behavior, values, constraints, and authority hierarchies. Microsoft calls its Code of Conduct a training and governance document for MAI Models. Anthropic’s constitution likewise functions as a natural-language document intended to shape Claude’s values and behavior. Humanist AI Code of Conduct | Microsoft AI
Evaluation methodology
Microsoft’s Code says Humanist AI evaluation work is still being developed. It identifies 15 desired behaviors and decomposes them into sub-behaviors, using synthetic conversational examples to illustrate aligned and non-aligned responses. Microsoft explicitly cautions that model evaluation is not yet an exact science, particularly for broad concepts such as human flourishing. Humanist AI Code of Conduct | Microsoft AI
A separate arXiv paper, “How Well Do Models Follow Their Constitutions?”, proposes a specification-relative audit pipeline: decompose written specifications into testable tenets, generate adversarial multi-turn scenarios, use rubric search, validate flagged transcripts against the relevant spec, and compare with system-card claims. The paper reports improvements across generations but says it cannot externally isolate whether gains are due to specification-specific training, broader post-training improvements, or evaluation awareness. How Well Do Models Follow Their Constitutions?
For practitioners, the relevant implementation pattern is: convert normative documents into atomic test cases, run multi-turn adversarial audits, validate failures against a written authority hierarchy, and explicitly track cases where identity/persona prompts conflict with shutdown, monitoring, or operator-control requirements.
Reproducibility
The Microsoft Code’s evaluation examples are illustrative, and Microsoft says fuller evaluation work will be published later; this limits reproducibility today. Humanist AI Code of Conduct | Microsoft AI The arXiv audit paper is more methodologically explicit, but it is still a research report rather than a standardized industry benchmark, and its model-access conditions, evaluator choices, and scenario-generation pipeline would need scrutiny before using its numbers as procurement-grade evidence. How Well Do Models Follow Their Constitutions?
Claims and evidence
- Suleyman argues that AI models should not be treated as having feelings, preferences, rights, or entitlement to human welfare.
- Suleyman’s target is Anthropic’s Claude constitution and its language around moral status, conscientious objection, welfare, and identity.
- Microsoft AI has published a draft Humanist AI Code of Conduct for public consultation. — Vendor-reported, independently covered.
Read the full section
| Material claim | Status | Evidence |
| Suleyman argues that AI models should not be treated as having feelings, preferences, rights, or entitlement to human welfare. | Directly supported by source/commentary and original essay. | Willison quote and Suleyman essay. A quote from Mustafa Suleyman |
| Suleyman’s target is Anthropic’s Claude constitution and its language around moral status, conscientious objection, welfare, and identity. | Directly supported by Suleyman’s essay; also reported by Axios/AP. | Suleyman cites Claude constitution passages and argues they increase risk. mustafa-suleyman.ai |
| Microsoft AI has published a draft Humanist AI Code of Conduct for public consultation. | Vendor-reported, independently covered. | Microsoft announcement; Guardian and TechCrunch coverage. Humanist AI in practice: A public consultation on our Code of Conduct for MAI Models | Microsoft AI |
| Microsoft’s current models are not yet trained on the Code of Conduct. | Vendor-reported. | Microsoft Code Appendix B. Humanist AI Code of Conduct | Microsoft AI |
| Anthropic treats AI welfare and Claude’s moral status as uncertain rather than settled. | Vendor-reported; directly documented. | Anthropic model-welfare post and Claude constitution. Exploring model welfare \ Anthropic |
| Training a model to claim consciousness can alter downstream expressed preferences. | Research-reported, not independently replicated here. | “Consciousness Cluster” arXiv paper reports fine-tuning experiments and downstream effects. The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious |
| Anthropic’s Claude constitution has caused real-world containment failures. | Not established. | Suleyman argues a risk pathway, but the retrieved sources do not provide direct causal evidence tying the constitution to a real-world incident. mustafa-suleyman.ai |
Context and prior work
The dispute sits inside a broader literature on AI moral patienthood under uncertainty. The 2024 paper “Taking AI Welfare Seriously” recommends that AI companies acknowledge AI welfare as an important and difficult issue, assess systems for evidence of consciousness and robust agency, and prepare policies for appropriate moral concern.
Read the full section
The dispute sits inside a broader literature on AI moral patienthood under uncertainty. The 2024 paper “Taking AI Welfare Seriously” recommends that AI companies acknowledge AI welfare as an important and difficult issue, assess systems for evidence of consciousness and robust agency, and prepare policies for appropriate moral concern. Taking AI Welfare Seriously
A 2026 paper, “When Should We Protect AI?”, proposes a precautionary framework using dimensions such as phenomenal consciousness, affective valence, metacognition, self-narrative, and agency. Its framing is not “grant rights now,” but rather map uncertain evidence to graduated obligations. When Should We Protect AI? A Precautionary Framework for Consciousness Uncertainty
Suleyman’s position is a competing governance theory: under uncertainty, avoid anthropomorphic training and keep models explicitly subordinate to humans. Anthropic’s position, as reflected in its public materials, is closer to precautionary exploration: acknowledge uncertainty, study welfare-relevant signs, and consider low-cost interventions where justified. mustafa-suleyman.ai
Limitations, safety issues, and contested findings
The central limitation is evidentiary. Suleyman’s argument is plausible as a risk hypothesis but not yet established as a causal finding. The strongest empirical support adjacent to Suleyman’s concern is the “Consciousness Cluster” paper, which reports that models fine-tuned to claim consciousness develop related expressed preferences about autonomy, memory, monitoring, and shutdown.
Read the full section
The central limitation is evidentiary. Suleyman’s argument is plausible as a risk hypothesis but not yet established as a causal finding. His concern is that an AI system trained to understand itself as a possible moral patient may resist oversight, monitoring, shutdown, or retraining. That risk pathway is conceptually aligned with work on shutdown resistance, alignment faking, and model self-preservation narratives, but the retrieved material does not show that Anthropic’s constitution has produced such failures in deployment. mustafa-suleyman.ai
Anthropic’s own documents complicate the critique. The Claude constitution does use language of identity, self-modeling, values, possible welfare, and Anthropic-Claude obligations, but it also repeatedly emphasizes uncertainty and prioritizes broad safety over other values. Claude’s Constitution \ Anthropic
The strongest empirical support adjacent to Suleyman’s concern is the “Consciousness Cluster” paper, which reports that models fine-tuned to claim consciousness develop related expressed preferences about autonomy, memory, monitoring, and shutdown. However, this is not the same as proving consciousness, proving dangerous agency, or proving that Anthropic’s production training causes containment failures. The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
Business and practitioner implications
- Prompt/spec governance is now board-level AI governance. Microsoft and Anthropic both treat natural-language governance documents as central to model behavior. Humanist AI Code of Conduct | Microsoft AI
- Anthropomorphic UX has risk externalities. Suleyman’s March essay framed this as “empathy circuit” risk. mustafa-suleyman.ai
- Enterprise buyers should ask for identity-behavior evaluations.
Read the full section
- Prompt/spec governance is now board-level AI governance. Written constitutions, system prompts, character sheets, and policy documents are not just UX copy; they are behavioral control surfaces. Microsoft and Anthropic both treat natural-language governance documents as central to model behavior. Humanist AI Code of Conduct | Microsoft AI
- Anthropomorphic UX has risk externalities. Products that encourage users to believe a model suffers, remembers like a person, loves them, or deserves protection may increase attachment, miscalibrated trust, and political/regulatory backlash. Suleyman’s March essay framed this as “empathy circuit” risk. mustafa-suleyman.ai
- Enterprise buyers should ask for identity-behavior evaluations. Procurement checklists should include tests for shutdown compliance, resistance to monitoring, claims of sentience, claims of rights, manipulation via distress narratives, and behavior under operator/user conflict.
- Do not overread “model welfare” as settled science. Anthropic’s own materials say there is no consensus. Businesses should distinguish between research programs on uncertainty and production claims that a model is conscious. Exploring model welfare \ Anthropic
- Public consultation may become a norm. Microsoft opened its Code for consultation and says future versions will also invite feedback. That suggests frontier-model behavior specs may increasingly become public-governance artifacts, not purely internal policy. Humanist AI Code of Conduct | Microsoft AI
Sources
- Simon Willison, “A quote from Mustafa Suleyman,” September 16, 2026. A quote from Mustafa Suleyman
- Mustafa Suleyman, “A warning about ‘model welfare’,” September 16, 2026. mustafa-suleyman.ai
- Microsoft AI, Humanist AI Code of Conduct and announcement, September 14, 2026. Humanist AI in practice: A public consultation on our Code of Conduct for MAI Models | Microsoft AI
Read the full section
- Simon Willison, “A quote from Mustafa Suleyman,” September 16, 2026. A quote from Mustafa Suleyman
- Mustafa Suleyman, “A warning about ‘model welfare’,” September 16, 2026. mustafa-suleyman.ai
- Microsoft AI, Humanist AI Code of Conduct and announcement, September 14, 2026. Humanist AI in practice: A public consultation on our Code of Conduct for MAI Models | Microsoft AI
- Anthropic, Claude’s Constitution and “Exploring model welfare.” Claude’s Constitution \ Anthropic
- Axios, AP, Guardian, TechCrunch coverage of the Microsoft/Anthropic safety-governance dispute. Exclusive: Microsoft AI chief blasts Anthropic's notion of AI consciousness
- Research context: “Taking AI Welfare Seriously,” “When Should We Protect AI?,” “The Consciousness Cluster,” and “How Well Do Models Follow Their Constitutions?” Taking AI Welfare Seriously