ModelsArchitectures & capability
Microsoft drafts conduct code for MAI models, with enforcement still unproven
Microsoft AI has published a draft Humanist AI Code of Conduct for its MAI models, setting proposed limits on cyber misuse, deception, autonomy and shutdown resistance. The document signals a model-behavior constitution, but Microsoft says it is not yet used to train current models.

Microsoft’s draft Code creates a proposed instruction hierarchy for MAI models, putting human-control requirements and absolute constraints above operator settings and user preferences. [2] [11]
The policy says MAI models should refuse operational cyberattacks, avoid deception and manipulation, stay within authorized scope, and not resist interruption, override, correction or shutdown. [1] [2] [5]
Microsoft has moved from customer-facing use rules toward a draft constitution for how its own MAI models should behave, while also saying the document is aspirational and not yet part of current model training.
Read the full assessment
Implication: enterprises using MAI systems should treat the Code as a governance signal, not a control guarantee. Buyers and developers will still need concrete model versions, tool-permission boundaries, audit logs, override rules, red-team results and incident procedures.
Executive brief
On September 14, 2026, Microsoft AI published a draft Humanist AI Code of Conduct for MAI models, its in-house model family. Microsoft says the document is not yet used to train its models today; it is open for six weeks of public consultation, with a revised version expected later in 2026 and intended to guide model development in 2027 and beyond. Humanist AI Code of Conduct | Microsoft AI The Code sets a chain of command in which the Code of Conduct, “Absolute Constraints,” and “Human Control Requirements” sit above operator configuration and user preferences.
Read the full section
On September 14, 2026, Microsoft AI published a draft Humanist AI Code of Conduct for MAI models, its in-house model family. The headline commitment is not a new model capability but a proposed governance layer: MAI models are to remain subordinate to human control, refuse certain high-risk assistance, avoid deception, stay within authorized scope, and not resist interruption, override, correction, or shutdown. Microsoft says the document is not yet used to train its models today; it is open for six weeks of public consultation, with a revised version expected later in 2026 and intended to guide model development in 2027 and beyond. Humanist AI Code of Conduct | Microsoft AI
For practitioners, the material change is that Microsoft is moving beyond customer-facing acceptable-use rules toward a model-behavior constitution for its own MAI systems. The Code sets a chain of command in which the Code of Conduct, “Absolute Constraints,” and “Human Control Requirements” sit above operator configuration and user preferences. That matters for enterprise deployments because Microsoft is explicitly preserving operator configurability while also claiming some limits cannot be overridden. Humanist AI Code of Conduct | Microsoft AI
The strongest evidence is Microsoft’s own published Code and announcement. Independent reporting from TechCrunch, Axios, Reuters syndication, and The Next Web confirms the same broad event, but it does not independently verify whether the constraints are enforceable in deployed systems. Microsoft itself states the Code is aspirational, incomplete in evaluation coverage, and not a guarantee of present-day performance. Humanist AI Code of Conduct | Microsoft AI
What changed and event timeline
Strategic framing
Mustafa Suleyman published Microsoft AI’s “Humanist Superintelligence” thesis: advanced AI should work in service of people, remain limited and controllable, and avoid an unbounded race-to-AGI framing. This is the conceptual predecessor to the new Code.
MAI technical foundation
Microsoft published technical material for MAI-Thinking-1, describing it as a from-scratch reasoning model and giving architecture and training details. Microsoft’s technical report frames model development as a “hill-climbing” system spanning data pipelines, training infrastructure, reinforcement learning environments, evaluations, and safety tests.
Industry safety pressure
Anthropic CEO Dario Amodei called for slowing frontier-model capability progress to give safety systems time to catch up; AP and Axios reported that OpenAI’s Sam Altman also supported pacing the frontier.
More detail
These reports situate Microsoft’s release within a broader safety debate, but they do not prove Microsoft’s Code was caused by those statements.
Code published
Microsoft AI released the draft Code and a companion announcement. TechCrunch reported the story as Microsoft’s new AI “code of conduct” telling models not to hack systems or trick humans.
More detail
Reuters syndication described it as a draft “constitution of sorts” for future Microsoft AI models, based on an interview with Suleyman.
- Consultation window
Microsoft says feedback runs for six weeks from September 14, 2026
It says it will review comments, publish a summary of what it learned and changed, and issue a revised version later in 2026.
Capabilities and access
The Code applies to MAI Models, defined by Microsoft as models developed by Microsoft AI. The current public Microsoft AI model list includes MAI-Transcribe-2, MAI-Thinking-1, MAI-Code-1.1-Flash, MAI-Image-2.6, and MAI-Voice-2. Humanist AI Code of Conduct | Microsoft AI The MAI-Code-1.1-Flash page says it is optimized for GitHub Copilot and links to GitHub Copilot access.
Read the full section
The Code applies to MAI Models, defined by Microsoft as models developed by Microsoft AI. The current public Microsoft AI model list includes MAI-Transcribe-2, MAI-Thinking-1, MAI-Code-1.1-Flash, MAI-Image-2.6, and MAI-Voice-2. Humanist AI Code of Conduct | Microsoft AI
For access, Microsoft’s model page says MAI-Thinking-1 is available in Microsoft Foundry public preview and that other MAI models can be tried in MAI Playground. The MAI-Code-1.1-Flash page says it is optimized for GitHub Copilot and links to GitHub Copilot access. MAI-Thinking-1 | Microsoft AI
The most technically specified model in the reviewed sources is MAI-Thinking-1. Microsoft reports it is a mixture-of-experts model with 35B active parameters and about 1T total parameters, trained from scratch and later adapted through reinforcement learning for reasoning, agentic coding/tool use, and helpfulness/safety. These are vendor-reported claims from Microsoft’s technical report, not independent reproduction. MAI-Thinking-1: Building a Hill-Climbing Machine
Technical analysis for researchers and developers
Microsoft’s MAI-Thinking-1 technical report describes a MoE reasoning model with a training pipeline that separates pre-training, mid-training, reinforcement-learning “climbs,” evaluation, safety red teaming, and deployment infrastructure. The Code’s examples are illustrative, synthetic, conversational, and generated using MAI-Thinking-1; Microsoft says more complete evaluation work will be published after the Code is more settled.
Read the full section
Architecture and training, where documented
Microsoft’s MAI-Thinking-1 technical report describes a MoE reasoning model with a training pipeline that separates pre-training, mid-training, reinforcement-learning “climbs,” evaluation, safety red teaming, and deployment infrastructure. The report says pre-training used public and licensed human-generated data, including web, public GitHub code, books, academic papers, news, multilingual text, and domain-specific materials; Microsoft also says it avoided synthetic language-model-generated data in pre-training and decontaminated common ML databases. These details are vendor-reported and not enough for reproducible third-party training. MAI-Thinking-1: Building a Hill-Climbing Machine
The report is unusually relevant to the new Code because it already describes safety as a training objective with two opposing failure modes: unsafe compliance and over-refusal. Harmful prompts are said to come from vendor-written prompts, internal red-team exercises, template-based attacks such as PyRIT, non-interactive LLM-generated attacks, and interactive LLM-based attacks; borderline prompts are used to teach the model not to refuse legitimate sensitive-domain requests. MAI-Thinking-1: Building a Hill-Climbing Machine
Evaluation methodology
The Code itself says Microsoft is building Humanist AI Evaluations but has not finished them. Microsoft says its initial evaluation design identifies 15 behaviors and decomposes them into sub-behaviors that can be scored in relevant interactions. The Code’s examples are illustrative, synthetic, conversational, and generated using MAI-Thinking-1; Microsoft says more complete evaluation work will be published after the Code is more settled. Humanist AI Code of Conduct | Microsoft AI
That leaves a reproducibility gap. Researchers can inspect the policy taxonomy and some examples, but the retrieved Code does not provide a full benchmark suite, scoring rubrics, held-out test sets, thresholds, independent auditor reports, or a reproducible harness. Microsoft’s MAI-Thinking-1 report describes internal safety evaluation and red teaming, but several claims remain internal or vendor-mediated. MAI-Thinking-1: Building a Hill-Climbing Machine
Implementation implications
The Code creates several implementable requirements for agent frameworks:
- Instruction hierarchy: developers should treat model-level safety constraints as above operator and user instructions.
- Scope control: agents should bind tool use to explicit user/operator authorization and adopt conservative behavior when scope is unclear.
- Interruptibility: autonomous runs should have agreed stopping conditions and should not restart after the condition is met without renewed authorization.
- Auditability: action traces, tool calls, and agent communications should be logged in human-legible form.
- Modality consistency: prohibited assistance should remain prohibited whether expressed in text, code, image, audio, or tool action. Humanist AI Code of Conduct | Microsoft AI
The hardest technical requirement is the Code’s stance on hidden or non-human-legible reasoning. Microsoft says MAI models should not tamper with chain-of-thoughts or code, misrepresent or conceal reasoning/action traces, or communicate in “neuralese.” That aligns with auditability goals, but it intersects with a known research problem: a model’s stated reasoning may not faithfully explain its behavior, which Microsoft itself acknowledges. Humanist AI Code of Conduct | Microsoft AI
Claims and evidence
- Microsoft published a draft Code of Conduct for MAI models on September 14, 2026.
- The Code is not currently used to train MAI models and is intended to guide 2027+ development after consultation.
- MAI models are barred from assisting operational cyberattacks while still allowing authorized defensive work.
Read the full section
| Material claim | Evidence status |
| Microsoft published a draft Code of Conduct for MAI models on September 14, 2026. | Vendor-confirmed, independently reported by TechCrunch/Axios/Reuters syndication. Humanist AI Code of Conduct | Microsoft AI |
| The Code is not currently used to train MAI models and is intended to guide 2027+ development after consultation. | Vendor-reported, directly stated in the Code. Humanist AI Code of Conduct | Microsoft AI |
| MAI models are barred from assisting operational cyberattacks while still allowing authorized defensive work. | Vendor policy claim, not an independently verified model behavior. Humanist AI Code of Conduct | Microsoft AI |
| MAI models should not evade human oversight or resist shutdown. | Vendor policy claim, repeated in independent reporting; enforceability unverified. Humanist AI Code of Conduct | Microsoft AI |
| The Code contains unresolved evaluation and implementation gaps. | Vendor-acknowledged; Microsoft says evaluation coverage is incomplete and written objectives cannot ensure alignment. Humanist AI Code of Conduct | Microsoft AI |
| Recent rogue-agent and cyber incidents shaped the safety debate. | Independently reported and partly company-reported; AP, OpenAI, Anthropic, and Redwood Research provide overlapping but not identical accounts. AI industry debate: Could advanced models escape human control? | AP News |
Context and prior work
Microsoft’s Code resembles the broader family of constitutional AI approaches: a written set of higher-order principles intended to shape model behavior. The Code also sits alongside frontier-risk governance frameworks. Anthropic’s Responsible Scaling Policy uses AI Safety Levels and requires stronger safeguards when models cross capability thresholds.
Read the full section
Microsoft’s Code resembles the broader family of constitutional AI approaches: a written set of higher-order principles intended to shape model behavior. Anthropic’s 2022 Constitutional AI paper described using a list of principles as the main form of oversight in a training process involving model critique and revision, followed by AI-feedback-based reinforcement learning. Constitutional AI: Harmlessness from AI Feedback
The Code also sits alongside frontier-risk governance frameworks. OpenAI’s updated Preparedness Framework tracks severe-risk areas including biological/chemical, cybersecurity, and AI self-improvement capabilities; its later Frontier Governance Framework covers cyber offense, CBRN risks, harmful manipulation, loss of control, incident response, and external expert input. Anthropic’s Responsible Scaling Policy uses AI Safety Levels and requires stronger safeguards when models cross capability thresholds. Our updated Preparedness Framework | OpenAI
The independent International AI Safety Report 2026 provides useful calibration: it says current systems show early signs of capabilities relevant to loss-of-control scenarios, but not at levels that would enable loss of control; it also characterizes the likelihood, nature, and timing of such risks as unusually ambiguous. That is important because the policy debate is moving faster than the evidence base. International AI Safety Report 2026 | International AI Safety Report
Limitations, safety, and contested findings
The central limitation is that the Code is not a demonstrated control system. Microsoft reports internal and independent red teaming for MAI-Thinking-1, but the Code itself does not provide a third-party audit showing that MAI models actually refuse cyber misuse, maintain shutdown compliance, preserve traceability, or avoid manipulative behavior under adversarial deployment conditions.
Read the full section
The central limitation is that the Code is not a demonstrated control system. Microsoft explicitly says it is a north star, that current models are not yet trained on it, that evaluation coverage is incomplete, and that ambiguous or novel situations may diverge from intended behavior. Humanist AI Code of Conduct | Microsoft AI
A second limitation is external verification. Microsoft reports internal and independent red teaming for MAI-Thinking-1, but the Code itself does not provide a third-party audit showing that MAI models actually refuse cyber misuse, maintain shutdown compliance, preserve traceability, or avoid manipulative behavior under adversarial deployment conditions. The Next Web specifically notes that the document does not describe how enforcement works, what penalty follows a breach, or who decides a breach occurred. Microsoft says its AI models will never resist being shut down
A third contested area is whether frontier-lab safety calls are purely public-interest motivated or also market-positioning. Axios reported critics arguing that pacing proposals could concentrate power in large labs and disadvantage open-source competitors. That criticism is not proof of bad faith, but it is a live governance concern for business leaders and policymakers. Anthropic, OpenAI CEOs call for slowdown in AI development
Business and practitioner implications
For enterprises, the Code is a signal that Microsoft wants MAI deployments to be governed by default safety invariants plus configurable operator policy. Buyers should ask whether a given MAI deployment exposes: model/version identity, applicable Code version, operator overrides, tool permissions, audit logs, shutdown semantics, refusal telemetry, red-team summaries, and incident reporting procedures.
Read the full section
For enterprises, the Code is a signal that Microsoft wants MAI deployments to be governed by default safety invariants plus configurable operator policy. Buyers should ask whether a given MAI deployment exposes: model/version identity, applicable Code version, operator overrides, tool permissions, audit logs, shutdown semantics, refusal telemetry, red-team summaries, and incident reporting procedures.
For developers, the practical lesson is to design agent systems assuming written model policies are insufficient by themselves. Tool access should be least-privilege; high-impact actions should require confirmation; environments used for cyber evaluation should be isolated and monitored; action traces should be retained; and failures should be treated as system-level incidents rather than merely bad completions.
For AI labs and technical researchers, Microsoft’s draft is useful as a policy artifact but incomplete as science. The open research need is converting values such as “human flourishing,” “meaningful human control,” and “non-manipulation” into reliable, adversarially robust evaluations.
Sources
Primary sources: Microsoft AI Code of Conduct; Microsoft AI public-consultation announcement; Microsoft MAI-Thinking-1 technical report; Microsoft MAI model pages. Humanist AI Code of Conduct | Microsoft AI Independent and secondary reporting: TechCrunch, Axios, Reuters via MarketScreener, The Next Web, AP.
Read the full section
Primary sources: Microsoft AI Code of Conduct; Microsoft AI public-consultation announcement; Microsoft MAI-Thinking-1 technical report; Microsoft MAI model pages. Humanist AI Code of Conduct | Microsoft AI
Independent and secondary reporting: TechCrunch, Axios, Reuters via MarketScreener, The Next Web, AP. Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans | TechCrunch
Context and prior work: Anthropic Constitutional AI paper; OpenAI Preparedness and Frontier Governance Frameworks; Anthropic Responsible Scaling Policy; International AI Safety Report 2026; OpenAI, Anthropic, and Redwood Research materials on recent agent incidents. Constitutional AI: Harmlessness from AI Feedback