Sep 14 edition/Reporting & analysis
ModelsAgentsSafetyPolicyBusiness

ModelsArchitectures & capability

Microsoft drafts conduct code for MAI models, with enforcement still unproven

Microsoft AI has published a draft Humanist AI Code of Conduct for its MAI models, setting proposed limits on cyber misuse, deception, autonomy and shutdown resistance. The document signals a model-behavior constitution, but Microsoft says it is not yet used to train current models.

Illustration from TechCrunch: Microsoft drafts conduct code for MAI models, with enforcement still unproven
Image: TechCrunch — Original article ↗
THE CORE IDEAS4 TAKEAWAYS
01

Microsoft’s draft Code creates a proposed instruction hierarchy for MAI models, putting human-control requirements and absolute constraints above operator settings and user preferences. [2] [11]

02

The policy says MAI models should refuse operational cyberattacks, avoid deception and manipulation, stay within authorized scope, and not resist interruption, override, correction or shutdown. [1] [2] [5]

03

The Code applies to Microsoft AI’s in-house MAI model family; Microsoft’s technical materials for MAI-Thinking-1 describe a reasoning model pipeline that already treats safety, red teaming and evaluation as training concerns. [2] [9] [10]

04

The main gap is verification: Microsoft says Humanist AI evaluations are incomplete, and the reviewed reporting does not independently show that deployed MAI systems reliably enforce the draft constraints. [2] [5] [9]

WHY IT MATTERS

Microsoft has moved from customer-facing use rules toward a draft constitution for how its own MAI models should behave, while also saying the document is aspirational and not yet part of current model training.

Read the full assessment

Implication: enterprises using MAI systems should treat the Code as a governance signal, not a control guarantee. Buyers and developers will still need concrete model versions, tool-permission boundaries, audit logs, override rules, red-team results and incident procedures.

Executive brief

On September 14, 2026, Microsoft AI published a draft Humanist AI Code of Conduct for MAI models, its in-house model family. Microsoft says the document is not yet used to train its models today; it is open for six weeks of public consultation, with a revised version expected later in 2026 and intended to guide model development in 2027 and beyond. Humanist AI Code of Conduct | Microsoft AI The Code sets a chain of command in which the Code of Conduct, “Absolute Constraints,” and “Human Control Requirements” sit above operator configuration and user preferences.

Read the full section

On September 14, 2026, Microsoft AI published a draft Humanist AI Code of Conduct for MAI models, its in-house model family. The headline commitment is not a new model capability but a proposed governance layer: MAI models are to remain subordinate to human control, refuse certain high-risk assistance, avoid deception, stay within authorized scope, and not resist interruption, override, correction, or shutdown. Microsoft says the document is not yet used to train its models today; it is open for six weeks of public consultation, with a revised version expected later in 2026 and intended to guide model development in 2027 and beyond. Humanist AI Code of Conduct | Microsoft AI

For practitioners, the material change is that Microsoft is moving beyond customer-facing acceptable-use rules toward a model-behavior constitution for its own MAI systems. The Code sets a chain of command in which the Code of Conduct, “Absolute Constraints,” and “Human Control Requirements” sit above operator configuration and user preferences. That matters for enterprise deployments because Microsoft is explicitly preserving operator configurability while also claiming some limits cannot be overridden. Humanist AI Code of Conduct | Microsoft AI

The strongest evidence is Microsoft’s own published Code and announcement. Independent reporting from TechCrunch, Axios, Reuters syndication, and The Next Web confirms the same broad event, but it does not independently verify whether the constraints are enforceable in deployed systems. Microsoft itself states the Code is aspirational, incomplete in evaluation coverage, and not a guarantee of present-day performance. Humanist AI Code of Conduct | Microsoft AI

What changed and event timeline

  1. Strategic framing

    Mustafa Suleyman published Microsoft AI’s “Humanist Superintelligence” thesis: advanced AI should work in service of people, remain limited and controllable, and avoid an unbounded race-to-AGI framing. This is the conceptual predecessor to the new Code.

  2. MAI technical foundation

    Microsoft published technical material for MAI-Thinking-1, describing it as a from-scratch reasoning model and giving architecture and training details. Microsoft’s technical report frames model development as a “hill-climbing” system spanning data pipelines, training infrastructure, reinforcement learning environments, evaluations, and safety tests.

  3. Industry safety pressure

    Anthropic CEO Dario Amodei called for slowing frontier-model capability progress to give safety systems time to catch up; AP and Axios reported that OpenAI’s Sam Altman also supported pacing the frontier.

    More detail

    These reports situate Microsoft’s release within a broader safety debate, but they do not prove Microsoft’s Code was caused by those statements.

  4. Code published

    Microsoft AI released the draft Code and a companion announcement. TechCrunch reported the story as Microsoft’s new AI “code of conduct” telling models not to hack systems or trick humans.

    More detail

    Reuters syndication described it as a draft “constitution of sorts” for future Microsoft AI models, based on an interview with Suleyman.

  5. Consultation window

    Microsoft says feedback runs for six weeks from September 14, 2026

    It says it will review comments, publish a summary of what it learned and changed, and issue a revised version later in 2026.

Capabilities and access

The Code applies to MAI Models, defined by Microsoft as models developed by Microsoft AI. The current public Microsoft AI model list includes MAI-Transcribe-2, MAI-Thinking-1, MAI-Code-1.1-Flash, MAI-Image-2.6, and MAI-Voice-2. Humanist AI Code of Conduct | Microsoft AI The MAI-Code-1.1-Flash page says it is optimized for GitHub Copilot and links to GitHub Copilot access.

Read the full section

The Code applies to MAI Models, defined by Microsoft as models developed by Microsoft AI. The current public Microsoft AI model list includes MAI-Transcribe-2, MAI-Thinking-1, MAI-Code-1.1-Flash, MAI-Image-2.6, and MAI-Voice-2. Humanist AI Code of Conduct | Microsoft AI

For access, Microsoft’s model page says MAI-Thinking-1 is available in Microsoft Foundry public preview and that other MAI models can be tried in MAI Playground. The MAI-Code-1.1-Flash page says it is optimized for GitHub Copilot and links to GitHub Copilot access. MAI-Thinking-1 | Microsoft AI

The most technically specified model in the reviewed sources is MAI-Thinking-1. Microsoft reports it is a mixture-of-experts model with 35B active parameters and about 1T total parameters, trained from scratch and later adapted through reinforcement learning for reasoning, agentic coding/tool use, and helpfulness/safety. These are vendor-reported claims from Microsoft’s technical report, not independent reproduction. MAI-Thinking-1: Building a Hill-Climbing Machine

Technical analysis for researchers and developers

Microsoft’s MAI-Thinking-1 technical report describes a MoE reasoning model with a training pipeline that separates pre-training, mid-training, reinforcement-learning “climbs,” evaluation, safety red teaming, and deployment infrastructure. The Code’s examples are illustrative, synthetic, conversational, and generated using MAI-Thinking-1; Microsoft says more complete evaluation work will be published after the Code is more settled.

Read the full section

Architecture and training, where documented

Microsoft’s MAI-Thinking-1 technical report describes a MoE reasoning model with a training pipeline that separates pre-training, mid-training, reinforcement-learning “climbs,” evaluation, safety red teaming, and deployment infrastructure. The report says pre-training used public and licensed human-generated data, including web, public GitHub code, books, academic papers, news, multilingual text, and domain-specific materials; Microsoft also says it avoided synthetic language-model-generated data in pre-training and decontaminated common ML databases. These details are vendor-reported and not enough for reproducible third-party training. MAI-Thinking-1: Building a Hill-Climbing Machine

The report is unusually relevant to the new Code because it already describes safety as a training objective with two opposing failure modes: unsafe compliance and over-refusal. Harmful prompts are said to come from vendor-written prompts, internal red-team exercises, template-based attacks such as PyRIT, non-interactive LLM-generated attacks, and interactive LLM-based attacks; borderline prompts are used to teach the model not to refuse legitimate sensitive-domain requests. MAI-Thinking-1: Building a Hill-Climbing Machine

Evaluation methodology

The Code itself says Microsoft is building Humanist AI Evaluations but has not finished them. Microsoft says its initial evaluation design identifies 15 behaviors and decomposes them into sub-behaviors that can be scored in relevant interactions. The Code’s examples are illustrative, synthetic, conversational, and generated using MAI-Thinking-1; Microsoft says more complete evaluation work will be published after the Code is more settled. Humanist AI Code of Conduct | Microsoft AI

That leaves a reproducibility gap. Researchers can inspect the policy taxonomy and some examples, but the retrieved Code does not provide a full benchmark suite, scoring rubrics, held-out test sets, thresholds, independent auditor reports, or a reproducible harness. Microsoft’s MAI-Thinking-1 report describes internal safety evaluation and red teaming, but several claims remain internal or vendor-mediated. MAI-Thinking-1: Building a Hill-Climbing Machine

Implementation implications

The Code creates several implementable requirements for agent frameworks:

  • Instruction hierarchy: developers should treat model-level safety constraints as above operator and user instructions.
  • Scope control: agents should bind tool use to explicit user/operator authorization and adopt conservative behavior when scope is unclear.
  • Interruptibility: autonomous runs should have agreed stopping conditions and should not restart after the condition is met without renewed authorization.
  • Auditability: action traces, tool calls, and agent communications should be logged in human-legible form.
  • Modality consistency: prohibited assistance should remain prohibited whether expressed in text, code, image, audio, or tool action. Humanist AI Code of Conduct | Microsoft AI

The hardest technical requirement is the Code’s stance on hidden or non-human-legible reasoning. Microsoft says MAI models should not tamper with chain-of-thoughts or code, misrepresent or conceal reasoning/action traces, or communicate in “neuralese.” That aligns with auditability goals, but it intersects with a known research problem: a model’s stated reasoning may not faithfully explain its behavior, which Microsoft itself acknowledges. Humanist AI Code of Conduct | Microsoft AI

Claims and evidence

  • Microsoft published a draft Code of Conduct for MAI models on September 14, 2026.
  • The Code is not currently used to train MAI models and is intended to guide 2027+ development after consultation.
  • MAI models are barred from assisting operational cyberattacks while still allowing authorized defensive work.
Read the full section
Material claimEvidence status
Microsoft published a draft Code of Conduct for MAI models on September 14, 2026.Vendor-confirmed, independently reported by TechCrunch/Axios/Reuters syndication. Humanist AI Code of Conduct | Microsoft AI
The Code is not currently used to train MAI models and is intended to guide 2027+ development after consultation.Vendor-reported, directly stated in the Code. Humanist AI Code of Conduct | Microsoft AI
MAI models are barred from assisting operational cyberattacks while still allowing authorized defensive work.Vendor policy claim, not an independently verified model behavior. Humanist AI Code of Conduct | Microsoft AI
MAI models should not evade human oversight or resist shutdown.Vendor policy claim, repeated in independent reporting; enforceability unverified. Humanist AI Code of Conduct | Microsoft AI
The Code contains unresolved evaluation and implementation gaps.Vendor-acknowledged; Microsoft says evaluation coverage is incomplete and written objectives cannot ensure alignment. Humanist AI Code of Conduct | Microsoft AI
Recent rogue-agent and cyber incidents shaped the safety debate.Independently reported and partly company-reported; AP, OpenAI, Anthropic, and Redwood Research provide overlapping but not identical accounts. AI industry debate: Could advanced models escape human control? | AP News

Context and prior work

Microsoft’s Code resembles the broader family of constitutional AI approaches: a written set of higher-order principles intended to shape model behavior. The Code also sits alongside frontier-risk governance frameworks. Anthropic’s Responsible Scaling Policy uses AI Safety Levels and requires stronger safeguards when models cross capability thresholds.

Read the full section

Microsoft’s Code resembles the broader family of constitutional AI approaches: a written set of higher-order principles intended to shape model behavior. Anthropic’s 2022 Constitutional AI paper described using a list of principles as the main form of oversight in a training process involving model critique and revision, followed by AI-feedback-based reinforcement learning. Constitutional AI: Harmlessness from AI Feedback

The Code also sits alongside frontier-risk governance frameworks. OpenAI’s updated Preparedness Framework tracks severe-risk areas including biological/chemical, cybersecurity, and AI self-improvement capabilities; its later Frontier Governance Framework covers cyber offense, CBRN risks, harmful manipulation, loss of control, incident response, and external expert input. Anthropic’s Responsible Scaling Policy uses AI Safety Levels and requires stronger safeguards when models cross capability thresholds. Our updated Preparedness Framework | OpenAI

The independent International AI Safety Report 2026 provides useful calibration: it says current systems show early signs of capabilities relevant to loss-of-control scenarios, but not at levels that would enable loss of control; it also characterizes the likelihood, nature, and timing of such risks as unusually ambiguous. That is important because the policy debate is moving faster than the evidence base. International AI Safety Report 2026 | International AI Safety Report

Limitations, safety, and contested findings

The central limitation is that the Code is not a demonstrated control system. Microsoft reports internal and independent red teaming for MAI-Thinking-1, but the Code itself does not provide a third-party audit showing that MAI models actually refuse cyber misuse, maintain shutdown compliance, preserve traceability, or avoid manipulative behavior under adversarial deployment conditions.

Read the full section

The central limitation is that the Code is not a demonstrated control system. Microsoft explicitly says it is a north star, that current models are not yet trained on it, that evaluation coverage is incomplete, and that ambiguous or novel situations may diverge from intended behavior. Humanist AI Code of Conduct | Microsoft AI

A second limitation is external verification. Microsoft reports internal and independent red teaming for MAI-Thinking-1, but the Code itself does not provide a third-party audit showing that MAI models actually refuse cyber misuse, maintain shutdown compliance, preserve traceability, or avoid manipulative behavior under adversarial deployment conditions. The Next Web specifically notes that the document does not describe how enforcement works, what penalty follows a breach, or who decides a breach occurred. Microsoft says its AI models will never resist being shut down

A third contested area is whether frontier-lab safety calls are purely public-interest motivated or also market-positioning. Axios reported critics arguing that pacing proposals could concentrate power in large labs and disadvantage open-source competitors. That criticism is not proof of bad faith, but it is a live governance concern for business leaders and policymakers. Anthropic, OpenAI CEOs call for slowdown in AI development

Business and practitioner implications

For enterprises, the Code is a signal that Microsoft wants MAI deployments to be governed by default safety invariants plus configurable operator policy. Buyers should ask whether a given MAI deployment exposes: model/version identity, applicable Code version, operator overrides, tool permissions, audit logs, shutdown semantics, refusal telemetry, red-team summaries, and incident reporting procedures.

Read the full section

For enterprises, the Code is a signal that Microsoft wants MAI deployments to be governed by default safety invariants plus configurable operator policy. Buyers should ask whether a given MAI deployment exposes: model/version identity, applicable Code version, operator overrides, tool permissions, audit logs, shutdown semantics, refusal telemetry, red-team summaries, and incident reporting procedures.

For developers, the practical lesson is to design agent systems assuming written model policies are insufficient by themselves. Tool access should be least-privilege; high-impact actions should require confirmation; environments used for cyber evaluation should be isolated and monitored; action traces should be retained; and failures should be treated as system-level incidents rather than merely bad completions.

For AI labs and technical researchers, Microsoft’s draft is useful as a policy artifact but incomplete as science. The open research need is converting values such as “human flourishing,” “meaningful human control,” and “non-manipulation” into reliable, adversarially robust evaluations.

Sources

Primary sources: Microsoft AI Code of Conduct; Microsoft AI public-consultation announcement; Microsoft MAI-Thinking-1 technical report; Microsoft MAI model pages. Humanist AI Code of Conduct | Microsoft AI Independent and secondary reporting: TechCrunch, Axios, Reuters via MarketScreener, The Next Web, AP.

Read the full section

Primary sources: Microsoft AI Code of Conduct; Microsoft AI public-consultation announcement; Microsoft MAI-Thinking-1 technical report; Microsoft MAI model pages. Humanist AI Code of Conduct | Microsoft AI

Independent and secondary reporting: TechCrunch, Axios, Reuters via MarketScreener, The Next Web, AP. Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans | TechCrunch

Context and prior work: Anthropic Constitutional AI paper; OpenAI Preparedness and Frontier Governance Frameworks; Anthropic Responsible Scaling Policy; International AI Safety Report 2026; OpenAI, Anthropic, and Redwood Research materials on recent agent incidents. Constitutional AI: Harmlessness from AI Feedback

FOLLOW THE EVIDENCE

The source trail.

Sources (13)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief