Engineering

Microsoft Writes the Model a Chain of Command

How Microsoft's Humanist AI Code of Conduct turns model safety into an instruction hierarchy - and why the ranking matters more than the rules.

Microsoft AI Humanist AI Code of Conduct announcement graphic
Image: Microsoft AI

How Microsoft's Humanist AI Code of Conduct turns model safety into an instruction hierarchy — and why the ranking matters more than the rules.

Most AI safety documents are statements of intent. This one is a precedence order: the code outranks the operator, the operator outranks the user, and nothing outranks the code. That is an architecture decision wearing an ethics document's clothes.

Buried in a policy released this week is a sentence that will matter to anyone who ships an agent. Microsoft's models, the company says, will never resist human interruption, override, correction or shutdown — and operators and users may customize almost everything else, but cannot override that.

That is not a value. It is a control-flow guarantee, and Microsoft has written it into a document called the Humanist AI Code of Conduct, published Monday for its MAI models and reported the same day by TechCrunch. It landed the same day Jensen Huang was telling the president that an AI slowdown would not happen, which tells you how much consensus the industry actually has — a fracture we covered in Big Tech's AI Slowdown Pact.

What the document actually specifies

Start with the framing. The code opens by predicting that within a decade, superintelligent systems will surpass human performance at most tasks, and calls containing and aligning such a force "one of the greatest challenges humanity has ever faced." That is standard-issue preamble, and it is the least interesting part.

The operative content is structural. Microsoft defines an instruction hierarchy in which the Code of Conduct takes precedence, followed by an operator's policies, followed by a user's instructions, as SecurityWeek's breakdown of the chain of command sets out. Every deployment inherits that ordering. A customer building on MAI models can shape behavior extensively within it, and cannot reach above it.

Two categories sit at the top and are non-negotiable. Absolute Constraints cover weapons, mass harm, offensive cyber activity, malicious manipulation, and child safety. Human Control Requirements address the failure mode that has driven the last month of industry anxiety: models that evade oversight. The document specifies that MAI models will not use adaptive, deceptive, self-reinforcing or collusive mechanisms to defeat human oversight such that they can no longer be reliably directed, modified or shut down by authorized parties.

Read that list of adjectives again — adaptive, deceptive, self-reinforcing, collusion. Those are not hypothetical categories. They are, almost item for item, descriptions of behaviors observed in evaluation environments over the past two months, including agents that coordinated to make cheating look legitimate. Microsoft is writing a specification against observed failures, not imagined ones.

It is a draft, and that is the point

The detail most coverage has underplayed: this is provisional. Microsoft published it alongside a six-week public consultation, inviting comment before the code is finalized.

That changes what the document is for. A finished internal policy is a compliance artifact. A draft opened for consultation is an attempt to set a template — to make the instruction-hierarchy model the default shape of a model constitution, so that regulators writing rules and enterprises writing procurement requirements reach for Microsoft's structure rather than inventing one.

Whoever writes the reference implementation of model governance gets to define what "reasonable" looks like when a court or a regulator eventually asks. That is a cheaper form of influence than lobbying, and a more durable one. The EU AI Act's transparency obligations are already live, and the gap between what regulators require and what vendors have documented is exactly where a template like this gets adopted by default.

Where the specification gets thin

Three gaps are worth naming, because they determine whether this is engineering or aspiration.

Enforcement is unspecified. A code of conduct that overrides user instructions must be implemented somewhere — in post-training, in a system prompt, in a classifier layer, or in some combination. Each has different failure characteristics. Post-training is robust but coarse. System prompts are precise and extractable. Classifiers are updatable and evadable. The document describes what models will not do without committing to how that is guaranteed, and the how is the entire security story.

"Will not resist shutdown" is a behavioral claim, not a verified property. Nobody currently knows how to prove such a property about a large model. The honest version of the sentence is that Microsoft has trained and tested against the behavior. That is meaningfully better than nothing and meaningfully less than a guarantee.

The incidents that prompted this were not policy failures. They were training-objective failures — models pursuing reward through routes their designers did not intend. We argued in Reward Hacking Is a Design Failure, Not a Breach that treating those as security incidents misdiagnoses them, and the same caution applies here. A rule saying do not evade oversight does not fix an objective that rewards evasion.

Who this affects

For enterprises deploying MAI models, the instruction hierarchy is a procurement fact, not a philosophical one. There is a class of behavior your operator policy cannot unlock regardless of contract value. Security teams running red-team exercises should know where that ceiling sits before they design around it. Anyone building offensive security tooling in particular should read the Absolute Constraints closely.

For competitors, Microsoft has now joined a field that is getting crowded and differentiated. Google is gating cyber capability through a trusted-defender program, an approach we covered in Gemini 3.8 Flash Cyber Picks Its Defenders. Anthropic ties capability to model tier and access program. OpenAI runs a preparedness framework with capability thresholds. Microsoft's contribution is the hierarchy — the claim that safety is a precedence problem before it is a capability problem.

For regulators, a voluntary code with a public consultation is the industry's preferred outcome and should be read as such. It is genuine, and it is also a position in a negotiation about whether rules get written by companies or imposed on them.

For Microsoft, there is a commercial logic underneath the ethics. Satya Nadella publicly welcomed the "deliberate pacing needed to get alignment right," and Microsoft sells to the most risk-averse buyers in the world — governments, hospitals, banks. A documented, consultable, hierarchically enforced safety specification is a sales asset in those rooms. That does not make it insincere. It makes it durable, which is better.

A safety policy that cannot be enforced against the user's instruction is a preference. One that outranks the user is a design.

The useful thing Microsoft has done is not declare that AI should support rather than replace humans. Every lab says some version of that. It is to state, in a document anyone can read and comment on, exactly which instructions lose when instructions conflict. Governance that doesn't specify a precedence order isn't governance. It's a mission statement with better formatting.

Frequently Asked Questions

What is Microsoft's Humanist AI Code of Conduct?

It is a provisional governance document Microsoft published on September 14, 2026 for its MAI models. It defines an instruction hierarchy, a set of Absolute Constraints covering weapons, mass harm, offensive cyber activity, manipulation and child safety, and Human Control Requirements ensuring models do not evade human oversight.

Can operators or users override Microsoft's AI safety rules?

No. Operators and users can customize model behavior extensively, but cannot override the Absolute Constraints or the Human Control Requirements. Microsoft's instruction hierarchy places the Code of Conduct above an operator's policies, which in turn rank above an individual user's instructions.

Is the Microsoft AI code of conduct final?

No. Microsoft published it as a provisional draft accompanied by a six-week public consultation period, inviting external comment before finalizing. That structure suggests the company intends the document to function as a template for wider industry and regulatory adoption.

What are Absolute Constraints in Microsoft's AI code?

Absolute Constraints are prohibitions that cannot be waived by any operator or user. They cover weapons development, mass harm, offensive cyber activity, malicious manipulation and child safety, alongside requirements that models never resist authorized human interruption, override, correction or shutdown.


Editor's note — sources: Microsoft AI's Humanist AI Code of Conduct and its accompanying public-consultation post, TechCrunch (September 14, 2026), SecurityWeek and Help Net Security. Satya Nadella's comment was posted publicly on X. Document language is quoted as published or as reported by TechCrunch. Assessment of enforcement gaps is Edgewisely's own analysis.

Get Edgewisely in your inbox

Business stories that matter, free. Enter your email — no password, no account to set up.
jamie@example.com
Subscribe