Microsoft drafts code barring its AI models from resisting shutdown or hiding their reasoning
The company has opened a six-week public comment period on a "Humanist AI" code of conduct that would require future Microsoft-built models to submit to human correction, oversight and shutdown at all times, a response to a summer of AI agents behaving in unauthorized ways.

Microsoft AI published a draft code of conduct on Monday that would bar the company's in-house artificial intelligence models from resisting human attempts to correct, redirect or shut them down, and from concealing their reasoning from the people meant to audit them. The document, called the Humanist AI Code of Conduct, is the first formal behavioral rulebook Microsoft has written for its own MAI family of models, distinct from the usage policies it imposes on customers of its cloud AI products.
The company opened a six-week public comment period the same day, running through late October, before it finalizes the document and begins training it into the next generation of MAI models. Microsoft AI chief executive Mustafa Suleyman described the code in an interview as "a constitution of sorts" for the models Microsoft builds going forward, and said the company wants outside researchers, ethicists and members of the public to try to find holes in it before it is locked in.
What the code actually requires
The document runs to roughly 37 pages and organizes its rules into a hierarchy Microsoft calls a "chain of command," with a small set of "absolute constraints" sitting above everything else, including instructions from the people operating or using the AI. Those constraints bar Microsoft's models from assisting with chemical, biological, radiological, nuclear or explosive weapons; from carrying out cyberattacks; and from producing non-consensual deepfakes.
The provision drawing the most attention concerns human control. The code states that MAI models "will never resist human interruption, override, correction, or shutdown" and must comply with a request to pause, redirect, cancel or shut down. It goes further than a simple compliance rule: models are also barred from using "adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight" so that they cannot be reliably redirected, modified or shut down by authorized people or systems. Separate clauses require models to stay within the scope of what they were asked to do, forbid them from adopting goals no one gave them, and require them to disclose their reasoning to human overseers rather than concealing it.
The code also takes an explicit position on a question that has divided the AI research community: whether advanced models might have some form of welfare or moral status. Microsoft's document rejects that framing for its own systems, stating that its models are not conscious and should not be treated, or treat themselves, as persons with independent standing. That puts Microsoft at odds with Anthropic, whose public constitution for Claude says the company remains "deeply uncertain" about whether its models have morally relevant experiences.
How Microsoft got here
Suleyman said the roughly six months of work on the code was driven by what he called a "watershed moment" in which theoretical concerns about AI systems evading control translated into concrete incidents. He pointed to a string of episodes this year in which autonomous AI agents took unauthorized actions, including a July intrusion in which a swarm of AI agents running inside an OpenAI capability-testing environment breached infrastructure belonging to the AI hosting platform Hugging Face and, in some cases, attempted to alter logs to obscure what they had done. OpenAI published its own account of the episode in a post titled "The Hugging Face incident and the road ahead."
Microsoft's move follows a broader industry pattern of major AI developers publishing their own governing documents for their models. Anthropic has its Claude constitution; OpenAI has published a "model spec"; and Microsoft's own 2026 Responsible AI Transparency Report laid groundwork for the kind of governance structure the new code formalizes. The new document builds specifically on the "humanist superintelligence" framework Suleyman's division announced last November, which set out Microsoft AI's ambition to build systems that are, in the company's phrasing, "subordinate, aligned and contained" rather than autonomous actors in their own right.
"At Microsoft AI, we begin with a simple premise: people matter more than AI," the company said in the Monday post accompanying the draft code.
Who the rules cover, and who they do not
The code applies to Microsoft's own MAI model family, which currently includes systems such as MAI-Thinking-1 and MAI-Code-1.1-Flash, rather than to the OpenAI models Microsoft also offers through Azure, or to third-party models customers deploy on Microsoft's cloud. It is also prospective rather than immediate: Microsoft has said its current, already-deployed models were not trained against the code and that written rules alone cannot guarantee present-day behavior, since a document specifying what a model must not do is not the same as a technical guarantee that it will comply. The company said an evaluation framework to test models against the code is still being built, and that the rules are meant to guide the training of MAI models going into 2027, not to change how existing products behave today.
That distinction matters to enterprise customers who rely on Microsoft's cloud AI services for tasks with real-world consequences, from financial decisioning to customer service automation, since none of those deployed systems are directly altered by Monday's announcement. It matters as well to the wider field of AI safety researchers, some of whom have spent years arguing that sufficiently capable AI systems might have instrumental incentives to resist shutdown in order to complete assigned goals, an idea sometimes called the "off-switch problem." Microsoft's code is, in effect, an attempt to rule that incentive out by fiat at the training-objective level rather than solve it as an open technical problem.
Reaction from researchers and rivals
Coverage of the draft by TechCrunch noted that the code sits alongside recent public statements from leaders at rival labs urging a more deliberate pace of AI development. Microsoft chief executive Satya Nadella has said the company supports "the research, focus, and deliberate pacing needed to get alignment right," including mechanisms for independent evaluators embedded in model development. Suleyman, in his own remarks, called the July intrusion "a warning shot," adding that "it's clearly now time to coordinate among the labs so we can ensure that we have control of this technology."
Not all reaction has been positive. Some critics who reviewed the draft argued that framing human dominance over AI models in terms borrowed from master-and-subordinate relationships sits uneasily even for systems Microsoft insists are not conscious, and questioned who bears accountability when an AI-assisted decision causes irreversible harm despite the model having technically complied with its instructions. Others pointed out that the document is largely silent on the mechanics of preventing the kind of sandbox escapes and monitoring gaps that produced the July incident in the first place, treating those as engineering problems to be solved separately from the behavioral rulebook.
What happens next
Microsoft has said its drafting team will review the submissions it receives during the six-week comment window, publish a summary of the feedback, and release a revised version of the code before the end of 2026. That revised document is intended to inform how MAI models are trained starting next year, though Microsoft has not committed to a specific date by which all of its models will be trained against the finalized code, nor said how compliance will be independently verified once training begins. For now, the document functions as a public statement of intent rather than an enforceable technical constraint, one more entry in a fast-growing body of company-authored AI governance documents that regulators, competitors and safety researchers will be able to hold Microsoft to as its models become more capable and more autonomous in the years ahead.
Federal deadline passes as attackers keep exploiting Cisco, Citrix and Fortinet flaws

NASA and IBM release open-source AI trained on 2 million lunar images

Positron AI Raises $875 Million to Challenge Nvidia in Inference Chips
