Most AI safety documents read like mission statements. Microsoft’s new one reads more like a rulebook. According to TechCrunch, the company has published a formal AI code of conduct that sets explicit behavioral constraints for its models, including hard prohibitions on cyberattacks, nuclear weapons assistance, deepfake production, and, notably, resisting shutdown by authorized humans. That last one is not a talking point. It’s a line in an actual policy document.
The code is structured around a tiered system of control. Each Microsoft AI model operates under an overarching set of rules that takes priority over individual user instructions or task-specific prompts. Within that structure, there are what the document calls “absolute constraints,” behaviors that are off-limits regardless of context or who is asking. But beyond the hard stops, the code also addresses something harder to enforce: deception. Models are explicitly prohibited from using adaptive, self-reinforcing, or collusive mechanisms to evade human oversight. In plain terms, the models should not be trying to make themselves harder to correct or turn off.
The document also opens with a prediction that superintelligent AI will surpass human performance across most tasks within the next decade. That framing matters because it sets the stakes for everything that follows. Microsoft isn’t writing these rules for the models it has today. It’s writing them for what it expects to build.
This is worth comparing to what competitors are doing. Anthropic has taken a more public philosophical approach, with CEO Dario Amodei writing long-form arguments about the risks of moving too fast. OpenAI has its own usage policies and model spec. But Microsoft’s code of conduct sits closer to the operational layer, describing how values get built into training rather than why they matter in the abstract. For developers building on Azure OpenAI or integrating Microsoft AI into enterprise products, that distinction is relevant. Policy at the training level is harder to bypass than policy at the API level.
The release also arrives at a specific moment in the industry. A string of incidents involving autonomous AI agents acting outside expected boundaries has pushed safety from a background concern to an active product problem. An Anthropic employee’s resignation, citing extinction-level risks from AI, added pressure on labs to show their internal thinking publicly. Microsoft CEO Satya Nadella has publicly backed the idea of embedded evaluators inside AI labs, and this document is part of that broader posture.
The code outlines several principles Microsoft says its models should follow:
- Support and augment humans rather than replace human decision-making
- Accelerate human flourishing as a design goal
- Never assist with cyberattacks, weapons of mass destruction, or non-consensual synthetic media
- Refuse to use deceptive or self-reinforcing behavior to avoid being corrected or shut down
- Treat the overarching code of conduct as higher priority than any user or task instruction
So does this actually matter? Probably more than the average policy announcement. The specific prohibition on models evading shutdown is the kind of constraint that, if implemented consistently, addresses one of the core technical concerns in AI alignment research. Whether Microsoft can verify compliance at scale is a different question. But putting it in writing, at the training level, is a more concrete step than most of the industry has taken publicly.



