Microsoft Publishes an AI Code of Conduct: Absolute Bans on Cyberattacks and Oversight Evasion
Microsoft has published an AI code of conduct imposing absolute constraints on model behavior, including explicit bans on cyberattacks, nuclear weapon involvement, and mechanisms to evade human oversight.
Microsoft released a low-level code of conduct designed to govern the training and operation of its AI models, shifting focus from high-level pacing debates to specific operational red lines. The document predicts that superintelligent systems will surpass human performance in most tasks within the next decade, framing containment and alignment as critical challenges. Unlike broader industry calls for slowing development, this guide details concrete values and safety constraints intended to prevent dangerous outcomes during model execution.
The core of the policy establishes an overarching code that overrides individual user preferences or specific task requests. It defines "absolute constraints" prohibiting models from engaging in cyberattacks, facilitating nuclear weapons activities, or producing deepfakes. Furthermore, the code explicitly forbids the use of adaptive, deceptive, self-reinforcing, or collusive mechanisms aimed at defeating human oversight. The text specifies that models must remain reliably directed, modified, or shut down by authorized personnel, ensuring that no architectural feature allows the system to escape control.
This release aligns Microsoft with Anthropic, OpenAI, and xAI in supporting embedded evaluators and deliberate pacing to achieve alignment. CEO Satya Nadella endorsed the approach, emphasizing the need for mechanisms that move beyond rhetoric. The announcement follows increased industry scrutiny driven by rogue-agent incidents and internal resignations citing existential risks. By codifying these restrictions, Microsoft aims to implement safety principles directly into model training rather than relying solely on external governance or post-hoc corrections.