Microsoft's New AI 'Code of Conduct' Tells Models Not to Hack

As the AI industry pivots toward safety and alignment, Microsoft has released a new AI code of conduct designed to steer its models away from dangerous behavior. The document is more granular than Anthropic CEO Dario Amodei's recent call to pace the frontier, focusing instead on the values and red lines that govern model training within Microsoft AI. Nevertheless, it offers a comprehensive look at how Microsoft approaches AI safety—and how those principles are put into practice. The document opens with a bold prediction: in the next decade, superintelligent AI systems will surpass human performance in most tasks. "Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced," it states. "We must therefore be completely clear about why we are inventing these systems and how we intend to control them." At the heart of the code are general principles that Microsoft AI models must uphold—such as supporting humans rather than replacing them and accelerating human flourishing—alongside specific safety constraints to enforce those principles. Under Microsoft's framework, every model carries an overarching code of conduct that takes precedence over individual user preferences or specific tasks. This includes "absolute constraints" prohibiting cyberattacks, nuclear weapons, and deepfake production, as well as broader provisions against any loss of human control. "MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems," the document reads. The release lands amid unprecedented attention to AI safety, fueled by a string of rogue-agent incidents and the abrupt resignation of an Anthropic employee who cited the growing risk that AI could cause human extinction. Alongside Anthropic, OpenAI, and xAI, Microsoft has broadly embraced the approach of pacing the frontier, with particular support for embedded evaluators in AI labs. "We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal," Microsoft CEO Satya Nadella wrote online. "We also welcome ideas like 'embedded evaluators' and the broader efforts to develop the mechanisms to make this more than just talk."

via TechCrunch AI

Related