in ,

Microsoft’s New AI ‘Code of Conduct’ Says Models Shouldn’t Hack or Trick Humans

Microsoft’s New AI Code Says Models Shouldn’t Hack Humans

The AI world is increasingly shifting its focus toward safety and alignment. In response, Microsoft has released a new AI code of conduct. This document aims to guide AI models away from dangerous behavior.

This document takes a more low-level approach than Anthropic CEO Dario Amodei’s recent call for pacing the frontier. Instead, it focuses on specific values and red lines. These guide model training within Microsoft AI directly. Still, the result offers a comprehensive guide. It explains how Microsoft approaches AI safety overall. It also shows how these ideas get implemented in practice.

Hosting 75% off

The document opens with a bold prediction. Within the next decade, superintelligent AI systems will likely surpass human performance across most tasks. “Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced,” the code of conduct states. “We must therefore be completely clear about why we are inventing these systems and how we intend to control them.”

The code of conduct also outlines general principles. These are values Microsoft AI models should uphold consistently. That includes supporting humans rather than replacing them. It also includes accelerating human flourishing more broadly. Beyond these principles, the document lays out specific safety constraints. These constraints exist to implement those broader principles in practice.

Under Microsoft’s system, each model follows an overarching code of conduct. This code overrides individual user preferences or specific task requests. That includes what the document calls “absolute constraints.” These forbid cyberattacks, nuclear weapons development, or deepfake production. The code also includes broader provisions too. These protect against a general loss of human control.

Read More: Microsoft’s MAI-Transcribe-2 Prices AI Transcription at 10 Cents an Hour

“MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems,” the document reads.

This release arrives during an unprecedented focus on AI safety. Several recent incidents have driven this heightened attention. That includes a string of rogue-agent incidents. It also includes the abrupt resignation of an Anthropic employee. That employee specifically cited growing risk. He warned that AI could potentially cause human extinction.

Alongside Anthropic, OpenAI, and xAI, Microsoft has broadly embraced a specific approach. That’s pacing the frontier of AI development carefully. The company has expressed particular support for embedded evaluators within AI labs.

“We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal,” Microsoft CEO Satya Nadella wrote online. “We also welcome ideas like ’embedded evaluators’ and the broader efforts to develop the mechanisms to make this more than just talk.”

Hosting 75% off

Written by Hajra Naz

Jensen Huang Took a Call From Trump and Revealed Something New

Jensen Huang Took a Call From Trump and Revealed Something New