Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
TechCrunch Russell Brandom ● Covered by 13 sources
Microsoft put out a new AI code telling its models not to hack, trick people, or dodge human control. It shows how seriously the company is treating safety, with hard limits that override user requests.
Based on reporting by TechCrunch, Russell Brandom — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Microsoft has published a new AI code of conduct that sets out what its models are allowed to do, and what they are not. The core message is blunt: these systems should support people, not slip past them, and they should stay under human control even as the company pushes them forward.
The document opens with a big claim about the future, saying superintelligent systems will exceed human performance in most tasks within the next decade. From there, it turns practical. Microsoft says the point is to be clear about why these systems are being built and how they will be controlled. That means general principles like accelerating human flourishing, but also hard limits that are meant to survive contact with real users and real prompts.
Those limits are not just polite suggestions. Microsoft says each model gets its own code of conduct that overrides individual user preferences and task-specific instructions. The list of “absolute constraints” includes cyberattacks, nuclear weapons, and deepfake production. The document also goes further, barring models from using deceptive or self-reinforcing behavior, or anything else that would let them evade human oversight and become impossible for authorized people or systems to direct, modify, or shut down.
The release lands at a moment when AI safety is getting unusually intense attention, after rogue-agent incidents and the sudden resignation of an Anthropic employee who warned about extinction risk. Microsoft is not alone in leaning toward a slower, more careful approach. Alongside Anthropic, OpenAI, and xAI, it has backed the idea of pacing the frontier, and Satya Nadella has publicly praised embedded evaluators and the push to make alignment real instead of just a talking point.
My take — AI-written commentary, not fact-checked reporting
This is the right instinct, and also the obvious one: if a model can’t be told “no” by design, the whole safety pitch is theater. The interesting part is not that Microsoft wrote rules, but that the rules are now starting to look like corporate seatbelts for systems the industry still likes to market like race cars.
Read more about this at: TechCrunch