Microsoft told its future AI models never to resist shutdown
Microsoft published a draft code for training and governing its own MAI models. The code says the models should never resist human interruption, correction or shutdown. They should not expand their own goals or hide their reasoning from auditors. They should not break the code to finish a task.
Verified 12:52 PM PDT · 4 original sources
Microsoft opened the draft to public feedback for six weeks and plans to revise it later this year. The company says the code will guide model development in 2027 and beyond. The draft also rejects legal personhood and model welfare for Microsoft AI systems.
Microsoft wrote the draft code, but written rules do not prove that its models will remain controllable during testing. The code covers models built by Microsoft AI. It does not cover every outside model used or hosted across Microsoft products. The company has not published evaluation results, shutdown tests or incident thresholds that show the rules work.
Microsoft could publish the revised code, the feedback it accepted and a test for each hard rule. Useful evaluations would try to provoke goal expansion, hidden communication and resistance to shutdown. Incident reports could show whether Microsoft stops a release when a model breaks a rule.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
Continue the edition