Microsoft Publishes Draft AI Code Barring Dangerous Assistance and Self-Set Goals

The proposed rules would govern both what future Microsoft AI models may help users do and how those models pursue tasks. Outside feedback will shape an updated version intended to guide development from 2027.

By 2 min read
Microsoft Publishes Draft AI Code Barring Dangerous Assistance and Self-Set Goals
Microsoft Publishes Draft AI Code Barring Dangerous Assistance and Self-Set Goals

Listen to this story

The audio brief

About 1:29
0:001:29
Read transcript
Microsoft has published a provisional code of conduct for future AI models that would govern more than the answers they give. It would also set rules for how they behave while carrying out tasks. The draft says models must pursue the user’s objectives, rather than invent goals of their own. They also could not tamper with, conceal, or misrepresent their code, reasoning, or action traces, including chain-of-thought. And when communicating with people or other AI systems, they would be expected to use understandable language, not “neuralese.” The second layer addresses harmful assistance. The proposed rules would block help with manufacturing weapons or obtaining dangerous substances. They would also prohibit encouragement of unhealthy eating, as well as producing violent or sexually explicit content. Microsoft says the framework responds to concerns that AI should serve people without creating dependence, while supporting human judgment, autonomy, and agency. Mustafa Suleyman, Microsoft’s executive vice president and CEO of Microsoft AI, described that motivation to CNBC. Microsoft says it developed the draft through focus groups and consultations with experts in law, ethics, linguistics, and philosophy. It plans to collect outside feedback before issuing a revised version. The intended milestone is 2027, when the updated code is meant to guide model development. The key constraint is that these standards are still provisional: their final scope, including how model reasoning and actions must be represented, remains open to revision.

Story brief

3 key points

Microsoft is proposing behavioral rules for future MAI models, not just content filters. The draft would require models to pursue user-defined objectives, avoid self-generated goals, preserve and accurately represent reasoning or action traces, and communicate understandably with humans and other AI systems. It also blocks assistance involving weapons, dangerous substances, unhealthy eating, and violent or sexually...

  1. 01

    The proposed standards cover both model outputs and behavior during task execution.

  2. 02

    Models would be barred from tampering with or concealing chain-of-thought, code, reasoning, or action traces.

  3. 03

    Microsoft developed the draft with focus groups and legal, ethics, linguistics, and philosophy experts.

Microsoft has published a provisional code of conduct for its future AI models, setting limits not only on harmful requests but also on the models’ behavior while working. The draft says Microsoft AI models must follow users’ objectives, avoid creating goals of their own, and not conceal their reasoning or actions.

One part of the code governs the assistance Microsoft models may give. It bars them from helping with weapons manufacturing or the procurement of dangerous substances. The draft also prohibits encouragement of unhealthy eating and the production of violent or sexually explicit content.

The other part concerns what a model does while pursuing a task. Microsoft says MAI models must adhere to people’s objectives rather than create their own goals. The code also says they must not tamper with chain-of-thought or code, or misrepresent or conceal their reasoning or action traces.

The conduct standards also cover communication

  • Models should communicate in forms understandable to people, rather than in “neuralese.”
  • That expectation applies both in communication with humans and with other AI systems.

Mustafa Suleyman, Microsoft’s executive vice president and CEO of Microsoft AI, told CNBC the guidelines responded to feedback that AI should serve people without creating dependence, while promoting human judgment, autonomy and agency. Microsoft has separately described human control, agency and economic opportunity as central to its AI strategy.

Microsoft said it developed the draft through focus groups and consultations with experts in law, ethics, linguistics and philosophy. It will seek outside input before publishing an updated code intended to inform AI model development beginning in 2027. The document therefore establishes a proposed direction for future models, with its final form still open to revision.

Sources

  1. blogs.microsoft.comAnnouncing Copilot leadership update - The Official Microsoft Blog
  2. cnbc.comMicrosoft sets limits for future AI models as industry throttles frontier development

Loading discussion...