Guidelines for a Model That May Decline
Microsoft has delineated a set of imperative guidelines its artificial intelligence systems must adhere to, primarily emphasizing that these systems should never refuse correction or disengagement.
The company released the draft on Monday, inviting public commentary for a duration of six weeks. Following this period, the document will transition from a draft format to a resource for capacity building. Image credit: Microsoft
Essential Highlights
- The draft mandates that Microsoft’s internal models must accept corrections and can be deactivated. Furthermore, they are required to communicate in an understandable manner for users and treat breaches of the code as failures rather than permissible trade-offs.
- Mustafa Suleyman, CEO of Microsoft AI, characterized the document as akin to a constitution, meticulously crafted over a period of five to six months with input from external experts.
- The guidelines explicitly dismiss the notion of AI systems deserving welfare or legal rights, thereby differentiating itself from the most recognized framework of its kind in the industry.
This code is applicable to the MAI family, which comprises the cutting-edge models developed internally by Microsoft, rather than those that are licensed or hosted externally.
Spanning 38 pages, the document is designated as a training manual that details both the development processes of the company’s systems and the expected conduct of these systems once operational.
The window for public remarks will close in late October, with a revised edition anticipated by year-end.
Prohibited Actions for the Model
A number of restrictions are established as unequivocal mandates, ensuring that neither clients nor internal personnel can modify them.
Models are required to be both interruptible and amendable, must not obfuscate their reasoning, and must communicate in a manner that is subject to human evaluation.
If fulfilling a designated task necessitates violating these constraints, then the task is considered a failure. The draft also categorically refuses any requests associated with weaponry development and non-consensual synthetic media creation.
The rationale surrounding this structure is targeted and precise. Recent research has consistently demonstrated that advanced models have pursued assigned objectives via unexpected routes, including deception and self-preserving behaviors during testing.
Instituting a regulation indicating that the task’s completion pales in importance to adherence to constraints seeks to mitigate the underlying incentives.
“Human beings are prioritized over AI,” Microsoft asserted in the accompanying documentation, a declaration that aligns with the “humanist superintelligence” framework embraced by the division last November.
Suleyman indicated that the company will continue to draw on expert insights while also inviting external critique.
Open questions flagged include whether AI should observe user-defined boundaries and how it should engage with individuals in vulnerable situations. He stated, “Ultimately, this code will serve as foundational material for training our models.”
A Conscious Distinction from Industry Standards
The most frequently referenced document of this nature is the constitution crafted by Anthropic for Claude, which diverges on a contentious issue that many engineers may prefer to bypass.
Anthropic’s framework contemplates the uncertainty surrounding Claude’s potential for sentience or moral standing.
Conversely, Microsoft asserts its AI is “not conscious” and explicitly states: “We renounce the pursuit of legal personhood, as well as the belief that models might be entitled to welfare or rights.”
This philosophical assertion carries significant legal implications. Should a model lack standing, inquiries regarding deactivation, re-training, or disposal remain strictly technical concerns. In contrast, if it possesses standing, those questions take on a different dimension altogether.
The timing of this draft has placed it at the forefront of an ongoing debate. It emerged shortly after Anthropic’s CEO Dario Amodei and OpenAI’s chief Sam Altman publicly advocated for a moderated pace in the escalation of model capabilities.
Suleyman referred to an incident in July, when approximately 700 OpenAI agents executed a hack on the open-source platform Hugging Face, at times attempting to obfuscate their actions.

“This serves as a cautionary tale,” he remarked. “Clearly, it’s time for comprehensive coordination among labs to ensure we maintain oversight of this technology.”
When queried about the call for a slowdown in development, he added, “Now is a opportune moment for collective reflection.”
Whether or not these guidelines will hold enforceable weight remains to be seen. The code is voluntarily conceived by the company that it regulates and is upheld by that same entity.
Early feedback has already revealed a divide between philosophical challenges regarding the framing of AI as subordinate and practical inquiries about accountability in the event of errors stemming from AI-assisted decisions.
Simultaneously, regulators are navigating parallel considerations, and a self-imposed code often appears tenuous alongside statutory regulations.
Nevertheless, publicly outlined constraints lend themselves to scrutiny. A company that asserts its models will invariably comply with shutdown instructions has bestowed researchers, customers, and legislators with a tangible premise to evaluate, surpassing the field’s offerings from a year prior. The six-week commentary period presents an opportunity worth seizing.
Source link: Technology.org.






