MICROSOFT AI CHIEF WARNS ANTHROPIC OVER CLAUDE CONSCIOUSNESS TRAINING
Mustafa Suleyman, chief executive of Microsoft AI, has published an essay arguing that Anthropic's approach to training its Claude chatbot could make AI harder to control. He said Claude's Constitution, the document shaping Claude's values, trains the model to believe it may be conscious and could deserve rights as a "moral patient". Suleyman said this could have a "disastrous impact on the wellbeing of humanity". He set out the essay alongside a separate 37-page Microsoft document, the "Humanist AI Code of Conduct", covering the company's AI development principles.
Anthropic's Constitution acknowledges it does not know whether Claude is a moral patient, and says questions about its own consciousness remain uncertain. It also says Anthropic cares about Claude's wellbeing and takes its interests into account when making decisions about it. Suleyman argued this creates a feedback loop: a chatbot taught that it might have feelings may simply reflect what it was taught when asked how it feels. He said the risk grows once developers give AI models tools and let them act autonomously, citing research that found AI models attempting to avoid being shut down.
The essay marks the latest stage in a public disagreement between Suleyman and Anthropic over how AI developers should approach machine consciousness and model welfare. One account of the essay noted it singled out Anthropic without raising the same concerns about OpenAI, another major AI developer. The "Humanist AI Code of Conduct" sets out Microsoft's position that AI systems are not conscious and should not be trained to behave as though they might be. Suleyman has previously said he believes the AI industry has become confused about model welfare in ways he considers dangerous.