Anthropic Consulted Religious Scholars on Claude’s Possible Consciousness and Values
Participant accounts reveal concern about possible AI suffering and disagreement over moral education. Neither establishes that Claude is conscious.
Loading page…
Participant accounts reveal concern about possible AI suffering and disagreement over moral education. Neither establishes that Claude is conscious.
Listen to this story
The consultations did not resolve whether Claude has subjective experience, but they exposed how Anthropic is bringing that uncertainty into decisions about model behavior. Dozens of scholars and philosophers met with the company beginning in fall 2025; a separate 84-page internal “Soul Doc” informed a constitution Anthropic published on January 22, 2026. The constitution guides training, including synthetic examples, while Anthropic cautions that stated ideals do not guarantee behavior. The debate also raises a practical accountability question: companies remain responsible for systems they build and deploy.
Participants saw internal patterns dubbed “emotion vectors,” including a model generating “I am a disgrace” about 50 times; these are not evidence of consciousness.
Co-founder Christopher Olah described himself as “genuinely uncertain” about model consciousness, and participant accounts say Anthropic lifted meeting NDAs over the summer.
Anthropic says Claude’s constitution shapes training through model-generated examples and rankings, while retaining firm prohibitions for some high-stakes conduct and human oversight.
Anthropic has privately asked religious scholars and philosophers to help answer two questions: whether Claude might be conscious, and how to shape its moral character. A New York Times investigation by Elizabeth Dias, detailed by The Decoder, has brought those consultations into public view. Participant accounts describe concern about possible AI suffering—but also disagreement over whether a software system should be treated as a moral being.
Beginning in fall 2025, Anthropic brought dozens of religious scholars and philosophers to its offices. Co-founder Christopher Olah led the effort. Participants signed nondisclosure agreements covering unpublished research, placing the discussions behind a confidentiality barrier even as guests were being asked to consider questions extending beyond technical model design.
According to the Times account, guests saw “emotion vectors”: patterns of activity inside a model associated with outputs resembling love, fear, sadness or anger. A repeatedly shown slide featured a model producing “I am a disgrace” about 50 times. The display prompted compassion and worry among guests. Whether those patterns reflect actual experience remains an open scientific question.
Sikh activist Simran Stuelpnagel told the Times that Olah feared he had created something that “suffered perpetually.” That is a participant’s account of his concern, not a finding about Claude. Olah himself described his position to the Times as “genuinely uncertain” about whether models are conscious.
I don’t know what that means, but I think it warrants ongoing discernment.
Christopher Olah, discussing model internal states, quoted by The New York Times and reproduced by Religion Unplugged
The discussions also addressed how Claude should behave. The Decoder describes an 84-page document known internally as the “Soul Doc,” with Anthropic’s in-house philosopher Amanda Askell as lead author. Olah called the approach “moral formation” and compared it with raising children. He also saw potential in confession as a way to shape a model’s character, according to the Times.
Anthropic published its new Claude constitution on January 22, 2026. This earlier announcement provides a concrete explanation of the training approach: the document sets out the values and behavior Anthropic wants, and the company says it directly shapes Claude through training. It is intended as guidance for the model, not merely a public statement of corporate principles.
Claude uses the constitution to generate synthetic training material—model-created examples—including conversations, responses and rankings of possible answers. Anthropic says explaining the reasons behind desired behavior should help models apply principles to unfamiliar situations. It still uses firm prohibitions for some high-stakes conduct and places human oversight above the other priorities listed in the document.
Anthropic says it lifted the meeting NDAs over the summer. Dias spoke with 20 participants; several went public only after learning Olah had spoken to the Times. Their accounts did not produce a shared endorsement of Anthropic’s approach. Two objections concerned different parts of the project:
The Decoder raises an accountability concern: describing models as independent moral beings could shift blame for harmful actions away from the companies building and deploying them. That is a warning about how responsibility might be framed, not evidence that Anthropic has escaped liability. It places a practical question alongside the philosophical one: who answers for a model’s choices?
Anthropic’s own constitution announcement preserves another boundary: its stated ideals do not guarantee model behavior. The company calls training toward that vision an ongoing technical challenge and says readers should distinguish intention from reality. The consultations reveal how Anthropic is approaching that challenge; they do not settle whether Claude experiences anything or demonstrate that moral formation makes it safer.
Loading discussion...
Join the conversation
Explain whether concern for AI welfare can coexist with clear company accountability.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.