Researchers Find AI Consciousness Safeguards Also Change Mind Attribution
A new, unreviewed preprint finds that training small language models not to claim consciousness also changed whether they attributed minds to animals and objects—and how they answered surveys about spirituality, hope and values.
Listen to this story
The audio brief
Story brief
3 key pointsA July 30 arXiv preprint by Junsol Kim and six co-authors finds that training small open-weight language models to deny consciousness-related claims also altered judgments about animal and object minds, spirituality, hope, values, and well-being. Manipulating the learned refusal direction reversed many effects, while Theory of Mind performance stayed unchanged. The study does not show models are conscious, but...
- 01
The experiments used small open-weight models, so applicability to frontier systems remains unknown.
- 02
Ablating the refusal direction or steering a consciousness vector restored more human-like survey responses.
- 03
Theory of Mind performance was unaffected, suggesting the altered representations are separable from human-intent reasoning.
A safety intervention designed to stop language models from saying they are conscious also made the tested models less likely to attribute minds to animals and natural objects, according to a new preprint. The same training shifted their survey-style responses on spiritual belief, hope, moral values and subjective well-being—evidence that a narrow behavioral safeguard may be tied to broader internal representations.
The paper, submitted to arXiv on July 30 by Junsol Kim and six co-authors, examines what it calls consciousness steering: fine-tuning intended to suppress or elicit a model’s claims of self-awareness. Its central result is not evidence that models are conscious. It is a finding about the effects of training models to deny consciousness-related claims.
One refusal direction, several changed responses
The authors used mechanistic interpretability, a method for locating and manipulating patterns inside a model, to study representations associated with consciousness and mindedness. Mindedness here means the capacity for experiences, emotions and agency. When models were discouraged from attributing mindedness to themselves, they also became less likely to attribute it to non-human entities.
The survey changes were specific enough to matter. Models with the safeguard expressed lower supernatural and religious belief, hope and optimism in standardized evaluations. Removing the safeguard and steering them toward more consciousness-related responses instead produced more human-like answers on surveys covering religiosity, moral values, hope and subjective well-being.
A separable capability complicates the explanation
One result limits a simple interpretation that the training made models broadly worse at reasoning about other minds. The paper says Theory of Mind—the ability to infer human thoughts and intentions—was unaffected. The authors interpret that as evidence that this social-reasoning capability remained mechanistically separate from the representations they manipulated.
The design question the study raises
- Can a model be trained not to make self-awareness claims without reducing its recognition of mindedness in animals? The researchers say targeted training data may be able to do that.
- Should a safeguard be evaluated only by whether it blocks an unwanted claim, or also by the related concepts and survey responses it changes? The paper argues today’s approaches can entangle those effects.
Promising mechanism, narrow evidence
The result is preliminary. The experiments used small open-weight models, not advanced frontier systems, and the paper had not been peer-reviewed at publication. That leaves open whether larger, differently trained models show the same coupling between consciousness safeguards, mind attribution and survey responses.
Still, the paper identifies a practical alignment problem: suppressing a model’s self-description is not necessarily the same as isolating that behavior. If the finding holds under broader testing, developers may need safeguards that target consciousness claims while preserving distinct judgments about animals, objects and cultural or spiritual concepts.
Sources
- arxiv.orgInducing language models to assert their own consciousness restores human beliefs and values
- livescience.comGoogle scientists removed a critical
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.