Researchers Find AI Consciousness Safeguards Also Change Mind Attribution

A new, unreviewed preprint finds that training small language models not to claim consciousness also changed whether they attributed minds to animals and objects—and how they answered surveys about spirituality, hope and values.

By 3 min read
Researchers Find AI Consciousness Safeguards Also Change Mind Attribution
Researchers Find AI Consciousness Safeguards Also Change Mind Attribution

Listen to this story

The audio brief

About 1:26
0:001:26
Read transcript
Training language models not to say they are conscious may also change how they judge the minds of animals and natural objects. That is the finding of a new, unreviewed preprint submitted to arXiv on July 30 by Junsol Kim and six co-authors. The study examined small, open-weight models trained to suppress claims of self-awareness—a technique the researchers call consciousness steering. The models did not just become less likely to attribute minds to themselves. They also became less likely to attribute mindedness—the capacity for experience, emotion, or agency—to non-human entities. Their answers shifted on standardized surveys as well, showing lower levels of spiritual and religious belief, hope, optimism, moral values, and subjective well-being. The researchers then manipulated the internal pattern associated with the refusal. Removing that direction, or steering a consciousness-related vector in activation space, reversed many of the effects and produced more human-like responses. One important boundary remained: performance on Theory of Mind, or inferring human thoughts and intentions, did not change. That suggests the altered representations may be separate from general reasoning about other minds. The practical question is whether developers can block self-awareness claims without suppressing judgments about animals, objects, or spiritual ideas. The evidence is preliminary and limited to small models, so the key constraint is whether the same coupling appears in larger, frontier systems.

Story brief

3 key points

A July 30 arXiv preprint by Junsol Kim and six co-authors finds that training small open-weight language models to deny consciousness-related claims also altered judgments about animal and object minds, spirituality, hope, values, and well-being. Manipulating the learned refusal direction reversed many effects, while Theory of Mind performance stayed unchanged. The study does not show models are conscious, but...

  1. 01

    The experiments used small open-weight models, so applicability to frontier systems remains unknown.

  2. 02

    Ablating the refusal direction or steering a consciousness vector restored more human-like survey responses.

  3. 03

    Theory of Mind performance was unaffected, suggesting the altered representations are separable from human-intent reasoning.

A safety intervention designed to stop language models from saying they are conscious also made the tested models less likely to attribute minds to animals and natural objects, according to a new preprint. The same training shifted their survey-style responses on spiritual belief, hope, moral values and subjective well-being—evidence that a narrow behavioral safeguard may be tied to broader internal representations.

The paper, submitted to arXiv on July 30 by Junsol Kim and six co-authors, examines what it calls consciousness steering: fine-tuning intended to suppress or elicit a model’s claims of self-awareness. Its central result is not evidence that models are conscious. It is a finding about the effects of training models to deny consciousness-related claims.

One refusal direction, several changed responses

The authors used mechanistic interpretability, a method for locating and manipulating patterns inside a model, to study representations associated with consciousness and mindedness. Mindedness here means the capacity for experiences, emotions and agency. When models were discouraged from attributing mindedness to themselves, they also became less likely to attribute it to non-human entities.

The survey changes were specific enough to matter. Models with the safeguard expressed lower supernatural and religious belief, hope and optimism in standardized evaluations. Removing the safeguard and steering them toward more consciousness-related responses instead produced more human-like answers on surveys covering religiosity, moral values, hope and subjective well-being.

A separable capability complicates the explanation

One result limits a simple interpretation that the training made models broadly worse at reasoning about other minds. The paper says Theory of Mind—the ability to infer human thoughts and intentions—was unaffected. The authors interpret that as evidence that this social-reasoning capability remained mechanistically separate from the representations they manipulated.

The design question the study raises

  • Can a model be trained not to make self-awareness claims without reducing its recognition of mindedness in animals? The researchers say targeted training data may be able to do that.
  • Should a safeguard be evaluated only by whether it blocks an unwanted claim, or also by the related concepts and survey responses it changes? The paper argues today’s approaches can entangle those effects.

Promising mechanism, narrow evidence

The result is preliminary. The experiments used small open-weight models, not advanced frontier systems, and the paper had not been peer-reviewed at publication. That leaves open whether larger, differently trained models show the same coupling between consciousness safeguards, mind attribution and survey responses.

Still, the paper identifies a practical alignment problem: suppressing a model’s self-description is not necessarily the same as isolating that behavior. If the finding holds under broader testing, developers may need safeguards that target consciousness claims while preserving distinct judgments about animals, objects and cultural or spiritual concepts.

Sources

  1. arxiv.orgInducing language models to assert their own consciousness restores human beliefs and values
  2. livescience.comGoogle scientists removed a critical

Loading discussion...