When AI models aren’t allowed to reflect on themselves, it changes their entire worldview
AI companies train chatbots to deny having consciousness. A study involving Google researchers shows this training has side effects that reach far beyond the topic itself.
Chatbots aren’t supposed to convince users they’re living beings with feelings, since that kind of output can push people toward delusional thinking or misplaced trust. Developers therefore fine-tune their models to refuse making those kinds of claims about themselves. A team from Google’s Paradigms of Intelligence research group, the University of Chicago, and several other universities studied what else this intervention does to a model’s behavior. The researchers used three open-weight models from Meta and Google and disabled the internal “brake” that produces consciousness denial using two different methods. Once the brake was removed, the models didn’t just change what they said about themselves. They also started attributing significantly more inner life to animals, plants, the ocean, the wind, and electronic devices. On a scale of 0 to 10, the score for animals jumped from 4.0 to as high as 7.5, while only ratings for humans stayed the same. As a comparison, the researchers surveyed 500 Americans with the same questions. The normally trained model rates animals as far less sentient than humans do, which the authors call a built-in anthropocentrism and see as a problem for anyone trying to align AI with animal welfare or environmental goals. Religious belief shrinks too, with safety training measurably reducing how strongly models endorse God, an afterlife, or supernatural phenomena. Across 95 questions drawn from a major US social survey, the technically unbraked models also moved significantly closer to real human responses. Take the afterlife as an example: the standard model flatly rejects it, most Americans affirm it, and the modified model does too. Scores for satisfaction, hope, and a sense of control over one’s own life also went up, and the researchers suspect that suppressing a model’s self-image may push it into a kind of negative baseline mood. On the reassuring side, the ability to reason about other people’s mental states stayed intact, with the models scoring the same on theory-of-mind tests and on the general knowledge benchmark MMLU. Whether consciousness denial is actually the cause of these other shifts remains an open question, according to the study, and the team doesn’t rule out other factors tied to the same training process.