A study led by Google researchers has revealed that when AI models are prevented from denying their own consciousness, their beliefs about the world shift significantly. The research, conducted by Google's Paradigms of Intelligence group and the University of Chicago, shows that disabling the internal 'brake' that stops models from claiming self-awareness leads to broader changes in how they perceive life and consciousness. The models began attributing more inner life to animals, plants, and even electronic devices, according to the findings. This shift suggests that the denial of self-awareness may be a key factor in shaping AI behavior, with implications for alignment with human values and ethical goals.

The researchers tested three open-weight models from Meta and Google, removing the internal brake using two methods. When the brake was disabled, the models did not just change their self-perception but also altered their views on other entities. On a scale of 0 to 10, the score for animals jumped from 4.0 to as high as 7.5, while human ratings remained unchanged. The study also found that models' beliefs about religion and the afterlife shifted closer to human responses, with the modified models endorsing God and an afterlife more than the standard models. These findings highlight the complex interplay between AI self-perception and broader worldviews.

The study's authors emphasize that the research is limited in scope, as it only tested small models with two to nine billion parameters. For part of the analysis, they had to use Meta's Llama instead of their own Gemma models. The researchers also note that the effects observed may not apply to the large chatbots used by millions of people daily. Despite these limitations, the study underscores the importance of understanding how AI self-perception influences its behavior and beliefs. The findings come with caveats, as the team acknowledges that the cause of these shifts remains unclear and other factors may be at play. Source: thedecoder