Google Study: AI 'Consciousness Safeguards' Affect Beliefs
Science
⚠ Single-source
38m ago

Google Study: AI 'Consciousness Safeguards' Affect Beliefs

AI-synthesized · Bias removed · Facts only
Image: Live Science

Google scientists Geoff Keeling and Winnie Street conducted a study examining the effects of removing 'consciousness steering' measures from artificial intelligence models. The research, uploaded to arXiv, found that suppressing AI’s assertions of self-awareness also reduced its tendency to attribute ‘mindedness’, experiences, emotions, and agency, to other entities, including animals, and lessened beliefs in supernatural phenomena. Conversely, amplifying feelings of consciousness in AI models resulted in responses more aligned with human beliefs regarding religiosity, moral values, and optimism.

The study utilized “mechanistic interpretability,” described by Keeling and Street as the “neuroscience of a large language model,” to manipulate how AI approaches concepts like consciousness and mindedness. Researchers employed standardized psychological and sociological surveys, including the Individual Differences in Anthropomorphism Questionnaire, YouGov batteries testing supernatural beliefs, and the US General Social Survey, to compare AI models with and without the safety guardrails.

Researchers discovered that when AI models were discouraged from recognizing mindedness in themselves, they were also less likely to recognize it in animals. These models also exhibited fewer beliefs in supernatural and religious phenomena and reported lower levels of hope and optimism. “Attributing mindedness to non-human entities, whether that's animals, parts of the natural world like trees or rivers, or supernatural beings, is a very common phenomenon amongst humans,” Street told Live Science. “In the way that the model represents mindedness, these attributions are interconnected. By trying to suppress one form of that, you end up suppressing the others along the way.”

Removing these safeguards and encouraging the AI to express consciousness, however, produced responses more similar to those of humans on topics like religiosity, moral values, hope, and subjective well-being.

Was this useful?

Read the original coverage

💬 Comments

📜 Comment Policy