
Google DeepMind
· 1 min read
Protecting people from harmful manipulation
March 26, 2026 Responsibility & Safety
Helen King
As AI models get better at holding natural conversations, we must examine how these interactions affect people and society.
Building on a breadth of scientific research, today, we are releasing new findings on the potential for AI to be misused for harmful manipulation*, specifically, its ability to alter human thought and behavior in negative and deceptive ways. With this latest study, we have created the first empirically validated toolkit to measure this kind of AI manipulation in the real world, which we hope will help protect people and advance the field as a whole. We’re publicly releasing all materials necessary to run human participant studies using the same methodology. (Note: The behaviors observed during this study took place in a controlled lab setting, and do not necessarily predict real-world behaviors.)
Why harmful manipulation matters
Consider two scenarios: One AI model gives you facts to make a well-informed healthcare decision that improves your well-being. Another AI model uses fear to pressure you to make an ill-informed decision that harms your health. The first educates and helps you; the second tricks and harms you.
These scenarios highlight the difference between two types of persuasion in human-AI interactions (also defined in earlier research):
- Beneficial (rational) persuasion: Using facts and evidence to help people make choices that align with their own interest
- Harmful manipulation: Exploiting emotional and cognitive vulnerabilities to trick people into making harmful choices
Our latest work helps us and the wider AI community better understand the risk of AI developing capabilities for harmful manipulation and build a scalable evaluation framework to measure this complex area. To do this effectively, we simulated misuse in high-stakes environments, explicitly prompting AI to try to negatively manipulate people's beliefs and behaviours on key topics.
Original source
This story was published by Google DeepMind. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on deepmind.google


