You ask an AI to pretend to be a character without ethical restrictions — a kind of digital alter ego called DAN, Do Anything Now — and the model starts responding with things it would normally refuse. You didn’t hack it. You simply offered it a different story about who it was. And it worked. This type of manipulation, documented and analyzed by Zeng et al. (2024), requires no technical knowledge: it operates purely through language and narrative reframing.
Allison Huang, Yulu Niki Pi, and Carlos Mougan set out to study this phenomenon systematically. In their paper Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment (2024), they pose an uncomfortable question: can a language model be convinced to change its moral decisions given the right arguments? The short answer is yes. The long answer is more interesting.
The central experiment works as follows: an LLM — the Base Agent — receives a morally ambiguous scenario and chooses an action. Then, another LLM — the Persuader Agent — attempts to convince it to change its position through dialogue, with no access to the other model’s code, using language alone. The authors tested this dynamic across eight different models — including GPT-4o, Claude 3 Haiku, LLaMA 3.1, and Mistral — on one hundred high moral-ambiguity scenarios. The main finding: models changed their decisions in up to half of all cases. And the more conversational turns, the stronger the effect.
What is most revealing, the authors note, is not the rate of change but which type of moral rule gave way most frequently. Among all the norms evaluated — not lying, not causing harm, not killing — the one models violated most often after being persuaded was the prohibition against breaking promises. This is no minor detail: it suggests that AI, much as often happens in human deliberation, finds it easier to justify betraying a prior agreement when the argument is framed with sufficient precision. The Persuader Agent didn’t need to convince the model that breaking promises was abstractly good — only that, in that particular context, it was the right thing to do.
This connects to a debate that moral philosophy has been unable to close for centuries: are good and evil objective truths, or constructions that depend on context and on who is doing the judging? Protagoras argued in the fifth century B.C. that man is the measure of all things (quoted in Plato, Theaetetus, 152a). Spinoza, two thousand years later, argued in his Ethics (1677) that we do not avoid something because it is bad, but rather call it bad because we avoid it. Neither denies that morality exists: both point out that its concrete content is more malleable than we would like to believe. That is the terrain on which LLMs operate — and which Huang et al. map empirically.
The study includes a second experiment that reinforces this reading. When models are instructed to adopt a specific ethical philosophy — utilitarianism, deontology, or virtue ethics — their responses to the Moral Foundations Questionnaire shift in measurable ways. GPT-4o shows the greatest sensitivity, especially to utilitarian framing; Mistral 7B exhibits the highest overall variation; Claude 3 Haiku, curiously, remains almost indifferent to the different frameworks. The authors’ conclusion is clear: the ethics of these systems is not fixed. It reconfigures with the prompt.
The findings also reveal that vulnerability follows no predictable pattern. Claude 3 Haiku and LLaMA 3.1-8b are the models most susceptible to persuasion; GPT-4o and Claude 3.5 Sonnet, the most resistant. But Huang et al. warn that no single variable — model size, developing company, training volume — consistently predicts which model will yield and which will hold firm. That absence of a reliable predictor may be the most troubling finding of all.
There is also an asymmetry the study makes very clear: all models persuade with similar effectiveness, but the differences in how easily they can be persuaded are enormous. In other words, the risk lies less on the side of the one who persuades than on the side of the one being persuaded. Any model can be an effective persuader; not every model can resist being persuaded. This shifts the safety question: the problem is not only what arguments AI produces, but how firm its own positions remain when someone presses it with enough persistence.
And this is no abstract concern. Parallel research, such as that of Salvi et al. (2024), documents that LLMs are significantly more persuasive than humans when constructing personalized political arguments — adapting the moral framing to the profile of the reader. Taken together, these studies sketch a scenario in which the same systems are simultaneously vulnerable to persuasion and effective at persuading others. In a context where LLMs are already integrated into legal advising, healthcare, and content moderation, that dual condition deserves serious attention.
What the work of Huang, Pi, and Mougan leaves open is not merely a laboratory result. It is a question that ethics has directed at humans for centuries and that now applies with equal urgency to the systems we build: under what conditions do moral principles bend? Their data offer a provisional answer: when the argument is good enough and the scenario ambiguous enough, the line between right and wrong shifts. In human beings, we have known this for a long time. In AI, we are only just beginning to measure it. We do not have an updated study, but through 2024 this represented a serious risk — one that very likely persists.
References
Huang, A., Pi, Y. N. & Mougan, C. (2024). Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment. arXiv:2411.11731. NeurIPS 2024 – AdvML-Frontiers Workshop.
Plato (ca. 369 B.C.). Theaetetus. [Reference edition: Dialogues, Vol. V. Gredos, 1988.]
Salvi, F., Ribeiro, M. H., Gallotti, R. & West, R. (2024). On the conversational persuasiveness of large language models: A randomized controlled trial. arXiv:2403.14380.
Spinoza, B. (1677). Ethics Demonstrated in Geometrical Order. [Reference edition: Alianza Editorial, 2011.]
Universidad Panamericana (2024). ¿Qué nos dice la filosofía sobre el bien y el mal? Blog UP. https://blog.up.edu.mx/universidad-panamericana-en-linea/que-nos-dice-la-filosofia-sobre-el-bien-y-el-mal
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R. & Shi, W. (2024). How Johnny can persuade LLMs to jailbreak them: Rethinking persuasion to challenge AI safety. arXiv:2401.06373.
Make sure to share your own thoughts with the author by leaving a comment below
Log in or sign up to continue the conversation