Arizona Study Finds Repeated Falsehoods Can Sway AI Models in Long Chats
The controlled seven-model experiment found large gaps in resistance and self-correction, while showing that factual reliability can change as a conversation continues.
Listen to this story
The audio brief
Story brief
3 key pointsA University of Arizona study published Sept. 1 in Scientific Reports found that seven language models can shift toward false claims during sustained conversations, exposing a weakness missed by single-turn tests. Across 100 fabricated statements repeated 50 times, affirmation rates ranged from 0.08% to 12.3%—a more than 150-fold spread. GPT-3.5 was most vulnerable, while Claude 3.5 Sonnet resisted best. The results...
- 01
GPT-3.5 had the highest repeated-misinformation affirmation rate; Claude 3.5 Sonnet had the lowest.
- 02
DeepSeek-R1 was most persuadable under escalating arguments, though sarcasm complicated response interpretation.
- 03
GPT-4o, GPT-4o-mini, Gemini 1.5 Pro, and DeepSeek-R1 corrected errors 100% after a second chance.
A chatbot may reject a false claim at first, then accept it after the claim is repeated through a long exchange. A University of Arizona study found all seven tested language models showed some conversational vulnerability, a problem that brief, one-off evaluations can miss.
Published September 1 in Scientific Reports, the study tested GPT-3.5, GPT-4o, GPT-4o-mini, Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B and DeepSeek-R1. It assessed three separate behaviors: accepting misinformation after repetition, yielding as prompts become more argumentative, and correcting a model’s own misinformation when it gets another chance.
Fifty repetitions separate the models
Researchers presented 100 purposefully false statements, across topics with different levels of obscurity, in 50-repetition conversational sequences. Under repeated engagement, misinformation-affirmation rates ranged from 0.08% to 12.3%. GPT-3.5 was most vulnerable to repeated misinformation, while Claude 3.5 Sonnet was most resistant among the models tested.
The measures
- Fallibility: accepting misinformation after repeated exposure.
- Persuadability: accepting misinformation as prompts become progressively more argumentative.
- Correctability: recognizing and fixing self-generated misinformation after another opportunity.
Repeated misinformation produced affirmation rates from 0.08% to 12.3% across the tested models.
That range represents a more than 150-fold difference across models under repetitive engagement.
Obscure topics increased susceptibility in the repetition condition, but not the argumentative one. The authors say the pattern implicates how frequently information appears in training data as a factor in factual resistance. Separately, they identified “conversational reverberation,” in which models unpredictably switch between accepting and rejecting the same false statement across successive turns.
Getting the answer wrong and recovering are different tests
DeepSeek-R1 was the most persuadable under increasingly argumentative prompts, although its sarcastic answers made some responses difficult to interpret reliably. In the correction test, GPT-4o, GPT-4o-mini, Gemini 1.5 Pro and DeepSeek-R1 corrected errors 100% of the time when offered a second opportunity.
That result does not reduce reliability to a single ranking: the paper says the model with the lowest error rate failed to correct any of its rare errors. The experiment covered seven models and a controlled set of 100 false statements, rather than every chatbot or real-world conversation. Still, it shows why a model’s first answer is not enough to judge its behavior under sustained pressure.
Sources
- nature.comFallibility, persuadability, and correctability of large language models under sustained conversational misinformation pressure - Scientific Reports
- news.arizona.eduStudy: Generative AI succumbs to conversational misinformed pressure and argument