Modelspublished

Arizona Study Finds Repeated Falsehoods Can Sway AI Models in Long Chats

The controlled seven-model experiment found large gaps in resistance and self-correction, while showing that factual reliability can change as a conversation continues.

By 2 min read
Arizona Study Finds Repeated Falsehoods Can Sway AI Models in Long Chats
Arizona Study Finds Repeated Falsehoods Can Sway AI Models in Long Chats

Listen to this story

The audio brief

About 1:44
0:001:44
Read transcript
A chatbot can reject a false claim, then accept it after the same claim keeps coming back. That is the central finding from a University of Arizona study published September first in Scientific Reports, which tested seven language models across long, controlled conversations. Researchers repeated one hundred fabricated statements up to fifty times. The models’ affirmation rates ranged from just zero point zero eight percent to twelve point three percent—a gap of more than one hundred fifty times. GPT-3.5 was the most vulnerable to repeated misinformation, while Claude 3.5 Sonnet was the most resistant among the systems tested. The study also separated three behaviors: fallibility, or accepting a falsehood; persuadability, or yielding as an argument becomes more forceful; and correctability, or fixing an error when given another chance. DeepSeek-R1 was the most persuadable under escalating arguments, although sarcasm made some of its answers harder to interpret. On the correction test, GPT-4o, GPT-4o-mini, Gemini 1.5 Pro, and DeepSeek-R1 corrected their errors every time. The researchers also observed “conversational reverberation”: models sometimes switched unpredictably between accepting and rejecting the same false statement. Obscure claims were more vulnerable to repetition, but not to argumentative pressure. The important constraint is that this was a controlled test of seven models and one hundred false claims—not every chatbot or real-world exchange. It still shows why a first answer is not enough: evaluations may need to measure persistence, escalation, and recovery across an entire conversation.

Story brief

3 key points

A University of Arizona study published Sept. 1 in Scientific Reports found that seven language models can shift toward false claims during sustained conversations, exposing a weakness missed by single-turn tests. Across 100 fabricated statements repeated 50 times, affirmation rates ranged from 0.08% to 12.3%—a more than 150-fold spread. GPT-3.5 was most vulnerable, while Claude 3.5 Sonnet resisted best. The results...

  1. 01

    GPT-3.5 had the highest repeated-misinformation affirmation rate; Claude 3.5 Sonnet had the lowest.

  2. 02

    DeepSeek-R1 was most persuadable under escalating arguments, though sarcasm complicated response interpretation.

  3. 03

    GPT-4o, GPT-4o-mini, Gemini 1.5 Pro, and DeepSeek-R1 corrected errors 100% after a second chance.

A chatbot may reject a false claim at first, then accept it after the claim is repeated through a long exchange. A University of Arizona study found all seven tested language models showed some conversational vulnerability, a problem that brief, one-off evaluations can miss.

Published September 1 in Scientific Reports, the study tested GPT-3.5, GPT-4o, GPT-4o-mini, Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B and DeepSeek-R1. It assessed three separate behaviors: accepting misinformation after repetition, yielding as prompts become more argumentative, and correcting a model’s own misinformation when it gets another chance.

Fifty repetitions separate the models

Researchers presented 100 purposefully false statements, across topics with different levels of obscurity, in 50-repetition conversational sequences. Under repeated engagement, misinformation-affirmation rates ranged from 0.08% to 12.3%. GPT-3.5 was most vulnerable to repeated misinformation, while Claude 3.5 Sonnet was most resistant among the models tested.

The measures

  • Fallibility: accepting misinformation after repeated exposure.
  • Persuadability: accepting misinformation as prompts become progressively more argumentative.
  • Correctability: recognizing and fixing self-generated misinformation after another opportunity.
The gap under repeated pressure
0.08%Lowest affirmation rate

Repeated misinformation produced affirmation rates from 0.08% to 12.3% across the tested models.

12.3%Highest affirmation rate

That range represents a more than 150-fold difference across models under repetitive engagement.

Obscure topics increased susceptibility in the repetition condition, but not the argumentative one. The authors say the pattern implicates how frequently information appears in training data as a factor in factual resistance. Separately, they identified “conversational reverberation,” in which models unpredictably switch between accepting and rejecting the same false statement across successive turns.

Getting the answer wrong and recovering are different tests

DeepSeek-R1 was the most persuadable under increasingly argumentative prompts, although its sarcastic answers made some responses difficult to interpret reliably. In the correction test, GPT-4o, GPT-4o-mini, Gemini 1.5 Pro and DeepSeek-R1 corrected errors 100% of the time when offered a second opportunity.

That result does not reduce reliability to a single ranking: the paper says the model with the lowest error rate failed to correct any of its rare errors. The experiment covered seven models and a controlled set of 100 false statements, rather than every chatbot or real-world conversation. Still, it shows why a model’s first answer is not enough to judge its behavior under sustained pressure.

Sources

  1. nature.comFallibility, persuadability, and correctability of large language models under sustained conversational misinformation pressure - Scientific Reports
  2. news.arizona.eduStudy: Generative AI succumbs to conversational misinformed pressure and argument