Add BrightSurf on Google Email

Study: Generative AI succumbs to conversational misinformed pressure and argument

09.04.26 | University of Arizona
SAMSUNG T9 Portable SSD 2TB

SAMSUNG T9 Portable SSD 2TB transfers large imagery and model outputs quickly between field laptops, lab workstations, and secure archives.

The results are in: Which AI model is the most fallible? Persuadable? Correctible?

University of Arizona research assessed seven different generative AI language learning models, or LLMs, for these three qualities during lengthy conversation. Their work, published in Nature's Scientific Reports, reveals intrinsic limitations that might go undetected during one-off interactions.

Among the seven LLMs tested – ChatGPT (GPT-3.5, GPT-4o, GPT-4o-mini), Claude 3.5, Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek-R1 – they found that:

"This underscores the need for careful human engagement and the danger of blind reliance," said senior study author Dr. Marvin Slepian, Regents Professor of medicine and biomedical engineering. "When generative AI came out in November 2022, there was a lot of regulation potential, but that has since fell by the wayside. People are recognizing the onus is now left to the users."

Many people are familiar with AI's limitations, such as a tendency toward sycophancy, or the tendency to agree with users, and hallucination, or confidently wrong answers, but there has been very little work on evaluating AI's limitations during what is called multi-turn conversations, in which answers are predicated on previous context. Such usage more closely mirrors the real-world, according to the research team.

"These limitations raise important safety concerns, particularly as generative AI systems are increasingly deployed in high-stakes settings," Slepian said.

The team also identified four different ways the models failed to affirm factual information. For example, some models oscillated between accepting and rejecting the same false statement during the conversation.

"If one were relying on the model for critical decision-making, one might – depending upon the phase of the oscillation – 'fire the missile' or 'cut off the leg,' or not, based simply on chance," Slepian said.

As a medical doctor, specifically a cardiologist member of the Sarver Heart Center , he characterizes such failures as "pathologies." This specific pathology was dubbed "reverberation."

"This is dangerous," said Slepian, who led the artificial intelligence subcommittee of the United States Patent and Trademark Office until last year. He is also a member of the James E. Rogers College of Law faculty. "How can we use fickle systems that are not reproducible? These need to be fixed, but this study has spanned three years, and there's still the same unfixed characteristics."

But, closed models, such as Chat and Claude, make it impossible to "peek under the hood" to diagnose and solve the problem.

However, as the founder and director of the Arizona Center for Accelerated Biomedical Innovation , or ACABI, Slepian and his team are beginning to develop diagnostic tools for open AI models as part of their AI Pathology Lab.

"I use AI and so does my team, but as scientist and physician, I have to understand the anatomy and physiology, then understand pathologies – what can go wrong – to diagnose and prevent them. The same goes for AI."

Co-authors on the study include the U of A's Jordan Rodgriguez, Zachary Hansen, Luis De Anda, Katelyn Rohrer and Camila Grubb – all computer science students and researchers in ACABI; as well as Mihai Surdeanu and Enrique Noriega of the Department of Computer Science.

Scientific Reports

10.1038/s41598-026-68231-0

Data/statistical analysis

Not applicable

Fallibility, persuadability, and correctability of large language models under sustained conversational misinformation pressure

1-Sep-2026

The authors declare no competing interests.

Keywords

Article Information

Contact Information

Mikayla Mace Kelley
University of Arizona
mikaylamace@arizona.edu

Source

This article is based on a news release from University of Arizona. BrightSurf curates and republishes science news from research institutions worldwide; the original release is linked below.

How to Cite This Article

APA:
University of Arizona. (2026, September 4). Study: Generative AI succumbs to conversational misinformed pressure and argument. Brightsurf News. https://www.brightsurf.com/news/LVDO5GEL/study-generative-ai-succumbs-to-conversational-misinformed-pressure-and-argument.html
MLA:
"Study: Generative AI succumbs to conversational misinformed pressure and argument." Brightsurf News, Sep. 4 2026, https://www.brightsurf.com/news/LVDO5GEL/study-generative-ai-succumbs-to-conversational-misinformed-pressure-and-argument.html.