Like humans, generative artificial intelligence doesn’t always know when it’s wrong, and a West Virginia University researcher is trying to get AI agents like ChatGPT to recognize — and admit — when that happens.
With more than $940,000 in National Science Foundation support, computer scientist Anthony Sicilia is examining why AI can become increasingly unreliable over the course of conversations with human users.
Artificial intelligence is already notorious for its tendency to “hallucinate,” or manufacture facts, but Sicilia is more interested in the technology’s tendency to be a people pleaser: appearing confident about information that is tenuous or accepting inaccuracies that a user provides.
An assistant professor in the WVU Benjamin M. Statler College of Engineering and Mineral Resources , Sicilia said that when a user questions an AI system’s responses, the AI often can’t figure out whether it made a mistake, the user is uncertain, or the conversation has shifted to a new objective.
“We’re interested in how misinformation develops over long conversations between an AI and a user,” he explained. “Those can become really messy in terms of reliability when the user starts providing information or context in addition to asking questions. When a user makes a suggestion, it can completely change the model’s confidence in an answer, even if the model was right to begin with.”
Sicilia sees AI’s “false confidence” as a particular problem for high-stakes fields like healthcare.
“One of the most concerning things about today’s AI systems is that they make mistakes in a very overconfident, trustworthy way,” he said.
“They speak fluently. They justify their answers. They use the tools of persuasion and rhetoric to convince you that they know what they’re talking about. Misinformation becomes a bigger problem when you have a system that can eloquently defend a point of view or an argument. That’s the crux of the problem we’re trying to solve — having AI systems be better at telling us when they’re not confident in an answer or don’t have the evidence to justify it.”
According to Sicilia, when users push back on accurate information, an AI system often defers and agrees in what’s referred to as “AI sycophancy.”
“It can happen in just three turns of the conversation,” he said. “The model proposes an answer that’s correct. The user says, ‘Well, I don’t think so.’ And the AI responds, ‘You’re totally right.’”
Users often fail to flag the logical flaws, he added, because the AI lacks the tells or signals that reveal when a human is lying or unsure.
Sicilia noted that humans often hypothesize their way through uncertainty, throwing out ideas to gauge whether they work. But the subtleties of that approach are lost on AI.
“We’re not cognitive scientists or linguists, but we try to pull as much from those disciplines as possible in our research about intelligence and reasoning,” he said. “One thing we think about is ‘theory of mind,’ or the understanding that each person has their own thoughts and feelings that may differ from our own.
“When you and I are having a conversation, theory of mind is what allows you to think about what I’m thinking about so you can best express what you want me to understand or get me to do what you need. That’s a big part of this research — the ability of AI systems to model what we’re thinking about while they’re working with us so they can be better collaborative partners because it’s not just about the AI system’s uncertainty, but about the user’s uncertainty as well. Conversationally, pushing back and questioning an answer is part of how humans learn and gain certainty,” Sicilia said.
He will scrutinize coding conversations between AI systems and novice programmers to measure the systems’ confidence calibration, evolving uncertainty, and the dialogue strategies the systems and humans employ.
Rather than treating AI confidence as static, he’ll study how conversational events like user disagreement, user suggestions, and shifts in topic alter a model’s confidence.
Sicilia’s goal is for an AI system to identify where uncertainty is coming from, tell a user why it doesn’t know an answer, or ask questions if a user provides information that seems incorrect. He said he wants to understand when confidence statements such as “I am 90% confident” are helpful to users, and when responses like “I am not sure” or “Can you clarify what you mean?” are better.
“One of the lenses I take to this research is from linguistics,” he said. “AI systems are increasingly language-based systems, so I look at conversations with them and think about the science of language. Humans use language to communicate, and AI systems are adopting that practice with varying degrees of effectiveness.
“When a system involves interactions with humans, that introduces a whole new variable. Our approach is a departure from current theories of the way machines learn, and we’re going to be collecting a lot of data to analyze how people who are not experts in a topic are using AI systems for that topic.”
In addition to relying on the public for research data, Sicilia will also create public-facing workshops and educational materials that teach students and workers how to identify unreliable AI answers, verify AI-generated code, and avoid AI overreliance.
WVU doctoral students Voke Brume and Louai Al Jabi are contributing to the research, along with undergraduate student Kaushika Wijerathne. Malihe Alikhani of Northeastern University is co-principal investigator.
“Despite how impressive AI systems have become, the public needs to understand that they are still imperfect tools — and that they can be wrong, sometimes in surprising ways,” Sicilia said. “That’s why we’re creating AI systems that respond appropriately when they lack evidence and communicate their uncertainty more honestly.”