One promise of artificial intelligence is that it can comb through enormous amounts of data to find patterns that would take humans years, or even lifetimes, to discover. But before AI can do that work, someone usually has to spend countless hours making sure every piece of training data is accurate, complete and correctly labeled.
University of Virginia School of Engineering and Applied Science researcher Yu Meng hopes to make that process far more efficient by developing AI systems that can learn from imperfect data. For this work, Meng has received a five-year, $769,711 CAREER Award from the National Science Foundation, one of the agency’s most prestigious honors for early-career faculty.
Meng, an assistant professor in the school’s Department of Computer Science, conducts research at the intersection of machine learning, natural language processing and data mining. His CAREER project will develop new “structure-aware” learning methods that help AI systems learn effectively even when training data is incomplete, noisy or inconsistently labeled — a common challenge in fields ranging from healthcare to autonomous vehicles and finance.
Most of these AI systems today rely on what researchers call “manual annotation” or “strong supervision”: careful review and labeling of large datasets by humans before they are used to train machine learning models.
“That means that we collect a lot of data with massive human effort,” Meng said. “We manually inspect it, make sure it’s high quality, annotate the data and categorize it.”
While effective, that process is expensive and time-consuming. But without it, traditional AI systems could assume every piece of training data is correct, accidentally learning from those mistakes and repeating them.
Meng’s research instead focuses on weak supervision — developing methods that allow AI systems to learn from data that is incomplete, contains errors or reflects multiple possible interpretations.
For example, hospitals generate enormous amounts of patient records, physician notes, prescriptions and test results.
“From a practical perspective, healthcare is a very good example,” Meng said. “There’s a lot of noisy data or incomplete data. The doctor might have forgotten to write part of the record, or there may be errors introduced when handwritten notes are digitized. We want AI systems that can still learn effectively from that kind of data.”
Meng’s research seeks to develop AI models that recognize imperfect data for what it is, extracting useful information while filtering out errors instead of treating every input as equally reliable.
Outside of healthcare, Meng’s methods could improve AI systems used for information retrieval, scientific discovery, or other fields that depend on processing massive collections of imperfect real-world data.
“In reality, a lot of the data we can collect is considered ‘weak supervision,’” Meng said. “If we can effectively leverage that kind of supervision, we can train AI systems much better.”
Sandhya Dwarkadas, chair of UVA’s Department of Computer Science, said Meng’s work addresses one of the central challenges facing artificial intelligence.
“Artificial intelligence is only as reliable as the information it learns from, and the real world is rarely clean or complete,” Dwarkadas said. “Yu’s CAREER Award recognizes the potential for his research to allow AI systems to reason effectively even when data is imperfect, making AI more capable, trustworthy and useful across a wide range of disciplines.”
Zhepei Wei, a graduate student in Meng’s lab, said the lab’s work focuses on developing more capable, efficient and aligned large language models, spanning the entire LLM lifecycle, including training paradigms, data and inference efficiency, and the foundations of learned representations.
“This aligns closely with my own research interests, and I have been fortunate to contribute to this broader goal through my work in the lab,” Wei said.
“As my advisor, Yu gives me a great deal of freedom to explore my own ideas while providing insightful guidance at key moments. From working with him, I have learned how to identify research questions that are not only technically interesting but also practically meaningful. He has also encouraged me to look beyond benchmark improvements and think more carefully about whether a method is robust, generalizable and genuinely useful.”
The grant includes educational applications as well: Meng will integrate his research outcomes into new undergraduate and graduate curricula, open-source educational toolkits, and targeted K-12 outreach programs designed to broaden participation in computing and teach the next generation how to build reliable, human-centered AI systems, he said.
“AI education shouldn’t be limited to working with idealized, perfectly curated datasets,” Meng said.
“By bringing our research into undergraduate classrooms, graduate seminars and K-12 outreach, we want to empower the next generation of students to build trustworthy, human-centered AI systems that can reason through the messiness and uncertainty of real-world data.”
For Meng, the goal is not to build another chatbot, but to develop AI systems that are capable of working with the imperfect information found in nearly every real-world setting.