Add BrightSurf on Google Email

Variations between labs can misinform scientific AI models, team reports

08.13.26 | Penn State
DJI Air 3 (RC-N2)

DJI Air 3 (RC-N2) captures 4K mapping passes and environmental surveys with dual cameras, long flight time, and omnidirectional obstacle sensing.


UNIVERSITY PARK, Pa. — To turn abundant carbon dioxide into valuable fuel, researchers need a fast and efficient way of determining which catalysts work best over the longest time. Artificial intelligence (AI) models have the potential to help guide catalyst selection, but as with internet chatbots, AI models are only as good as the data put into them.

By convening four laboratories from across the nation to test an experimental carbon monoxide-producing catalyst, a key first step in turning carbon dioxide into fuels, scientists at SLAC National Accelerator Laboratory , alongside two researchers from Penn State, have demonstrated the importance of generating highly reproducible experimental data when building AI models for investigations in science. They published the results in Nature Catalysis .

“An AI model is only as good as the data used to create it, which will undoubtedly come from multiple sources,” said Robert Rioux , Friedrich G. Helfferich Professor of Chemical Engineering at Penn State and co-author on the study. “This study represents the first round-robin study focused on heterogeneous catalysis that quantifies the uncertainty from studies done between labs, and the impact this uncertainty will have on future data-driven AI modeling.”

With a good AI model, researchers can enter conditions such as temperature, length of time of the reaction, and catalyst formulation, then run the simulation and see a prediction of how well the catalyst performs. They can then confirm the predictions with a few well-designed experiments, ultimately speeding up catalyst discovery and implementation at a global scale.

In addition to saving time and money, such models can also explore conditions that are difficult to achieve in the lab. Most lab catalysis studies can only look at short time periods — days — but catalyst deactivation occurs over the course of months or years due to buildup of impurities and repeated exposure to high temperatures.

AI models need large amounts of high-quality data for training. To generate the data, the four labs performed a set of round-robin experiments, in which multiple laboratories conduct the same tests to evaluate reproducibility using previously agreed upon protocols and the same rhodium-based catalyst.

To the researchers’ surprise, achieving the same results from four labs working independently was harder than anticipated. When they got together to share their results, they realized that they had a problem.

Each of the four research teams produced results that varied in the amounts of carbon monoxide and methane, an undesirable side product, produced. The computer would not be able to learn from four sets of data that contain different outcomes.

“It was a bit of an eye-opener,” said SLAC staff scientist Adam Hoffman, senior author of the study. “This experience shines light on the practical challenges of including real-world data into machine learning models.”

Painstakingly, the teams evaluated their methods. Through rigorous testing, they found a handful of sources of mismatch, with one of the biggest contributors to the variability coming down to how hard the mixture was shaken or stirred.

“Our findings are a reminder to exercise caution about what information we feed into a machine-learning model, and how the consistency of experimental data can influence the reliability of the outcomes,” said Selin Bac, a postdoctoral researcher at the University of California, Santa Barbara, and first author on the study.

With further standardization across the four labs — which in addition to SLAC included groups at Penn State, Stanford University and University of California, Santa Barbara — the results began to look more consistent. The team outlined several recommendations to strengthen experimental reproducibility, including enhancing the consistency of reactor design, operating protocols and experimental conditions.

“We contributed to the characterization of the catalytic materials used in the round-robin study,” Rioux said. “Our findings reinforce the idea that high-quality experimental data dictates the success of any AI-based model by determining model accuracy, reliability and scope.”

Hoffman said he hopes the study will help experimentalists and data scientists who are designing AI models to consider how small variations in experimental design across labs can lead to problems with reproducibility and impact on long-term predictions for AI modes.

“We see this work as a guide for the community as to how to think about designing experiments for inclusion in machine learning models,” Hoffman said.

Greg Barber, assistant professor of chemistry at Penn State Altoona and affiliate researcher in the Institute of Energy and the Environment and co-author on the paper, also contributed to this research.

This work was supported in part by the U.S. Department of Energy's Office of Science under award number FWP 101064 . Testing equipment was supplied in part by Co-ACCESS, part of the SUNCAT Center for Interface Science and Catalysis, a joint research center supported by SLAC National Accelerator Laboratory and Stanford University. This content is solely the responsibility of the authors and does not necessarily represent the views of the Department of Energy.

Editor’s note: A version of this story originally appeared on the SLAC National Accelerator Laboratory website .

Nature Catalysis

10.1038/s41929-026-01559-y

Experimental study

Not applicable

Quantifying uncertainty in catalyst activity and deactivation during CO2 hydrogenation via round-robin testing for data-driven modelling

31-Jul-2026

Keywords

Article Information

Contact Information

Ty Tkacik
Penn State
tct5204@psu.edu

Source

This article is based on a news release from Penn State. BrightSurf curates and republishes science news from research institutions worldwide; the original release is linked below.

How to Cite This Article

APA:
Penn State. (2026, August 13). Variations between labs can misinform scientific AI models, team reports. Brightsurf News. https://www.brightsurf.com/news/8J4EK5WL/variations-between-labs-can-misinform-scientific-ai-models-team-reports.html
MLA:
"Variations between labs can misinform scientific AI models, team reports." Brightsurf News, Aug. 13 2026, https://www.brightsurf.com/news/8J4EK5WL/variations-between-labs-can-misinform-scientific-ai-models-team-reports.html.