A multimodal large language model framework converts scattered charts, reports, and regulatory records into a comprehensive dataset for assessing plastic chemical production, hazards, environmental behavior, and oversight
More than 16,000 chemicals are known or suspected to occur in plastic materials and products, yet essential information about many of them remains fragmented across scientific papers, market reports, images, databases, and regulatory documents. This lack of accessible data makes it difficult for scientists and policymakers to identify high-priority substances and manage chemical risks associated with plastics.
Researchers from Dalian University of Technology have now developed an artificial intelligence framework that can extract and organize large amounts of information about plastic chemicals from both text and images. The system achieved an average F1 score of 92.8%, indicating a high level of accuracy and completeness in the extracted data.
“ Managing plastic pollution requires more than knowing which chemicals are present. We also need to understand how much is produced and traded, how these substances behave in the environment, what hazards they may pose, and whether they are adequately regulated, ” said corresponding author Jingwen Chen. “ Our framework brings these previously disconnected forms of information together in a structured and scalable way. ”
The researchers used the open-source multimodal models Qwen2-VL-72B and Qwen2.5-32B to analyze 36,289 online images and 969 literature sources . These materials included charts, figures, industry reports, and scientific publications in both Chinese and English.
The AI pipeline extracted 44,249 records covering 814 plastic chemicals and nine economic indicators , including production volume, production capacity, market size, consumption, and import and export volumes and values. Each record was organized with contextual details such as time, location, units, and data source.
The team then combined these results with manually curated records and computational predictions. The completed dataset includes toxicity information for 4,551 chemicals, environmental behavior data for approximately 9,770 chemicals, and regulatory information for 13,684 chemicals across 111 inventories and regulatory lists from 26 countries and regions. Graph attention network models were also used to predict physicochemical and hazard-related properties for more than 9,000 substances.
The analysis identified 2,144 chemicals with persistence, bioaccumulation, and toxicity characteristics and 3,951 chemicals with persistence, mobility, and toxicity characteristics. Many had not previously been classified in these categories.
The regulatory comparison also revealed important gaps. Of the plastic chemicals examined, 7,024 were absent from all regulatory inventories included in the study. Among them, 1,958 showed predicted PBT or PMT properties. Six unregulated toxic chemicals also exceeded the annual production threshold commonly used to define high-production-volume substances.
The dataset could support material flow analysis, exposure modeling, chemical prioritization, regulatory screening, and the development of safer alternatives. As a demonstration, the researchers used the economic data to examine flows of polytetrafluoroethylene, or PTFE, in China, showing increasing stocks and waste generation between 2001 and 2025.
The authors emphasize that economic data can vary across years, regions, and reporting systems, and the dataset should therefore be treated as a transparent baseline rather than a definitive record of global chemical use.
The study demonstrates how multimodal AI can transform scattered and unstructured information into usable evidence for chemical management. The researchers suggest that future AI agents could automate the full process, from data collection and extraction to analysis and reporting, helping governments and industries respond more effectively to the chemical dimensions of plastic pollution.
===
Journal reference: Cui Y; Zhang Y; Wang H; et al. Multimodal large language models for big data synthesis: Addressing the data bottleneck of plastic chemicals. AI Environ. 2026, 1(2): 83-92. DOI: 10.66178/aie-0026-0011
https://www.the-newpress.com/aie/article/doi/10.66178/aie-0026-0011
About the Journal:
Artificial Intelligence & Environment is an international multidisciplinary platform for communicating advances in fundamental and applied research on the intersection of environmental science and artificial intelligence (AI). It is dedicated to serving as an innovative, efficient and professional platform for researchers in the cross-discipline fields of earth and environmental sciences, big data science and AI around the world to deliver findings from this rapidly expanding field of science. It is a peer-reviewed, open-access journal that publishes critical review, original research, rapid communication, view-point, commentary and perspective papers.
Follow us on Facebook , X , and Bluesky .
Experimental study
Multimodal large language models for big data synthesis: Addressing the data bottleneck of plastic chemicals
15-Jun-2026