Existing machine learning (ML) models deployed in wastewater treatment plants (WWTPs) typically require extensive data collection for training or continuous data acquisition for model updates, posing the dual limitations of cost and operational maintenance.
In a new study published in Water & Ecology , a research team led by Hong-Cheng Wang from Harbin Institute of Technology, Shenzhen, challenged the prevailing assumption that bigger datasets always produce better predictions. The study provides a transferable strategy for data-efficient machine learning in environmental engineering, demonstrating that intelligent data selection can match the performance of exhaustive datasets at a fraction of the resource cost.
The team developed a three-step framework that distills massive water quality monitoring datasets into a lean, representative core. First, Koopman analysis—an approach for identifying characteristics of nonlinear dynamic systems—is applied to sub-datasets to extract Koopman matrices and their singular values, which encode the principal dynamic features. Second, density-based spatial clustering (DBSCAN) groups these features into distinct dynamic regimes. Third, representative data from each regime are used to train machine learning architectures including BiLSTM, Transformer, and XGBoost for effluent quality prediction.
Applied to a full-scale WWTP in southern China, the method partitioned 75 sub-datasets into 36 distinct dynamic categories. The resulting representative dataset comprised only 66.41% of the original data—a 43.59% reduction, yet yielded models that performed statistically similarly to those trained on the full dataset. In contrast, models trained on conventional scenario-classified datasets showed significantly higher prediction errors and divergent error distributions.
“Our framework quantifies exactly how much data can be reduced and explains why such a reduction does not harm model performance,” says Wang. “The Koopman singular values essentially cluster the water quality dynamics, so selecting representatives from each category captures all necessary system behaviors.”
The research further introduces a Low Data Requirement (LDR) operational framework for WWTPs. “Under this scheme, incoming monitoring data are screened against existing representative datasets; only data representing novel dynamic patterns are retained for model updates,” says Wang.
Notably, the strategy reduces storage burdens and energy consumption while preserving modeling accuracy.
“From an engineering standpoint, this is about mining the right data, not hoarding all of it,” adds Wang. “The framework is particularly valuable for facilities with high data acquisition costs or limited historical records, and its underlying Koopman approach is generalizable to other water systems and environmental infrastructures.”
Nevertheless, the team acknowledged that the method requires consistent underlying dynamic principles across datasets and careful selection of Koopman matrix dimensions to avoid overfitting.
###
Contact the author:
Hong-Cheng Wang
-State Key Laboratory of Urban Water Resource and Environment, School of Eco-Environment, Harbin Institute of Technology, Shenzhen 518055, China
wanghongcheng@hit.edu.cn
Yu-Qi Wang
17877784587@163.com
The publisher KeAi was established by Elsevier and China Science Publishing & Media Ltd to unfold quality research globally. In 2013, our focus shifted to open access publishing. We now proudly publish more than 200 world-class, open access, English language journals, spanning all scientific disciplines. Many of these are titles we publish in partnership with prestigious societies and academic institutions, such as the National Natural Science Foundation of China (NSFC).
Water & Ecology
Computational simulation/modeling
Not applicable
Panning for Gold: Finding Representative Water Quality Dynamic Data for Data-Efficient Machine Learning through Koopman Analysis
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.