Medical researchers, clinicians and pharmaceutical scientists often study small molecules called metabolites, created in living cells for variety of reasons, for example when they break down food, chemicals or drugs through the metabolic process. This area of study is known as metabolomics and it can be used to identify new signs of disease, measure the efficiency of medical treatment or study how diet and nutrition affect the body.
While many of these molecules can be identified through the use of mass spectrometry, a lab test that measures the weight and charge of tiny particles, the vast majority of molecules remain unknown. These unidentified particles are referred to as “dark matter”, or the “dark metabolome”.
Dark matter, while detectable, can’t be matched to known molecular structures. These unknown particles leave behind a mass of data that is inaccessible for biological interpretation, creating a gap in knowledge of the genetic biome, cellular health, diseases and drug response.
But a new tool, the Molecular Community Network (MCN), developed by Professor Vladimir Boginski and an interdisciplinary team of researchers, could illuminate the identity of molecules in the “dark matter”. Their work has been published in the most recent issue of Cell Reports Methods .
“There are roughly 8.4 million observed mass spectra of known and unknown molecules combined in public repositories, and it is estimated that the dark matter comprises up to 90% of the entire observed molecular space,” Boginski says. “Thus, it is crucial to develop methods that would allow one to systematically investigate this vast molecular space and potentially discover new molecules.”
Traditional molecular networking does not resolve the problem. It connects molecules only when their similarity scores, calculated from observed mass spectra, exceed a predetermined similarity threshold. This means that molecules that are biologically related but fall below the calculated cutoff can be separated from one another, resulting in fragmented molecular families and limiting the chance for molecular discovery.
Unlike these existing networking methods, the MCN uses an algorithm that divides the entire molecular network into natural communities with a vast amount of strong links within groups rather than between them. At that point, the strongest connections are retained to keep the communities connected. Rather than breaking molecular families into pieces, this process reveals the data structure that’s already present.
“This approach allows us to have almost every molecule in the network linked to at least one neighbor,” Boginski says. “Moreover, these links are typically between molecules from similar molecular families.”
This, in turn, facilitates the process of annotation propagation, or the prediction of the identity of unknown molecules by looking at their known neighbors within a network community.
“Since nearly 95% of molecules are now connected and assigned to network communities, we now have a much wider and richer search space for molecular discovery,” Boginski says. “We have shown that this approach can indeed find previously unknown molecules based on their positions in the molecular community network.”
With the aid of the MCN, Boginski and his team of co-authors already discovered a new class of bile acids. These particular bile acids are made by gut microbes and one of them only appears in early infants. Their discovery was first predicted by the MCN and then confirmed by the researchers in the lab.
Boginski says that this discovery is just the tip of the iceberg, with the potential for many similar discoveries with the MCN.
“Biomarker discovery often stalls at the point where a molecule that differentiates sick from healthy is detected by its mass spectrum, but it may not be known what this molecule is,” he says. “MCN can potentially help convert some of those ‘dead ends’ into identifiable molecules, and because this method runs on mass spectrum data that has already been collected, roughly 8.4 million molecules in public repositories are now mapped and open to reanalysis, increasing the potential for new biomarker discovery.”
Cell Reports Methods
Ordering molecular diversity in untargeted metabolomics via molecular community networking
20-Jul-2026