An international team of researchers, including Earth-Life Science Institute (ELSI) at the Institute of Science Tokyo, has developed a protein language model that brings together two fundamental sources of information about proteins: their amino acid sequences and their three-dimensional structures. The model provides researchers with a new way to map relationships across the protein universe and investigate how proteins have evolved over billions of years.
The research was led by Prof. Rachel Kolodny and PhD candidate Guy Yanai of the University of Haifa, Prof. Nir Ben-Tal and graduate student Gabriel Axel of Tel Aviv University, and Specially Appointed Associate Professor Liam M. Longo of ELSI. Kolodny also spent five months as a visiting researcher at ELSI developing approaches to analyse the new model. The findings were published in Proceedings of the National Academy of Sciences (PNAS).
Thousands of protein families are responsible for carrying out nearly every function within living cells. A fundamental question in evolutionary biochemistry is how these proteins are related to one another and where they came from in the first place.
Scientists traditionally organise proteins into hierarchical groups based on their relatedness, somewhat like the genus and species classifications used for living organisms. These carefully curated systems contain decades of scientific knowledge, but advances in artificial intelligence are now creating new ways of exploring relationships across the vast protein universe.
Protein language models can convert a protein into a numerical representation known as an "embedding". One way to think of an embedding is as a kind of postcode: proteins with similar properties tend to receive nearby addresses. Researchers can then visualise these relationships to produce a "protein world map" (Figure 1).
However, there is a complication. Proteins contain information in both their amino acid sequences and their three-dimensional structures, and the relationship between the two is not straightforward. Proteins with unrelated sequences can sometimes adopt similar structures, while similar or even identical sequences can produce very different structures.
Most protein language models have approached protein sequence and structure separately. Even models that use both kinds of information do not necessarily place the sequence and structure of the same protein at the same location on a protein map (Figure 2).
The researchers developed a model called Contrastive Learning Sequence-Structure, or CLSS, designed to produce highly similar embeddings for both the sequence and structure representations of a protein.
CLSS uses an approach called contrastive learning. During training, the model receives protein sequences and their corresponding structures and learns to produce similar embeddings for sequence-structure pairs while separating unrelated pairs. The result is a shared map in which a protein representation occupies a similar location within the protein world map, regardless of whether its sequence or its structure was used.
When compared with other state-of-the-art protein language models, CLSS successfully brought sequence and structure information together in a cohesive map (Figure 2). Its representations also closely reproduced relationships recorded in the expert-curated ECOD and CATH protein classification systems, even though those classifications were not provided to the model during training.
The model also performed strongly in classification tests, demonstrating that combining sequence and structure information can produce more informative representations of proteins.
"This gives us a way to look at the protein universe through sequence and structure at the same time, rather than treating them as separate worlds," said Longo. "What is particularly exciting for us is the possibility of using these maps to uncover large-scale evolutionary patterns that are difficult to recognise using conventional approaches."
Most existing protein language models require a complete sequence or structure to produce meaningful representations. CLSS showed that, in many cases, short sequence fragments can be positioned meaningfully alongside complete sequences and structures.
Protein fragments are particularly important for understanding evolution. Small pieces of proteins have been repeatedly reused and rearranged throughout evolutionary history, and some may even have served as building blocks for the earliest protein domains. Similar fragments appearing in otherwise different proteins can therefore provide clues to ancient evolutionary relationships.
The maps produced by CLSS also revealed broader patterns across protein space. When the researchers overlaid biological properties onto the maps, for example, proteins associated with organic cofactors were concentrated in particular regions, whereas metal-binding proteins were more widely distributed.
Such patterns illustrate how global protein maps can be used not only to classify proteins but also to explore relationships between their sequence, structure, function, and evolutionary history.
Ultimately, the researchers envision unified sequence-structure representations opening new possibilities for database searches, protein engineering, and the reconstruction of evolutionary trajectories. By bringing different kinds of biological information into the same map, CLSS offers another way to explore how the diversity of proteins found in life today emerged over nearly four billion years of evolution.
Reference
Guy Yanai a , Gabriel Axel b , Liam M. Longo c,d,* , Nir Ben-Tal b,* , and Rachel Kolodny a,* , Contrastive learning unites sequence and structure in a global representation of protein space, Proceedings of the National Academy of Sciences of the United States of America , DOI: 10.1073/pnas.2532702123
a. Department of Computer Science, University of Haifa, Haifa 3303221, Israel
b. Department of Biochemistry and Molecular Biology, School of Neurobiology, Biochemistry and Biophysics, George S. Wise Faculty of Life Sciences, Tel Aviv University, Ramat Aviv 6997801, Israel
c. Earth-Life Science Institute, Institute of Science Tokyo, Tokyo 152-8550, Japan
d. Blue Marble Space Institute of Science, Seattle, WA 98104
More information
Earth-Life Science Institute (ELSI) is one of Japan’s ambitious World Premiere International research centers, whose aim is to achieve progress in broadly inter-disciplinary scientific areas by inspiring the world’s greatest minds to come to Japan and collaborate on the most challenging scientific problems. ELSI’s primary aim is to address the origin and co-evolution of the Earth and life.
Institute of Science Tokyo (Science Tokyo) was established on October 1, 2024, following the merger between Tokyo Medical and Dental University (TMDU) and Tokyo Institute of Technology (Tokyo Tech), with the mission of “Advancing science and human wellbeing to create value for and with society.”
World Premier International Research Center Initiative (WPI) was launched in 2007 by Japan's Ministry of Education, Culture, Sports, Science and Technology (MEXT) to foster globally visible research centers boasting the highest standards and outstanding research environments. Numbering more than a dozen and operating at institutions throughout the country, these centers are given a high degree of autonomy, allowing them to engage in innovative modes of management and research. The program is administered by the Japan Society for the Promotion of Science (JSPS).
Proceedings of the National Academy of Sciences
Data/statistical analysis
Not applicable
Contrastive learning unites sequence and structure in a global representation of protein space
3-Aug-2026