An international team of researchers, including Earth-Life Science Institute (ELSI) at the Institute of Science Tokyo, has developed a protein language model that brings together two fundamental sources of information about proteins: their amino acid sequences and their three-dimensional structures. The model provides researchers with a new way to map relationships across the protein universe and investigate how proteins have evolved over billions of years.
The research was led by Prof. Rachel Kolodny and PhD candidate Guy Yanai of the University of Haifa, Prof. Nir Ben-Tal and graduate student Gabriel Axel of Tel Aviv University, and Specially Appointed Associate Professor Liam M. Longo of ELSI. Kolodny also spent five months as a visiting researcher at ELSI developing approaches to analyze the new model. The findings were published in Proceedings of the National Academy of Sciences (PNAS).
Thousands of protein families are responsible for carrying out nearly every function within living cells. A fundamental question in evolutionary biochemistry is how these proteins are related to one another and where they came from in the first place.
Scientists traditionally organise proteins into hierarchical groups based on their relatedness, somewhat like the genus and species classifications used for living organisms. These carefully curated systems contain decades of scientific knowledge, but advances in artificial intelligence are now creating new ways of exploring relationships across the vast protein universe.
Protein language models can convert a protein into a numerical representation known as an “embedding”. One way to think of an embedding is as a kind of postcode: proteins with similar properties tend to receive nearby addresses. Researchers can then visualise these relationships to produce a “protein world map”.
However, there is a complication. Proteins contain information in both their amino acid sequences and their three-dimensional structures, and the relationship between the two is not straightforward. Proteins with unrelated sequences can sometimes adopt similar structures, while similar or even identical sequences can produce very different structures.
Most protein language models have approached protein sequence and structure separately. Even models that use both kinds of information do not necessarily place the sequence and structure of the same protein at the same location on a protein map.
The researchers developed a model called Contrastive Learning Sequence-Structure, or CLSS, designed to produce highly similar embeddings for both the sequence and structure representations of a protein.
CLSS uses an approach called contrastive learning. During training, the model receives protein sequences and their corresponding structures and learns to produce similar embeddings for sequence-structure pairs while separating unrelated pairs. The result is a shared map in which a protein representation occupies a similar location within the protein world map, regardless of whether its sequence or its structure was used.
When compared with other state-of-the-art protein language models, CLSS successfully brought sequence and structure information together in a cohesive map. Its representations also closely reproduced relationships recorded in the expert-curated ECOD and CATH protein classification systems, even though those classifications were not provided to the model during training.
The model also performed strongly in classification tests, demonstrating that combining sequence and structure information can produce more informative representations of proteins.
This gives us a way to look at the protein universe through sequence and structure at the same time, rather than treating them as separate worlds. What is particularly exciting for us is the possibility of using these maps to uncover large-scale evolutionary patterns that are difficult to recognize using conventional approaches.”
Liam M. Longo, Specially Appointed Associate Professor, ELSI
Most existing protein language models require a complete sequence or structure to produce meaningful representations. CLSS showed that, in many cases, short sequence fragments can be positioned meaningfully alongside complete sequences and structures.
Protein fragments are particularly important for understanding evolution. Small pieces of proteins have been repeatedly reused and rearranged throughout evolutionary history, and some may even have served as building blocks for the earliest protein domains. Similar fragments appearing in otherwise different proteins can therefore provide clues to ancient evolutionary relationships.
The maps produced by CLSS also revealed broader patterns across protein space. When the researchers overlaid biological properties onto the maps, for example, proteins associated with organic cofactors were concentrated in particular regions, whereas metal-binding proteins were more widely distributed.
Such patterns illustrate how global protein maps can be used not only to classify proteins but also to explore relationships between their sequence, structure, function, and evolutionary history.
Ultimately, the researchers envision unified sequence-structure representations opening new possibilities for database searches, protein engineering, and the reconstruction of evolutionary trajectories. By bringing different kinds of biological information into the same map, CLSS offers another way to explore how the diversity of proteins found in life today emerged over nearly four billion years of evolution.
Source:
Journal reference:
Yanai, G., et al. (2026). Contrastive learning unites sequence and structure in a global representation of protein space. Proceedings of the National Academy of Sciences. DOI: 10.1073/pnas.2532702123. https://www.pnas.org/doi/10.1073/pnas.2532702123