HSE Researchers Train Neural Network to Predict Protein–Protein Interactions More Accurately

Scientists at the AI and Digital Science Institute of the HSE Faculty of Computer Science have developed a model capable of predicting protein–protein interactions with 95% accuracy. GSMFormer-PPI integrates three types of protein data (including information about protein surface properties) to analyse relationships between proteins, rather than simply combining datasets as in previous models. The solution could accelerate the discovery of disease molecular mechanisms, biomarkers, and potential therapeutic targets. The paper has been published in Scientific Reports.
Almost all cellular processes depend on interactions between proteins. Cells use these interactions to transmit signals, initiate and regulate chemical reactions, and form molecular complexes essential for proper functioning. When such interactions are disrupted, cellular processes can malfunction, potentially leading to disease.
Therefore, to study disease mechanisms and identify therapeutic targets, it is important for scientists to understand which proteins can interact and which cannot. Determining this experimentally is difficult: when dozens or hundreds of proteins are considered, the number of possible pairs becomes too large to test individually. As a result, biologists use machine learning methods to predict these interactions based on the structure and properties of molecules.
HSE researchers have developed the GSMFormer-PPI system, which takes into account three types of data for each protein in a candidate pair: the amino acid sequence, the three-dimensional structure, and the properties of the molecular surface. To process this information, the authors used existing models that convert this data into numerical representations. A protein language model analyses the amino acid sequence—the order of amino acids that make up the protein. The three-dimensional structure of the protein is represented as a graph, in which amino acids are treated as nodes and their spatial contacts as edges; this representation is processed by a graph neural network. In addition, a separate algorithm captures protein surface properties—the shape and physicochemical characteristics of the regions through which proteins recognise one another.
These numerical representations of proteins were then fed into a transformer module developed by the authors—a neural network that jointly analyses different types of protein data. In contrast to many previous approaches, where features were often simply concatenated into a single vector, this model does not combine them mechanically but instead captures the relationships between them.
Maria Poptsova
'When proteins interact, their surface is particularly important: it is through the surface that molecules recognise one another, and it is where the physicochemical properties that determine binding are concentrated. In our model, we sought to incorporate this information alongside the protein’s sequence and three-dimensional structure and not merely concatenate these features but enable the algorithm to analyse the relationships between them. This is what allowed us to predict protein–protein interactions more accurately,' comments one of the authors, Maria Poptsova, Director of the Centre for Biomedical Research and Technology at the HSE FCS AI and Digital Science Institute.

The researchers tested the new model’s performance on the PINDER dataset, a large database of known protein interactions. In these experiments, GSMFormer-PPI achieved an accuracy of 95.7%, outperforming popular graph-based models such as GCN and GAT. The researchers also tested a simpler version of GSMFormer-PPI—without the module that analyses relationships between different types of data. This version performed worse, demonstrating that it is not only the protein data itself but also how the model integrates and compares it that drives its accuracy.
Additional tests showed that all three types of data—sequence, spatial structure, and surface properties—are essential for accurate predictions. When the researchers removed any one component, prediction accuracy declined. In other words, the model performs better precisely because it considers the protein on multiple levels simultaneously. In the future, such systems could help identify protein pairs more efficiently when studying disease mechanisms and searching for drug targets.
The work was supported by a grant for research centres in AI provided by the Ministry of Economic Development of the Russian Federation and implemented at HSE University.
See also:
Scientists Train Neural Network to Generate Process Plans from 3D Models
Researchers at the HSE FCS AI and Digital Science Institute have developed CAD2TechSpec, a framework that converts 3D models of mechanical parts into machining process plans—step-by-step instructions for machine tools. The solution aims to reduce the time required for the design and preparation of technical process documentation in mechanical engineering, aircraft manufacturing, and other high-tech industries. The study findings have been published in PeerJ Computer Science.
Biologists Discover 'Molecular Fingerprint' of Preeclampsia
Researchers at HSE University employed a new method to model hypoxia in placental cells during pregnancies complicated by preeclampsia and identified molecular markers of tissue hypoxia. Since hypoxia is one of the key mechanisms underlying preeclampsia, these findings are important for a more accurate and timely diagnosis of the disease and for the development of effective treatment methods. The paper has been published in Placenta.
‘Hedgehog’ Versus ‘Relatives’: Researchers Measure How the Brain Responds to Unexpected Words During Natural Speech
Russian neurophysiologists, including researchers from HSE University, have demonstrated the feasibility of using event-related fields (ERFs) to study brain activity during natural speech perception. The researchers showed that this approach can be applied not only to individual words but also to continuous speech. Their findings indicate that words whose meanings differ significantly from the preceding context require longer processing times. The study also reveals that the brain processes function words in two stages: first, it identifies their grammatical role and then uses this information to predict the next word. The study has been published in Frontiers in Human Neuroscience.
HSE Researchers Create New Corpus of Early Child Speech in Russian
Researchers at the HSE Centre for Language and Brain have presented RusLan-M, an open multimedia corpus that makes it possible to trace the development of early child speech in Russian from first words to the emergence of complex grammatical constructions. The database contains around 41 hours of video recordings and more than 35,000 child utterances. The new resource will help researchers study more precisely how children acquire Russian and, in the longer term, develop more reliable tools for assessing speech development. The study has been published in Language Resources and Evaluation.
Hybrid Intelligence: Competencies in the Age of AI Discussed at Technoprom-2026
Artificial intelligence is not creating new professions, but rather transforming the nature of existing ones. This was the conclusion reached by participants in the panel session ‘Hybrid Intelligence: Digital and Human Drivers of Development,’ organised by the Institute for Statistical Studies and Economics of Knowledge (ISSEK) at HSE University as part of the 13th International Forum of Technological Development (Technoprom-2026). The experts discussed how the nature of work is changing, which skills are becoming increasingly sought after, and what prevents companies from fully capitalising on new technologies.
Scientists Develop New Solution for 6G Communication Systems
A terahertz neuromorphic circuit developed by scientists at HSE University could make 6G communication systems both more accurate and energy-efficient. The circuit enables indoor tracking of mobile devices with an accuracy of up to 99%. The results were presented at PIERS 2026, an international symposium on Photonics and Electromagnetism held in China.
Scientists Develop Algorithm for More Reliable Processors in Data Centres
Researchers from HSE MIEM and Samara University have developed the LRF-3D algorithm to automatically bypass idle nodes in three-dimensional networks-on-chip. Thanks to its hierarchical architecture, the algorithm outperforms existing solutions in both speed and path accuracy, improving processor reliability for use in data centres, supercomputers, and AI computing. The source code and test results are publicly available.
Researchers Rank Recommendation Algorithms Using Sports Tournament Model
Researchers from the AI and Digital Science Institute at the HSE Faculty of Computer Science have developed an approach for selecting recommendation algorithms more effectively. Their approach uses pairwise comparisons of algorithms to create a tournament table, with the overall ranking based on their performance across all datasets in the tournament. This can reduce the number of algorithms that need to be tested when developing new services, saving both time and money. The study was presented at the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026).
Researchers Develop Method for Direct Generation of Regulatory DNA
Researchers at HSE University have developed a model for generating promoters and enhancers—DNA sequences that regulate gene activity. The model works directly with DNA nucleotides, without first transforming them into a continuous numerical representation. This solution could be useful for applications in synthetic biology and gene therapy. The study results were presented at the ICLR 2026 Workshop ‘Generative AI in Genomics (Gen^2): Barriers and Frontiers.’
Researchers at HSE University and Sber Train Neural Networks to Better Predict User Preferences
The HSE FCS AI and Digital Science Institute and Sber have introduced a new architecture for recommendation systems that combines two classes of models, enabling algorithms to better predict users’ interests and needs. A preprint of the paper has been published on arxiv.org and presented at Urban ML.


