In a study recently published in Nature Communications, researchers from Berlin, Potsdam, and Jena present a new method for analyzing the epigenome, the genome’s control system that determines which genes are switched on and off. The machine-learning method identifies differentially methylated DNA regions without sample labels – a prerequisite for many existing algorithms. This makes it possible to identify previously hidden biological patterns as well as new subgroups of cells or diseases. Analyses of blood cells, pancreatic cancer, and brain tumors confirms known biological relationships and reveals new regulatory processes. The metilene3 method opens up new possibilities for better understanding disease mechanisms and identifying potential biomarkers for diagnostics and research.
The activity of our genes is not determined by DNA sequence alone. The attachment of small chemical compounds – known as methyl groups – influences which genes are active, and which remain silenced. DNA methylation is thus a central component of the epigenome.
Changes to the epigenome play a crucial role in the development of our body, influence our aging process, and are relevant to numerous diseases such as cancer. To understand such changes, researchers specifically search for differentially methylated DNA regions (DMRs). However, existing methods usually require that the samples under investigation be assigned to known groups – such as healthy or diseased tissue. With complex clinical datasets, however, this information is often unknown.
New method detects differences even without known groups
With metilene3, a new software tool, researchers have now developed a method that overcomes this limitation. The software can compare DNA methylation patterns both between predefined groups (supervised mode) and unlabeled samples (unsupervised mode).
In the latter mode, the software searches for DMRs without samples having to be pre-classified into groups such as “healthy” or “diseased.” The software autonomously segments the genome based on methylation signals, grouping the samples fully automatically. This automatic classification makes it possible to visualize epigenetic similarities and developmental relationships between the samples. Previously unknown cell types or disease subgroups can thus be identified, as well as those regions in which the samples both resemble and differ from already known diseases or cell types. At the same time, the biological differences remain traceable, since every similarity and difference can be attributed to specific methylation patterns.
Practical test using real data sets
Afterwards, the researchers tested the method on various biological data sets. Using human blood cells, metilene3 reconstructed the known developmental pathways of various immune cell types based on their DNA methylation alone. In addition, the software identified regulatory DNA regions linked to transcription factors that control the identity of these cell types. The results demonstrate that the method not only distinguishes between cell groups but also identifies functionally relevant regulatory elements.
This approach also proved to be extremely powerful when applied to tumor data. In datasets on glioblastomas, metilene3 identified various molecular subgroups of the tumors and even detected individual samples with unusual biological properties.
In tissue samples from pancreatic cancer, the software was also able to trace the step-by-step progression from healthy tissue through precancerous lesions to the tumor. In the process, the researchers discovered DNA regions in which the binding sites of the transcription factors NF-κB and NFAT occur together particularly frequently. These regions could play an important role in the development of pancreatic cancer and serve as a key starting point for further investigations. However, this biological hypothesis must now be experimentally verified.
“The cancer-related changes in DNA methylation identified by the machine learning method allow us to draw direct conclusions about molecular disruptions in transcription factors,” emphasizes Prof. Helene Kretzmer of the Hasso Plattner Institute in Potsdam. Especially for medical questions, she notes, the requirements for the interpretability of predictions are particularly high.
Our method opens up new possibilities for discovering previously hidden biological relationships and therapeutic approaches.”
Prof. Steve Hoffmann, Leibniz Institute on Aging – Fritz Lipmann Institute (FLI), Jena
New opportunities for research and medicine
With metilene3, researchers now have a tool that significantly expands the analysis of complex DNA methylation data. Particularly in the case of heterogeneous tissue samples or clinical datasets, where biological groups cannot always be defined in advance, the method can help to uncover previously hidden biological relationships. The study’s authors therefore see great potential in the software for researching aging processes, cancer, and other diseases, as well as for identifying new biomarkers. In the future, the tool should also be used for other sequencing technologies and single-cell analyses to make the analysis of epigenetic data even more widely applicable.
“The better we understand epigenetic patterns, the more precisely we can decipher the biological processes behind them. The software provides an important basis for this and could help us to derive new hypotheses for biomedical research from large datasets,” says Kretzmer.
Source:
Journal reference:
Zhu, Z., et al. (2026). metilene3: identifying DMRs across multiple conditions with auto-classification. Nature Communications. DOI: 10.1038/s41467-026-74931-y. https://www.nature.com/articles/s41467-026-74931-y