• Illumina Connected Multiomics
  • News
  • 09/10/2026

Turn 5-base DNA methylation data into clearer biological hypotheses

How QC, DMR analysis, pathway interpretation, and variant evidence can help turn 5-base DNA methylation data into clearer biological hypotheses.

The Illumina 5-base DNA Prep kit measures the four standard DNA bases along with methylated cytosine, enabling researchers to profile DNA methylation while also capturing DNA variant information from a single sequencing workflow. Once the data is generated, however, researchers often face the harder question: which methylation changes are likely to inform interpretation, and which are simply part of a long results table?

Illumina Connected Multiomics analysis software provides a collaborative environment where each step of a multiomic analysis can stay connected within a single workflow. This helps researchers and collaborators build, revisit, and refine the biological story as the analysis progresses. Designed to support Illumina multiomic assay innovations, it provides the ideal exploratory analysis solution for the 5-base DNA methylation assay.

The workflow below shows how researchers can move from raw methylation calls to candidate biological signals using a connected analysis approach in Connected Multiomics.

How to analyze 5-base DNA methylation data: an AML example

Throughout this blog, we use a human acute myeloid leukemia (AML) cohort with 24 samples: three healthy controls and 21 samples representing different molecular subtypes. The samples are processed with the Illumina 5-base DNA Prep kit, and 5-base DNA methylation calls are generated using a DRAGEN 5-base pipeline. Readers can explore the example dataset in a Connected Multiomics trial to see how the concepts discussed here appear in the software.

The analysis goal is to compare methylation patterns across the AML subtypes, identify differentially methylated regions, interpret the affected biology, and integrate those results with DNA variants to prioritize candidate genes. The analysis workflow has five core steps: assess data quality, explore sample structure, identify DMRs, add biological context, and integrate methylation results with DNA variant calls. Each step answers a specific question, helping you understand the data structure and prioritize meaningful biological signals.

Step 1: Assess data quality before downstream analysis

Low-quality samples can introduce noise into every downstream step, so start by reviewing sample quality metrics such as sequencing quality, methylation calling quality, and M-bias.

Look for patterns across all samples in the dataset. For sequencing quality and methylation calling metrics, lower duplicate rates, higher mapping rates, and stronger coverage generally support more reliable interpretation, while spike-in controls help confirm that conversion and methylation calling are performed as expected.

Figure 1: QC metrics plots help identify samples that may need closer inspection before downstream DMR analysis. Each data point represents a sample.

Use the M-bias plot to check whether methylation levels stay relatively consistent across genomic positions on reads. Some unevenness at the first and last 10 bases can be expected because of sequencing artifacts. In the AML dataset, one sample with M-bias and poorer sequencing or methylation calling metrics is removed before statistical testing to reduce downstream noise.

Figure 2: M-bias plots visualize methylation levels and coverage across positions on sequencing reads. Each line represents a sample.

Why it matters: QC decisions shape every result that follows. Reviewing quality early helps explain why samples are kept or removed and gives downstream DMR analysis a stronger foundation.

Step 2: Explore sample structure before statistical testing

After QC, explore how the high-quality samples relate to one another before statistical testing. Dimensionality reduction can reveal clustering patterns, potential outliers, and variables that may be driving methylation differences.

Start with principal component analysis (PCA) to identify the major sources of variation across samples. In the AML example, PCA performed on the most variable CpG sites does not clearly separate molecular subtypes, suggesting that subtype-specific methylation differences are relatively subtle and do not drive the largest sources of variation in the dataset. This is why PCA is typically followed by differential methylation analysis, where statistical testing is used to identify methylation changes associated with specific subtypes.

To further explore sample relationships, PCA can also be performed on aggregated genomic features, such as promoter regions, or disease-relevant regions, rather than individual CpG sites. This regional view may better capture the underlying methylation differences across subtypes and reveal biologically meaningful clustering patterns.

Uniform Manifold Approximation and Projection (UMAP) can then be used as a complementary dimensionality reduction approach to visualize local sample relationships that may not be apparent in PCA. Because UMAP is sensitive to parameter selection, it is best used as an exploratory tool, with careful parameter tuning and iterative evaluation to ensure that the observed patterns are meaningful.

Why it matters: Exploratory analysis lays the foundation for differential methylation analysis by helping you identify relevant comparisons and interpret DMR results in context, rather than treating statistical findings as a black box.

Step 3: Identify differentially methylated regions

After you understand the sample structure, move into differential methylation analysis to identify genomic regions where methylation levels differ between biological conditions. These differentially methylated regions, or DMRs, may be hypomethylated or hypermethylated in one group relative to another.

In the AML example, the analysis compares DNMT3A R882-mutant samples with KMT2A-rearranged samples. These AML subtypes are known to have distinct methylation patterns, with DNMT3A-mutant samples typically showing widespread hypomethylation due to impaired methyltransferase activity.

Begin by reviewing the DMR report, focusing on both the magnitude of methylation differences and their statistical significance. Negative methylation difference values indicate regions that are hypomethylated in the DNMT3A group relative to the KMT2A group. To help prioritize the most relevant findings, visualize the DMRs using volcano and Manhattan plots.

The volcano plot helps narrow a long list of DMRs to regions showing both large effect sizes and strong evidence of differential methylation. The Manhattan plot provides a genomic view of DMR distribution, highlighting whether DMR signals cluster in specific genomic regions or appear across the genome.

In this AML analysis, a larger number of hypomethylated DMRs are observed in the DNMT3A-mutant samples, consistent with the known biology of this AML subtype, and the broad distribution of DMR signals across multiple chromosomes suggests genome-wide methylation changes rather than a single genomic hotspot. This agreement between the statistical results and the expected disease biology increases confidence that the identified DMRs reflect meaningful subtype-associated epigenetic differences.

Why it matters: DMR analysis can produce long result lists. The value comes from finding regions that are statistically meaningful, biologically relevant, and aligned with the study question.

Figure 3: Volcano plots help prioritize DMRs by effect size and statistical significance.
Figure 4: Manhattan plots show where DMRs occur across the genome.

Step 4: Add biological context

Once you have a focused set of candidate DMRs, add biological context. Annotation links DMRs to nearby or overlapping genes and helps determine whether a DMR falls into a promoter, gene body, or intergenic region. This information helps prioritize DMRs that may have functional relevance and provides a starting point for understanding how methylation changes could influence gene regulation.

Pathway analysis can then show whether DMR-associated genes converge on biological processes more often than expected by chance. In the AML example, this helps assess whether subtype-associated DMRs point to pathways or functions relevant to AML biology.

For deeper interpretation, look at pathway diagrams to see where DMR-associated genes fall within a pathway. When multiple genes within the same pathway show methylation changes, that pattern may support a hypothesis about coordinated pathway involvement in disease.

Why it matters: Significant regions are only the starting point. By linking DMRs to genes, pathways, and curated biological knowledge, researchers can move from “this region changed” to “this change may point to a relevant biological process.”

Step 5: Integrate methylation and variant evidence

The final step is to connect methylation results with DNA variant calls. Because a 5-base DNA methylation workflow can generate both methylation and small variant information from the same sequencing data, researchers can examine these molecular layers together rather than in isolation. By identifying genes that harbor both DNA variants and nearby methylation changes, it becomes possible to prioritize disease-relevant candidate genes supported by multiple lines of evidence.

Before integration, filter and annotate variants according to the study objectives, using criteria such as variant quality, read depth, population frequency, predicted functional impact, or sample group membership. Then use the same gene annotation model for both DMRs and variants so that methylation and variant findings can be mapped to genes in a consistent, reproducible way.

In the AML example, intersecting annotated DMRs with annotated variants highlights genes supported by both epigenetic and genetic alterations, helping to refine a long list of findings into a focused set of biological candidates for further investigation and validation.

Why it matters: Integrated analysis can sharpen candidate prioritization. When methylation and variant evidence converge on the same gene or pathway, researchers gain greater confidence that the identified signals are biologically relevant and worthy of downstream follow-up.

From complex 5-base data to clearer biological hypotheses

A complete 5-base DNA methylation analysis works best when each step builds on the last. By moving from QC to exploratory analysis, DMR detection, biological interpretation, and variant integration, researchers can narrow complex methylation data into a focused set of candidate regions, pathways, and genes. In the AML example, this approach connects subtype-associated methylation changes with biological context and supporting variant evidence, creating a clearer path from data review to hypothesis generation.

See the workflow in action, then try it yourself

Ready to apply the workflow to your own study? Illumina Connected Multiomics helps bring QC, visualization, DMR detection, pathway interpretation, and variant evidence into a connected multiomic analysis experience, so you can spend less time reconciling outputs and more time evaluating which biological signals deserve follow-up.

Continue learning, then apply the workflow

See how Illumina Connected Multiomics can support your research goals.

---

For Research Use Only. Not for use in diagnostic procedures.
M-GL-04758