- Home
- News & Updates
- From FFPE Noise to Somatic Signal: Insights from Broad's Evaluation of DRAGEN-powered WES Workflows
From FFPE Noise to Somatic Signal: Insights from Broad's Evaluation of DRAGEN-powered WES Workflows
A workflow comparison showing how modeling FFPE noise during variant calling, using sample-specific error profiles and contextual variant evaluation, improves variant calling outcomes
Overview
In a recent study published in npj Genomic Medicine, researchers from Broad Clinical Labs, MD Anderson Cancer Center, Dana-Farber Cancer Institute, Massachusetts General Hospital, and Harvard Medical School compared an updated somatic whole-exome sequencing workflow built on DRAGEN™ to their prior multi-tool pipeline using standardized benchmarking samples and FFPE tumor cohorts. In the study, the DRAGEN-based workflow was associated with improved variant calling results, including lower false positive rates and higher analytical sensitivity in InDel detection in benchmarking scenarios without matched normals, alongside strong performance in tumor–normal workflows. These improvements were supported by an approach that addresses FFPE artifacts during variant calling through sample-specific error modeling and contextual evaluation of candidate variants, along with a lab-matched systematic noise model.
FFPE samples are inherently challenging. The fixation process introduces DNA damage and artifacts that appear as false mutations, making it difficult to distinguish true somatic variants. This challenge varies across samples depending on fixation conditions, block age, and tissue type. Conventional pipelines manage this complexity by combining tools such as BWA-mem, MuTect, and Strelka2, followed by post hoc filtering steps like panels of normals and base orientation bias filters. While these reduce noise, they do not address the root issue: artifacts are introduced early and handled only after variant calling, limiting the effectiveness of static, population-level filters.
In contrast, Broad's updated workflow reflects a shift in how FFPE noise is handled. By modeling FFPE-associated artifacts within the variant calling process, DRAGEN enables improved resolution of challenging genomic regions, supports tumor-only analysis in scenarios where matched normals are not available, and adapts to lab-specific noise profiles.
In this article, we explore the challenges FFPE introduces in variant calling, how DRAGEN addresses them, and what Broad’s validation means for laboratories working with similar tumor samples.
FFPE is embedded in standard pathology practice
Formalin-fixing tumor tissues is common practice in pathology because it preserves tissue morphology, enables long-term storage at room temperature, and supports a range of downstream molecular analyses. Approximately 1 billion FFPE blocks are stored in hospitals and tissue banks worldwide [3, 4], representing an enormous upstream reservoir of annotated tumor material available for retrospective and prospective profiling. As NGS methods become increasingly capable of extracting usable molecular data from these samples, reliably analyzing FFPE material is a core operational requirement for sequencing laboratories.
While FFPE is the standard sample type for oncology research, formalin fixation presents major challenges to NGS analysis.
DRAGEN is purpose-built for FFPE
Conventional somatic variant calling bioinformatics pipeline treat FFPE artifacts as a filtering problem. They rely on a panel of normals (PoN), flag strand-biased calls, and remove any known artifact contexts post-hoc. These steps reduce noise at the margins but do not address the fundamental issue - Artifacts enter the pipeline early and are managed only after variant calling.
Broad’s V6 bioinformatics workflow uses DRAGEN v4.3.11, which provides FFPE-aware variant calling capabilities including sample-specific nucleotide error modeling, local haplotype assembly, and lab-matched systematic noise filtering. Rather than filtering FFPE noise after calling, DRAGEN models FFPE-associated artifacts within the variant calling process.
DRAGEN handles FFPE-aware variant calling at three levels, each addressing a specific failure mode that conventional pipelines leave unresolved:
Table 1: Showing challenges in conventional FFPE bioinformatics workflows versus how DRAGEN handles them
Notes:
Parameter values shown reflect Broad’s validated configuration (DRAGEN v4.3.11) and may differ from newer DRAGEN defaults. See the user guide for further details.
Rather than forcing a tradeoff between analytical sensitivity and specificity, DRAGEN is designed to balance both by modeling noise at the sample level and evaluating variants in haplotype context.
What does this mean for your laboratory? Your bioinformatics and interpretation team receives a cleaner VCF from the start that can be routed directly for tertiary analysis without additional filtering steps. Fewer FFPE-induced false positives to triage means faster time to reporting, reduced computational overhead from maintaining post-hoc filtering pipelines, and lower server and storage costs associated with processing bloated variant call sets.
Your team’s effort can focus on uncovering the somatic biology behind the candidate variants, identifying biologically relevant events, and driving biomarker discovery with meaningful research impact.
How Broad Solved their FFPE Challenge with DRAGEN
Broad Clinical Labs operates a large-scale genomics research program that utilizes FFPE-derived tumor samples as a primary input. Their datasets include diverse cohorts such as glioblastoma (GBM), breast cancer (BrCa), chronic lymphocytic leukemia (CLL), and angiosarcoma. These samples reflect the full spectrum of fixation conditions, block age, and tumor cellularity typically observed in FFPE-derived sequencing data.
During the development of V6 workflow, an updated somatic whole-exome sequencing (WES) workflow the limitations of the prior workflow (V2) became apparent. The V2 assay used the Illumina Nextera Rapid Capture Exome v1.2 kit with 76 bp paired-end sequencing on HiSeq 2000/2500, while V6 assay uses the Twist Alliance Clinical Research Exome with 151 bp paired-end sequencing on NovaSeq 6000. Computationally, V2 relied on a multi-tool architecture including BWA-mem for alignment, MuTect for SNV detection, and Strelka2 for InDel detection. Noise filtering depended on a large external panel of normals constructed from 355 WES samples and 8,334 TCGA reference data sets. In addition, InDel detection required matched normal input (which introduced germline filtering effects that reduced sensitivity for certain variants), and the overall framework was not optimized to address FFPE-associated artifacts.
To address these challenges, Broad selected DRAGEN v4.3.11 as the new computational engine that provided them with a single integrated tool spanning alignment through variant calling. This approach replaces reliance on large external panels of normals with a lab-matched systematic noise model derived directly from representative datasets.
To evaluate performance, the study included two distinct comparisons. A controlled computational comparison ran both the V2 pipeline and DRAGEN on the same V6 sequencing reads (generated from HapMap cell line pools and NA12878 replicates), isolating the software as the only variable. Separately, an FFPE GBM comparison evaluated concordance when the same extracted tumor DNA was re-sequenced using both assay platforms and analyzed with their respective pipelines. The controlled comparison confirmed that DRAGEN's integrated variant calling is a major driver of performance gains, independent of upstream assay improvements.
Table 2: V2 and V6 Computational Workflow Comparison
Notes:
* Sensitivity metrics reflect both pipelines run on identical V6 sequencing data (Twist Alliance Clinical Research Exome, 151 bp PE, NovaSeq 6000) generated from standardized benchmarking samples. V2 data in the FP comparison used Illumina Nextera Rapid Capture Exome v1.2 (76 bp PE, HiSeq 2000/2500). Only the computational workflow differs in the sensitivity comparison.
** FP rates shown for reference from the full assay comparison (V2 data + V2 pipeline vs. V6 data + DRAGEN), where both sequencing chemistry and computational workflow differ. The controlled comparison also showed significantly lower FPs with DRAGEN (p = 5.7 x 10-11) but specific per-Mb rates were not reported separately.
Beyond improvements in benchmarking metrics, the DRAGEN-powered workflow also revealed biologically meaningful variants that were previously filtered out by their earlier V2 pipeline.
What Broad found in their FFPE cohort that they had not found before
When Broad compared the V6 and V2 workflows - each applied to data generated from the same extracted DNA but sequenced using their respective assay platforms - they identified 156 SNVs and 77 InDels unique to V6 across 11 GBM FFPE tumor-normal pairs. 95.5% of the V6-only SNVs had VAF ≥ 10% (significantly above the noise floor), indicating that these calls represent strong somatic signals. PCR validation independently confirmed 8 of 9 tested V6-only variants, with the majority absent from matched normal samples, supporting their classification as true somatic events. These results indicate that the V6 workflow was recovering real biological signals that their previous pipeline had filtered out [1].
Figure: Results from Broad Labs' validation study published in npj Genomic Medicine (2026), comparing somatic variant calling performance between their V2 and V6 bioinformatics workflows across 11 GBM FFPE tumor-normal pairs. A) The number of somatic variants identified by V2 only, V6 only, or both workflows across each sample, illustrating how many additional variants the V6 DRAGEN-powered workflow recovered. B) The variant allele fraction distribution of those calls, demonstrating that V6-only variants clustered at high VAF, consistent with true somatic events rather than low-confidence noise. C) PCR validation results for 29 selected variants, confirming that V6-only calls were predominantly true positives. D) The mutational spectrum across all 11 samples, with the dominant C>T signature consistent with known GBM biology and distinguishable from FFPE deamination artifacts. E) The landscape of coding variants in known GBM driver genes across the cohort, highlighting variants uniquely identified by V6 in genes including PTEN, EGFR, and SMC1A
A subset of variants Broad recovered were biologically significant and potentially relevant to known cancer biology. For example, three PTEN variants - p.V119F, p.R130* (nonsense), and p.R378fs (frameshift deletion) - had been filtered out by V2's read-remapping filter, triggered by a highly conserved PTEN pseudogene in the same genomic region. An EGFR missense variant, p.D587N, which had supporting reads at VAF <1% in both workflows, was filtered out in V2 by MuTect's nearby_gap_events and alt_allele_in_normal filters. An SMC1A variant, p.R799Q, was also lost to the same remapping filter. These four variants excluded by read-remapping had a median VAF of 32%, indicating strong, high-confidence calls that were incorrectly removed. All of these were very well-supported variants present in their sequencing data that the conventional variant calling pipeline had failed to report [1].
Additionally, V6 (DRAGEN-based) workflow confirmed a previously identified hypermutator phenotype in one sample (GBM_S7), driven by a somatic POLD1 p.R689W variant. The V6 workflow's improved sensitivity and specificity enabled more comprehensive characterization of the mutational landscape in this hypermutated sample. Sample GBM_S7 accounted for 67.4% of all SNVs and 80.9% of all InDels identified across the cohort of 11 GBM samples. This distinction is significant in genomic research for oncology. Tumor mutational burden (TMB) is widely used as a quantitative metric for characterizing mutational profiles. Misattributing a hypermutator signature to FFPE noise could lead to inaccurate TMB estimation and confound downstream analyses. Broad's V6 workflow, powered by DRAGEN's FFPE-aware variant calling, demonstrated the analytical sensitivity and specificity required to distinguish true biological signal from FFPE-related noise [1].
Taken together, these findings highlight broader implications for genomics research workflows.
Broad’s results raise a practical question for groups analyzing FFPE-derived sequencing data using conventional pipelines: are true somatic variants being filtered out before they are incorporated into downstream analysis?
In this evaluation, variants previously filtered out by the V2 pipeline’s read-remapping and variant calling filters were shown to represent biologically meaningful signals, underscoring how pipeline architecture directly influences which variants are retained for analysis.
By reducing artifact-driven noise while preserving true variant calls, the DRAGEN-based workflow enables the generation of higher-confidence VCF outputs with artifact modeling integrated within the variant calling process, reducing the need for extensive post-hoc filtering pipelines. This shift reduces the analytical burden associated with post hoc noise correction and allows greater focus on interpreting biologically meaningful variation and exploring cohort-level genomic patterns.
As noted in the study:
“DRAGEN’s filtering approach helped reduce candidate variant lists in lower-quality specimens and speed up tertiary interpretation and reporting, while still achieving adequate sensitivity.”
Tsuji, Rickles-Young et al., npj Genomic Medicine, 2026
To learn more about how DRAGEN approaches FFPE-aware somatic variant calling:
- Try DRAGEN on your own data: Access the DRAGEN demo at https://www.illumina.com/destination/dragen-demo.html
- Read the full Broad validation: Tsuji J, Rickles-Young M, Abreu J, et al. npj Genomic Medicine (2026). https://doi.org/10.1038/s41525-026-00569-w
- Explore the DRAGEN somatic pipeline documentation: Full parameter guide and somatic mode options at https://help.dragen.illumina.com
- Speak with your Illumina representative: To discuss how DRAGEN fits into your laboratory's somatic WES workflow
References
[1] Tsuji J, Rickles-Young M, Abreu J, et al. Clinical validation of a high-performance somatic exome sequencing assay: from target-enrichment strategy to variant calling. npj Genomic Medicine (2026). https://doi.org/10.1038/s41525-026-00569-w
[2] Scheffler K, et al. Somatic small-variant calling methods in Illumina DRAGEN Secondary Analysis. bioRxiv (2023). https://doi.org/10.1101/2023.03.23.534011
[3] Parse Biosciences. The scale of global FFPE archives. https://www.parsebiosciences.com
[4] Bruker Spatial Biology. FFPE tissue banks and sequencing access. https://www.brukerspatialbiology.com
[5] Coriell Institute. Institutional pathology archive scale and FFPE block management. https://www.coriell.org
[6] Illumina. FFPE sample volumes in academic medical centers. https://www.illumina.com
[7] New England Biolabs. Overcoming challenges in FFPE DNA library prep. https://www.neb.com/en-us/nebinspired-blog/overcoming-challenges-in-ffpe-dna-library-prep
[8] Lexogen. Fresh frozen vs FFPE samples for next-generation sequencing studies. https://www.lexogen.com/blog/fresh-frozen-vs-ffpe-samples-for-next-generation-sequencing-studies/
[9] Biocompare. Overcoming FFPE challenges in molecular biology research. https://www.biocompare.com/Editorial-Articles/618927-Overcoming-FFPE-Challenges-in-Molecular-Biology-Research/
[10] Guo Q, et al. The mutational signatures of formalin fixation on the human genome. Nature Communications 13, 4487 (2022). https://doi.org/10.1038/s41467-022-32041-5
For Research Use Only. Not for use in diagnostic procedures.
M-GL-04503