Back to Search Start Over

Comparative analysis of whole-genome sequencing pipelines to minimize false negative findings.

Authors :
Hwang KB
Lee IH
Li H
Won DG
Hernandez-Ferrer C
Negron JA
Kong SW
Source :
Scientific reports [Sci Rep] 2019 Mar 01; Vol. 9 (1), pp. 3219. Date of Electronic Publication: 2019 Mar 01.
Publication Year :
2019

Abstract

Comprehensive and accurate detection of variants from whole-genome sequencing (WGS) is a strong prerequisite for translational genomic medicine; however, low concordance between analytic pipelines is an outstanding challenge. We processed a European and an African WGS samples with 70 analytic pipelines comprising the combination of 7 short-read aligners and 10 variant calling algorithms (VCAs), and observed remarkable differences in the number of variants called by different pipelines (max/min ratio: 1.3~3.4). The similarity between variant call sets was more closely determined by VCAs rather than by short-read aligners. Remarkably, reported minor allele frequency had a substantial effect on concordance between pipelines (concordance rate ratio: 0.11~0.92; Wald tests, Pā€‰<ā€‰0.001), entailing more discordant results for rare and novel variants. We compared the performance of analytic pipelines and pipeline ensembles using gold-standard variant call sets and the catalog of variants from the 1000 Genomes Project. Notably, a single pipeline using BWA-MEM and GATK-HaplotypeCaller performed comparable to the pipeline ensembles for 'callable' regions (~97%) of the human reference genome. While a single pipeline is capable of analyzing common variants in most genomic regions, our findings demonstrated the limitations and challenges in analyzing rare or novel variants, especially for non-European genomes.

Details

Language :
English
ISSN :
2045-2322
Volume :
9
Issue :
1
Database :
MEDLINE
Journal :
Scientific reports
Publication Type :
Academic Journal
Accession number :
30824715
Full Text :
https://doi.org/10.1038/s41598-019-39108-2