Back to Search Start Over

The Lithuanian reference genome LT1 - a human de novo genome assembly with short and long read sequence and Hi-C data

Authors :
Dan M. Bolser
Sungwon Jeon
Jeremy S. Edwards
Yeo Jin Kim
Changjae Kim
Asta Blazyte
Ahn J
Jong Bhak
Hajin Kim
Changhan Yoon
Publication Year :
2021
Publisher :
Cold Spring Harbor Laboratory, 2021.

Abstract

We present LT1, the first high-quality human reference genome from the Baltic States. LT1 is a female de novo human reference genome assembly constructed using 57× of ultra-long nanopore reads and 47× of short paired-end reads. We also utilized 72 Gb of Hi-C chromosomal mapping data to maximize the assembly’s contiguity and accuracy. LT1’s contig assembly was 2.73 Gbp in length comprising of 4,490 contigs with an N50 value of 13.4 Mbp. After scaffolding with Hi-C data and extensive manual curation, we produced a chromosome-scale assembly with an N50 value of 138 Mbp and 4,699 scaffolds. Our gene prediction quality assessment using BUSCO identify 89.3% of the single-copy orthologous genes included in the benchmarking set. Detailed characterization of LT1 suggested it has 73,744 predicted transcripts, 4.2 million autosomal SNPs, 974,000 short indels, and 12,330 large structural variants. These data are shared as a public resource without any restrictions and can be used as a benchmark for further in-depth genomic analyses of the Baltic populations.

Details

Database :
OpenAIRE
Accession number :
edsair.doi...........d0a28764f21ea852e0d56c4a598d096f
Full Text :
https://doi.org/10.1101/2021.04.05.438426