Dergiler / Turkish Journal of Electrical Engineering and Computer Sciences / 2020 / Cilt: 28 - Sayı: 1

Automatic characterization of copy number polymorphism using high throughput sequencing

Sayfa
253–261
DOI
—

Abstract

Genome structural variation, broadly defined as alterations longer than 50 bp, are important sources forgenetic variation among humans, including those that cause complex diseases such as autism, developmental delay, andschizophrenia. Although there has been considerable progress in characterizing structural variation since the beginningsof the 1000 Genomes Project, one form of structural variation called segmental duplications (SDs) remained largelyunderstudied in large cohorts. This is mostly because SDs cannot be accurately discovered using the alignment filesgenerated with standard read mapping tools. Instead, they can only be found when multiple map locations are considered.There is still a single algorithm available for SD discovery, which includes various tools and scripts that are not portableand are difficult to use. Additionally, this algorithm relies on a priori information for regions where no structuralvariations are discovered in large number of genomes. Therefore, there is a need for fully automated, portable, anduser-friendly tools to make SD characterization a part of genome analyses. Here we introduce such an algorithm andefficient implementation, called mrCaNaVaR, that aims to fill this gap in genome analysis toolbox.