bioRxiv ScienceSearch

Biology subjects

Finkers, R.

Publications and source records attributed to Finkers, R..

4 recordsLinked to original sources

Haplotype assembly of autotetraploid potato using integer linear programming

Haplotype assembly of polyploids is an open issue in plant genomics. Recent experimental studies on highly heterozygous autotetraploid potato have shown that available methods are not delivering satisfying results in practice. We propose an optimal method to assemble haplotypes of highly heterozygous polyploids from Illumina short sequencing reads. Our method is based on a generalization of the existing minimum fragment removal (MFR) model to the polyploid case and on new integer linear programs (ILPs) to reconstruct optimal haplotypes. We validate our methods experimentally by means of a combined evaluation on simulated and real data based on 83 previously sequenced autotetraploid potato cultivars. Results on simulated data show that our methods produce highly accurate haplotype assemblies, while results on real data confirm a sensible improvement over the state of the art. Binaries for Linux are available at: http://github.com/ComputationalGenomics/HaplotypeAssembler.

bioinformatics

Family-based haplotype estimation and allele dosage correction for polyploids using short sequence reads

DNA sequence reads contain information about the genomic variants located on a single chromosome. By extracting and extending this information (using the overlaps of the reads), the haplotypes of an individual can be obtained. Adding parent-offspring relationships to the read information in a population can considerably improve the quality of the haplotypes obtained from short reads, as pedigree information can compensate for spurious overlaps (due to sequencing errors) and insufficient overlaps (due to shallow coverage). This improvement is especially beneficial for polyploid organisms, which have more than two copies of each chromosome and are therefore more difficult to be haplotyped compared to diploids. We develop a novel method, PopPoly, to estimate polyploid haplotypes in an F1-population from short sequence data by considering the transmission of the haplotypes from the parents to the offspring. In addition, PopPoly employs this information to improve genotype dosage estimation and to call missing genotypes in the population. Through realistic simulations, we compare PopPoly to other haplotyping methods and show its better performance in terms of phasing accuracy and the accuracy of phased genotypes. We apply PopPoly to estimate the parental and offspring haplotypes for a tetraploid potato cross with 10 offspring, using Illumina HiSeq sequence data of 9 genomic regions involved in plant maturity and tuberisation.

bioinformatics

TriPoly: a haplotype estimation approach for polyploids using sequencing data of related individuals

Knowledge of \"haplotypes\", i.e. phased and ordered marker alleles on a chromosome, is essential to answer many questions in genetics and genomics. By generating short pieces of DNA sequence, high-throughput modern sequencing technologies make estimation of haplotypes possible for single individuals. In polyploids, however, haplotype estimation methods usually require deep coverage to achieve sufficient accuracy. This often renders sequencing-based approaches too costly to be applied to large populations needed in studies of Quantitative Trait Loci (QTL).\n\nWe propose a novel haplotype estimation method for polyploids, TriPoly, that combines sequencing data with Mendelian inheritance rules to infer haplotypes in parent-offspring trios. Using realistic simulations of short- read sequencing data for potato (Solanum tuberosum) and banana (Musa acuminata) trios, we show that TriPoly yields more accurate progeny haplotypes at low coverages compared to the existing methods that work on single individuals.

bioinformatics

Exploiting Next Generation Sequencing to solve the Haplotyping puzzle in Polyploids: a Simulation study

Haplotypes are the units of inheritance in an organism, and many genetic analyses depend on their precise determination. Methods for haplotyping single individuals use the phasing information available in Next Generation sequencing reads, by matching overlapping sNPs while penalizing post hoc nucleotide corrections made. Haplotyping diploids is relatively easy, but the complexity of the problem increases drastically for polyploid genomes, which are found in both model organisms and in economically relevant plant and animal species. While a number of tools are available for haplotyping polyploids, the effects of the genomic makeup and the sequencing strategy followed on the accuracy of these methods have hitherto not been thoroughly evaluated.\n\nWe developed the simulation pipeline haplosim to evaluate the performance of haplotype estimation algorithms for polyploids: HapCompass, HapTree and SDhaP, in settings varying in sequencing approach, ploidy levels and genomic diversity, using tetraploid potato as the model. Our results show that sequencing depth is the major determinant of haplotype estimation quality, that 1kb PacBio CCS reads and Illumina reads with large insert-sizes are competitive, and that all methods fail to produce good haplotypes when ploidy levels increase. Comparing the three methods, HapTree produces the most accurate estimates, but also consumes the most resources. There is clearly room for improvement in polyploid haplotyping algorithms.

bioinformatics