bioRxiv ScienceSearch

Biology subjects

Le, T.

Publications and source records attributed to Le, T..

2 recordsLinked to original sources

Discovery of tandem and interspersed segmental duplications using high throughput sequencing

MotivationSeveral algorithms have been developed that use high throughput sequencing technology to characterize structural variations. Most of the existing approaches focus on detecting relatively simple types of SVs such as insertions, deletions, and short inversions. In fact, complex SVs are of crucial importance and several have been associated with genomic disorders. To better understand the contribution of complex SVs to human disease, we need new algorithms to accurately discover and genotype such variants. Additionally, due to similar sequencing signatures, inverted duplications or gene conversion events that include inverted segmental duplications are often characterized as simple inversions; and duplications and gene conversions in direct orientation may be called as simple deletions. Therefore, there is still a need for accurate algorithms to fully characterize complex SVs and thus improve calling accuracy of more simple variants. ResultsWe developed novel algorithms to accurately characterize tandem, direct and inverted interspersed segmental duplications using short read whole genome sequencing data sets. We integrated these methods to our TARDIS tool, which is now capable of detecting various types of SVs using multiple sequence signatures such as read pair, read depth and split read. We evaluated the prediction performance of our algorithms through several experiments using both simulated and real data sets. In the simulation experiments, using a 30x coverage TARDIS achieved 96% sensitivity with only 4% false discovery rate. For experiments that involve real data, we used two haploid genomes (CHM1 and CHM13) and one human genome (NA12878) from the Illumina Platinum Genomes set. Comparison of our results with orthogonal PacBio call sets from the same genomes revealed higher accuracy for TARDIS than state of the art methods. Furthermore, we showed a surprisingly low false discovery rate of our approach for discovery of tandem, direct and inverted interspersed segmental duplications prediction on CHM1 (less than 5% for the top 50 predictions). AvailabilityTARDIS source code is available at https://github.com/BilkentCompGen/tardis, and a corresponding Docker image is available at https://hub.docker.com/r/alkanlab/tardis/ Contactfhormozd@ucdavis.edu and calkan@cs.bilkent.edu.tr

bioinformatics

Estimating heterogeneous treatment effects by balancing heterogeneity and fitness

Estimating heterogeneous treatment effects is an important problem in many medical and biological applications since treatments may have different effects on the prognoses of different patients. Recently, several recursive partitioning methods have been proposed to identify the subgroups that with different responds to a treatment, and they rely on a fitness criterion to minimize the error between the estimated treatment effects and the unobservable true effects. In this paper, we propose that a heterogeneity criterion, which maximizes the differences of treatment effects among the subgroups, also needs to be considered. Moreover, we show that better performances can be achieved when the fitness and the heterogeneous criteria are considered simultaneously. Selecting the optimal splitting points then becomes a multi-objective problem; however, a solution that achieves optimal in both aspects are often not available. To solve this problem, we propose a multi-objective splitting procedure to balance both criteria. The proposed procedure is computationally efficient and fits naturally into the existing recursive partitioning framework. Experimental results show that the proposed multi-objective approach performs consistently better than existing ones.\n\nAuthor summaryThe effects of a treatment are often not the same for different individuals with different gene expressions. Learning to predict the heterogeneous treatment effects from clinical and expression data is an important step towards personalized medical treatment. Existing computational methods are not ideal for the task because they do not address the interpretability of the model and do not consider the limited sample sizes in biological and medical applications. Our method addresses these issues and achieves superior performance in analyzing the treatment effects of radiotherapy on breast cancer patients.

bioinformatics