bioRxiv ScienceSearch

Biology subjects

Amini, H.

Publications and source records attributed to Amini, H..

2 recordsLinked to original sources

Discovery of tandem and interspersed segmental duplications using high throughput sequencing

MotivationSeveral algorithms have been developed that use high throughput sequencing technology to characterize structural variations. Most of the existing approaches focus on detecting relatively simple types of SVs such as insertions, deletions, and short inversions. In fact, complex SVs are of crucial importance and several have been associated with genomic disorders. To better understand the contribution of complex SVs to human disease, we need new algorithms to accurately discover and genotype such variants. Additionally, due to similar sequencing signatures, inverted duplications or gene conversion events that include inverted segmental duplications are often characterized as simple inversions; and duplications and gene conversions in direct orientation may be called as simple deletions. Therefore, there is still a need for accurate algorithms to fully characterize complex SVs and thus improve calling accuracy of more simple variants. ResultsWe developed novel algorithms to accurately characterize tandem, direct and inverted interspersed segmental duplications using short read whole genome sequencing data sets. We integrated these methods to our TARDIS tool, which is now capable of detecting various types of SVs using multiple sequence signatures such as read pair, read depth and split read. We evaluated the prediction performance of our algorithms through several experiments using both simulated and real data sets. In the simulation experiments, using a 30x coverage TARDIS achieved 96% sensitivity with only 4% false discovery rate. For experiments that involve real data, we used two haploid genomes (CHM1 and CHM13) and one human genome (NA12878) from the Illumina Platinum Genomes set. Comparison of our results with orthogonal PacBio call sets from the same genomes revealed higher accuracy for TARDIS than state of the art methods. Furthermore, we showed a surprisingly low false discovery rate of our approach for discovery of tandem, direct and inverted interspersed segmental duplications prediction on CHM1 (less than 5% for the top 50 predictions). AvailabilityTARDIS source code is available at https://github.com/BilkentCompGen/tardis, and a corresponding Docker image is available at https://hub.docker.com/r/alkanlab/tardis/ Contactfhormozd@ucdavis.edu and calkan@cs.bilkent.edu.tr

bioinformatics

Concurrent Spatiotemporal Daily Land Use Regression Modeling and Missing Data Imputation of Fine Particulate Matter Using Distributed Space Time Expectation Maximization

Graphical Abstract\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC=\"FIGDIR/small/354852_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (37K):\norg.highwire.dtl.DTLVardef@caef72org.highwire.dtl.DTLVardef@12e5f38org.highwire.dtl.DTLVardef@16d6379org.highwire.dtl.DTLVardef@9d9daa_HPS_FORMAT_FIGEXP M_FIG C_FIG Land use regression (LUR) has been widely applied in epidemiologic research for exposure assessment. In this study, for the first time, we aimed to develop a spatiotemporal LUR model using Distributed Space Time Expectation Maximization (D-STEM). This spatiotemporal LUR model examined with daily particulate matter [≤] 2.5 m (PM2.5) within the megacity of Tehran, capital of Iran. Moreover, D-STEM missing data imputation was compared with mean substitution in each monitoring station, as it is equivalent to ignoring of missing data, which is common in LUR studies that employ regulatory monitoring stations data. The amount of missing data was 28% of the total number of observations, in Tehran in 2015. The annual mean of PM2.5 concentrations was 33 g/m3. Spatiotemporal R-squared of the D-STEM final daily LUR model was 78%, and leave-one-out cross-validation (LOOCV) R-squared was 66%. Spatial R-squared and LOOCV R-squared were 89% and 72%, respectively. Temporal R-squared and LOOCV R-squared were 99.5% and 99.3%, respectively. Mean absolute error decreased 26% in imputation of missing data by using the D-STEM final LUR model instead of mean substitution. This study reveals competence of the D-STEM software in spatiotemporal missing data imputation, estimation of temporal trend, and mapping of small scale (20 x 20 meters) within-city spatial variations, in the LUR context. The estimated PM2.5 concentrations maps could be used in future studies on short- and/or long-term health effects. Overall, we suggest using D-STEM capabilities in increasing LUR studies that employ data of regulatory network monitoring stations.\n\nHighlights- First Land Use Regression using D-STEM, a recently introduced statistical software\n- Assess D-STEM in spatiotemporal modeling, mapping, and missing data imputation\n- Estimate high resolution (20x20 m) daily maps for exposure assessment in a megacity\n- Provide both short- and long-term exposure assessment for epidemiological studies

epidemiology