bioRxiv ScienceSearch

Biology subjects

Santhosh Girirajan

Publications and source records attributed to Santhosh Girirajan.

5 recordsLinked to original sources

Novel metrics to measure coverage in whole exome sequencing datasets reveal local and global non-uniformity

Whole Exome Sequencing (WES) is a powerful clinical diagnostic tool for discovering the genetic basis of many diseases. A major shortcoming of WES is uneven coverage of sequence reads over the exome targets contributing to many low coverage regions, which hinders accurate variant calling. In this study, we devised two novel metrics, Cohort Coverage Sparseness (CCS) and Unevenness (UE) Scores for a detailed assessment of the distribution of coverage of sequence reads. Employing these metrics we revealed non-uniformity of coverage and low coverage regions in the WES data generated by three different platforms. This non-uniformity of coverage is both local (coverage of a given exon across different platforms) and global (coverage of all exons across the genome in the given platform). The low coverage regions encompassing functionally important genes were often associated with high GC content, repeat elements and segmental duplications. While a majority of the problems associated with WES are due to the limitations of the capture methods, further refinements in WES technologies have the potential to enhance its clinical applications.

Genomics

Quantitative assessment of eye phenotypes for functional genetic studies using Drosophila melanogaster

About two-thirds of the vital genes in the Drosophila genome are involved in eye development, making the fly eye an excellent genetic system to study cellular function and development, neurodevelopment/degeneration, and complex diseases such as cancer and diabetes. We developed a novel computational method, implemented as Flynotyper software (http://flynotyper.sourceforge.net), to quantitatively assess the morphological defects in the Drosophila eye resulting from genetic alterations affecting basic cellular and developmental processes. Flynotyper utilizes a series of image processing operations to automatically detect the fly eye and the individual ommatidium, and calculates a phenotypic score as a measure of the disorderliness of ommatidial arrangement in the fly eye. As a proof of principle, we tested our method by analyzing the defects due to eye-specific knockdown of Drosophila orthologs of 12 neurodevelopmental genes to accurately document differential sensitivities of these genes to dosage alteration. We also evaluated eye images from six independent studies assessing the effect of overexpression of repeats, candidates from peptide library screens, and modifiers of neurotoxicity and developmental processes on eye morphology, and show strong concordance with the original assessment. We further demonstrate the utility of this method by analyzing 16 modifiers of sine oculis obtained from two genome-wide deficiency screens of Drosophila and accurately quantifying the effect of its enhancers and suppressors during eye development. Our method will complement existing assays for eye phenotypes and increase the accuracy of studies that use fly eyes for functional evaluation of genes and genetic interactions.

Genetics

Improving the Power of Structural Variation Detection by Augmenting the Reference

The uses of the Genome Reference Consortiums human reference sequence can be roughly categorized into three related but distinct categories: as a representative species genome, as a coordinate system for identifying variants, and as an alignment reference for variation detection algorithms. However, the use of this reference sequence as simultaneously a representative species genome and as an alignment reference leads to unnecessary artifacts for structural variation detection algorithms and limits their accuracy. We show how decoupling these two references and developing a separate alignment reference can significantly improve the accuracy of structural variation detection, lead to improved genotyping of disease related genes, and decrease the cost of studying polymorphism in a population.

Bioinformatics

Gene discovery and functional assessment of rare copy-number variants in neurodevelopmental disorders

Rare copy-number variants (CNVs) are a significant cause of neurodevelopmental disorders. The sequence architecture of the human genome predisposes certain individuals to deletions and duplications within specific genomic regions. While assessment of individuals with different breakpoints has identified causal genes for certain rare CNVs, deriving gene-phenotype correlations for rare CNVs with similar breakpoints has been challenging. We present a comprehensive review of the literature related to genetic architecture that is predisposed to recurrent rearrangements, and functional evaluation of deletions, duplications, and candidate genes within rare CNV intervals using mouse, zebrafish, and fruit fly models. It is clear that phenotypic assessment and complete genetic evaluation of large cohorts of individuals carrying specific CNVs and functional evaluation using multiple animal models are necessary to understand the molecular genetic basis of neurodevelopmental disorders.

Genomics

The impact of comorbidity of intellectual disability on estimates of autism prevalence among children enrolled in US special education

ObjectivesWhile recent studies suggest a converging role for genetic factors towards risk for nosologically distinct disorders including autism, intellectual disability (ID), and epilepsy, current estimates of autism prevalence fail to take into account the impact of comorbidity of these neurodevelopmental disorders on autism diagnosis. We aimed to assess the effect of potential comorbidity of ID on the diagnosis and prevalence of autism by analyzing 11 years of special education enrollment data.\n\nDesignPopulation study of autism using the United States special education enrollment data from years 2000-2010.\n\nSettingUS special education.\n\nParticipantsWe analyzed 11 years (2000 to 2010) of longitudinal data on approximately 6.2 million children per year from special education enrollment.\n\nResultsWe found a 331% increase in the prevalence of autism from 2000 to 2010 within special education, potentially due to a diagnostic recategorization from frequently comorbid features like ID. In fact, the decrease in ID prevalence equaled an average of 64.2% of the increase of autism prevalence for children aged 3-18 years. The proportion of ID cases potentially undergoing recategorization to autism was higher (p=0.007) among older children (75%) than younger children (48%). Some US states showed significant negative correlations between the prevalence of autism compared to that of ID while others did not, suggesting differences in state-specific health policy to be a major factor in categorizing autism.\n\nConclusionsOur results suggest that current ascertainment practices are based on a single facet of autism-specific clinical features and do not consider associated comorbidities that may confound diagnosis. Longitudinal studies with detailed phenotyping and deep molecular genetic analyses are necessary to completely understand the cause of this complex disorder. Future studies of autism prevalence should also take these factors into account.\n\nSTRENGTHS AND LIMITATIONSO_LIWe present a large-scale population study of autism prevalence using longitudinal data from 2000-2010 on approximately 6.2 million children enrolled in US special education.\nC_LIO_LIWe provide one possible but compelling explanation for increase in autism prevalence and show that current ascertainment of autism is based on single facet of clinical features without considering other comorbid features such as intellectual disability\nC_LIO_LIWe are not able to dissect the exact frequency of comorbid features over time as US special education allows for enrollment under only one diagnostic category and does not document comorbidity information.\nC_LIO_LIThis study examines how comorbidity of related phenotypes such as ID can impact estimates of autism prevalence and does not assess other factors reported to impact autism prevalence.\nC_LI

Genomics