bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Human genomic regions with exceptionally high or low levels of population differentiation identified from 911 whole-genome sequences

BackgroundPopulation differentiation has proved to be effective for identifying loci under geographically-localized positive selection, and has the potential to identify loci subject to balancing selection. We have previously investigated the pattern of genetic differentiation among human populations at 36.8 million genomic variants to identify sites in the genome showing high frequency differences. Here, we extend this dataset to include additional variants, survey sites with low levels of differentiation, and evaluate the extent to which highly differentiated sites are likely to result from selective or other processes.\n\nResultsWe demonstrate that while sites of low differentiation represent sampling effects rather than balancing selection, sites showing extremely high population differentiation are enriched for positive selection events and that one half may be the result of classic selective sweeps. Among these, we rediscover known examples, where we actually identify the established functional SNP, and discover novel examples including the genes ABCA12, CALD1 and ZNF804, which we speculate may be linked to adaptations in skin, calcium metabolism and defense, respectively.\n\nConclusionsWe have identified known and many novel candidate regions for geographically restricted positive selection, and suggest several directions for further research.

Genomics

iCAGES: integrated CAncer GEnome Score for comprehensively prioritizing cancer driver genes in personal genomes

All cancers arise as a result of the acquisition of somatic mutations that drive the disease progression. A number of computational tools have been developed to identify driver genes for a specific cancer from a group of cancer samples. However, it remains a challenge to identify driver mutations/genes for an individual patient and design drug therapies. We developed iCAGES, a novel statistical framework to rapidly analyze patient-specific cancer genomic data, prioritize personalized cancer driver events and predict personalized therapies. iCAGES includes three consecutive layers: the first layer integrates contributions from coding, non-coding and structural variations to infer driver variants. For coding mutations, we developed a radial support vector machine using manually curated mutations to predict their driver potential. The second layer identifies driver genes, by using information from the first layer and integrating prior biological knowledge on gene-gene and gene-phenotype networks. The third layer prioritizes personalized drug treatment, by classifying potential driver genes into different categories and querying drug-gene databases. Compared to currently available tools, iCAGES achieves better performance by correctly classifying point coding driver mutations (AUC=0.97, 95% CI: 0.97-0.97, significantly better than the second best tool with P=0.01) and genes (AUC=0.93, 95% CI: 0.93-0.94, significantly better than MutSigCV with P<1x10-15). We also illustrated two examples where iCAGES correctly nominated two targeted drugs for two advanced cancer patients with exceptional response, based on their somatic mutation profiles. iCAGES leverages personal genomic information and prior biological knowledge, effectively identifies cancer driver genes and predicts treatment strategies. iCAGES is available at http://icages.usc.edu.

Genomics

Ongoing human chromosome end extension driven by a primate ancestral genomic region revealed by analysis of BioNano genomics data

The majority of human chromosome ends remain incompletely assembled due to their highly repetitive structure. In this study, we use BioNano data to anchor and extend chromosome ends from two European trios as well as two unrelated Asian genomes. BioNano assembled chromosome ends are structurally divergent from the reference genome, including both missing sequence (10%) and extensions(22%). These extensions are heritable and in some cases divergent between Asian and European samples. Six ninths of the extension sequence in NA12878 can be confirmed and filled by nanopore data. We identify two sequence families in these sequences which have undergone substantial duplication in multiple primate lineages. We show that these sequence families have arisen from progenitor interstitial sequence on the ancestral primate chromosome 7. Comparison of chromosome end sequences from 15 species revealed that chromosome end missing sequence matches the corresponding phylogenetic relationship and revealed a rate of chromosome extension per chromosome of 0.0020 bp per year in average.

genomics

Genome-wide identification of directed gene networks using large-scale population genomics data

Identification of causal drivers behind regulatory gene networks is crucial in understanding gene function. We developed a method for the large-scale inference of gene-gene interactions in observational population genomics data that are both directed (using local genetic instruments as causal anchors, akin to Mendelian Randomization) and specific (by controlling for linkage disequilibrium and pleiotropy). The analysis of genotype and whole-blood RNA-sequencing data from 3,072 individuals identified 49 genes as drivers of downstream transcriptional changes (P < 7 x 10-10), among which transcription factors were overrepresented (P = 3.3 x 10-7). Our analysis suggests new gene functions and targets including for SENP7 (zinc-finger genes involved in retroviral repression) and BCL2A1 (novel target genes possibly involved in auditory dysfunction). Our work highlights the utility of population genomics data in deriving directed gene expression networks. A resource of trans-effects for all 6,600 genes with a genetic instrument can be explored individually using a web-based browser.

genomics

Genome-wide mapping reveals conserved and diverged R-loop activities in the unusual genetic landscape of the African trypanosome genome

R-loops are stable RNA-DNA hybrids that have been implicated in transcription initiation and termination, as well as in telomere homeostasis, chromatin formation, and genome replication and instability. RNA Polymerase (Pol) II transcription in the protozoan parasite Trypanosoma brucei is highly unusual: virtually all genes are co-transcribed from multigene transcription units, with mRNAs generated by linked trans-splicing and polyadenylation, and transcription initiation sites display no conserved promoter motifs. Here, we describe the genome-wide distribution of R-loops in wild type mammal-infective T. brucei and in mutants lacking RNase H1, revealing both conserved and diverged functions. Conserved localisation was found at centromeres, rRNA genes and retrotransposon-associated genes. RNA Pol II transcription initiation sites also displayed R-loops, suggesting a broadly conserved role despite the lack of promoter conservation or transcription initiation regulation. However, the most abundant sites of R-loop enrichment were within the intergenic regions of the multigene transcription units, where the hybrids coincide with sites of polyadenylation and nucleosome-depletion. Thus, instead of functioning in transcription termination, most T. brucei R-loops act in a novel role, promoting RNA Pol II movement or mRNA processing. Finally, we show there is little evidence for correlation between R-loop localisation and mapped sites of DNA replication initiation.

genomics

Whole-Genome Genomics Correlates of Response To Anti-PD1 Therapy in Relapsed/Refractory Natural Killer/T Cell Lymphoma

AbstractThis study aims to identify recurrent genetic alterations in relapsed or refractory (RR) natural-killer/T-cell lymphoma (NKTL) patients who have achieved complete response (CR) with programmed cell death 1 (PD-1) blockade therapy. Seven of the eleven patients treated with pembrolizumab achieved CR while the remaining four had progressive disease (PD). Using whole genome sequencing (WGS), we found recurrent clonal structural rearrangements (SR) of the PD-L1 gene in four of the seven (57%) CR patients pretreated tumors. These PD-L1 SRs consist of inter-chromosomal translocations, tandem duplication and micro-inversion that disrupted the suppressive function of PD-L1 3UTR. Interestingly, recurrent JAK3-activating (p.A573V) mutations were also validated in two CR patients tumors that did not harbor the PD-L1 SR. Importantly, these mutations were absent in the four PD cases. With immunohistochemistry (IHC), PD-L1 positivity could not discriminate patients who archived CR (range: 6%-100%) from patients who had PD (range: 35%-90%). PD-1 blockade with pembrolizumab is a potent strategy for RR NKTL patients and genomic screening could potentially accompany PD-L1 IHC positivity to better select patients for anti-PD-1 therapy.

genomics

Genome-wide analyses supported by RNA-Seq reveal non-canonical splice sites in plant genomes

Most eukaryotic genes comprise exons and introns thus requiring the precise removal of introns from pre-mRNAs to enable protein biosynthesis. U2 and U12 spliceosomes catalyze this step by recognizing motifs on the transcript in order to remove the introns. A process which is dependent on precise definition of exon-intron borders by splice sites, which are consequently highly conserved across species. Only very few combinations of terminal dinucleotides are frequently observed at intron ends, dominated by the canonical GT-AG splice sites on the DNA level.\n\nHere we investigate the occurrence of diverse combinations of dinucleotides at predicted splice sites. Analyzing 121 plant genome sequences based on their annotation revealed strong splice site conservation across species, annotation errors, and true biological divergence from canonical splice sites. The frequency of non-canonical splice sites clearly correlates with their divergence from canonical ones indicating either an accumulation of probably neutral mutations, or evolution towards canonical splice sites. Strong conservation across multiple species and non-random accumulation of substitutions in splice sites indicate a functional relevance of non-canonical splice sites. The average composition of splice sites across all investigated species is 98.7% for GT-AG, 1.2% for GC-AG, 0.06% for AT-AC, and 0.09% for minor non-canonical splice sites. RNA-Seq data sets of 35 species were incorporated to validate non-canonical splice site predictions through gaps in sequencing reads alignments and to demonstrate the expression of affected genes. We conclude that bona fide non-canonical splice sites are present and appear to be functionally relevant in most plant genomes, if at low abundance.

genomics

Using Bayesian multilevel whole-genome regression models for partial pooling of estimation sets in genomic prediction

Estimation set size is an important determinant of genomic prediction accuracy. Plant breeding programs are characterized by a high degree of structuring, particularly into populations. This hampers establishment of large estimation sets for each population. Pooling populations increases estimation set size but ignores unique genetic characteristics of each. A possible solution is partial pooling with multilevel models, which allows estimating population specific marker effects while still leveraging information across populations. We developed a Bayesian multilevel whole-genome regression model and compared its performance to that of the popular BayesA model applied to each population separately (no pooling) and to the joined data set (complete pooling). As example we analyzed a wide array of traits from the nested association mapping maize population. There we show that for small population sizes (e.g., < 50), partial pooling increased prediction accuracy over no or complete pooling for populations represented in the estimation set. No pooling was superior however when populations were large. In another example data set of interconnected biparental maize populations either partial or complete pooling were superior, depending on the trait. A simulation showed that no pooling is superior when differences in genetic effects among populations are large and partial pooling when they are intermediate. With small differences, partial and complete pooling achieved equally high accuracy. For prediction of new populations, partial and complete pooling had very similar accuracy in all cases. We conclude that partial pooling with multilevel models can maximize the potential of pooling by making optimal use of information in pooled estimation sets.

Genetics

Integrative tissue-specific functional annotations in the human genome provide novel insights on many complex traits and improve signal prioritization in genome wide association studies

Extensive efforts have been made to understand genomic function through both experimental and computational approaches, yet proper annotation still remains challenging, especially in non-coding regions. In this manuscript, we introduce GenoSkyline, an unsupervised learning framework to predict tissue-specific functional regions through integrating high-throughput epigenetic annotations. GenoSkyline successfully identified a variety of non-coding regulatory machinery including enhancers, regulatory miRNA, and hypomethylated transposable elements in extensive case studies. Integrative analysis of GenoSkyline annotations and results from genome-wide association studies (GWAS) led to novel biological insights on the etiologies of a number of human complex traits. We also explored using tissue-specific functional annotations to prioritize GWAS signals and predict relevant tissue types for each risk locus. Brain and blood-specific annotations led to better prioritization performance for schizophrenia than standard GWAS p-values and non-tissue-specific annotations. As for coronary artery disease, heart-specific functional regions was highly enriched of GWAS signals, but previously identified risk loci were found to be most functional in other tissues, suggesting a substantial proportion of still undetected heart-related loci. In summary, GenoSkyline annotations can guide genetic studies at multiple resolutions and provide valuable insights in understanding complex diseases. GenoSkyline is available at http://genocanyon.med.yale.edu/GenoSkyline.

Bioinformatics

Annotation Regression for Genome-Wide Association Studies with an Application to Psychiatric Genomic Consortium Data

Although genome-wide association studies (GWAS) have been successful at finding thousands of disease-associated genetic variants (GVs), identifying causal variants and elucidating the mechanisms by which genotypes influence phenotypes are critical open questions. A key challenge is that a large percentage of disease-associated GVs are potential regulatory variants located in noncoding regions, making them difficult to interpret. Recent research efforts focus on going beyond annotating GVs by integrating functional annotation data with GWAS to prioritize GVs. However, applicability of these approaches is challenged by high dimensionality and heterogeneity of functional annotation data. Furthermore, existing methods often assume global associations of GVs with annotation data. This strong assumption is susceptible to violations for GVs involved in many complex diseases. To address these issues, we develop a general regression framework, named Annotation Regression for GWAS (ARoG). ARoG is based on finite mixture of linear regression models where GWAS association measures are viewed as responses and functional annotations as predictors. This mixture framework addresses heterogeneity of effects of GVs by grouping them into clusters and high dimensionality of the functional annotations by enabling annotation selection within each cluster. ARoG further employs permutation testing to evaluate the significance of selected annotations. Computational experiments indicate that ARoG can discover distinct associations between disease risk and functional annotations. Application of ARoG to autism and schizophrenia data from Psychiatric Genomics Consortium led to identification of GVs that significantly affect interactions of several transcription factors with DNA as potential mechanisms contributing to these disorders.

Bioinformatics

290 Metagenome-assembled Genomes from the Mediterranean Sea: Ongoing Effort to Generate Genomes from the Tara Oceans Dataset

The Tara Oceans Expedition has provided large, publicly-accessible microbial metagenomic datasets from a circumnavigation of the globe. Utilizing several size fractions from the samples originating in the Mediterranean Sea, we have used current assembly and binning techniques to reconstruct 290 putative high-quality metagenome-assembled bacterial and archaeal genomes, with an estimated completion of [&ge;]50%, and an additional 2,786 bins, with estimated completion of 0-50%. We have submitted our results, including initial taxonomic and phylogenetic assignments, for the putative high-quality genomes to open-access repositories for the scientific community to use in ongoing research.

Microbiology

Genome-wide prediction of microRNAs in Zika virus genomes reveals possible interactions with human genes involved in the nervous system development

Zika virus (ZIKV) is a member of the family Flaviviridae. In 2015, ZIKV triggered a large epidemic in Brazil and spread across Latin America. In November of that year, the Brazilian Ministry of Health reported a 20-fold increase in cases of neonatal microcephaly, which corresponds geographically and temporally to the ZIKV outbreak. ZIKV was isolated from the brain tissue of a fetus diagnosed with microcephaly, and recent studies in mice models revealed that ZIKV infection may cause brain defects by influencing brain cell developments. Unfortunately, the mechanisms by which ZIKV alters neurophysiological development remain unknown. MicroRNAs (miRNAs) are small noncoding RNAs that regulate post-transcriptional gene expression by translational repression. In order to gain insight into the possible role of ZIKV-mediated miRNA signaling dysfunction in brain-tissue development, we computationally predicted new miRNAs encoded by the ZIKV genome and their effective hybridization with transcripts from human genes previously shown to be involved in microcephalia. The results of these studies suggest a possible role of these miRNAs on the expression of human genes associated with this disease. Besides, a new ZIKV miRNA was predicted in the 3stem loop (3 SL) of the 3untranslated region (3UTR) of the ZIKV genome, suggesting the role of the 3UTR of flaviviruses as a source of miRNAs.

Microbiology

Mitochondrial genome variation affects the mutation rate of the nuclear genome in Drosophila melanogaster

Mutations are the raw material for evolutionary change. While the mutation rate has been thought constant between individuals, recent research has shown that poor genetic condition can elevate the mutation rate. Mitonuclear genetic conflict is a potential source of poor genetic condition, and considering the high mutation rate of mitochondrial genomes, there should be ample scope for mitochondrial mutations to interfere with genetic condition, with concomitant effects on the nuclear mutation rate. Moreover, because theory suggests mitochondrial genetic effects will often be male-biased, such effects could be more strongly felt in males than females. Here, by mating irradiated male Drosophila melanogaster to isogenic females bearing six distinct mitochondrial haplotypes, we tested whether mitochondrial genetic variation affects DNA repair capacity, and whether effects of mutation load on reproductive function are shaped by interactions between sex and mitochondrial haplotype. We found mitochondrial genetic effects on DNA repair, and that the mutational variance of reproductive fitness was higher in males bearing haplotypes characterized by high female fitness. These results suggest that mitochondrial genome variation may affect the mutation rate, and that induced mutations interact more strongly with male than female reproductive function. The potential for haplotype-specific effects on the nuclear mutation rate has broad implications for evolutionary dynamics, such as the accumulation of genetic load, adaptive potential, and the evolution of sexual dimorphism.

evolutionary biology

Genome-wide Enhancer Maps Differ Significantly in Genomic Distribution, Evolution, and Function

Non-coding gene regulatory enhancers are essential to transcription in mammalian cells. As a result, a large variety of experimental and computational strategies have been developed to identify cis-regulatory enhancer sequences. In practice, most studies consider enhancers identified by only a single method, and the concordance of enhancers identified by different methods has not been comprehensively evaluated. Here, we assess the similarities of enhancer sets identified by ten representative strategies in four biological contexts and evaluate the robustness of downstream conclusions to the choice of identification strategy. All pairs of enhancer sets we evaluated overlap significantly more than expected by chance; however, we also found significant dissimilarity between enhancer sets in their genomic characteristics, evolutionary conservation, and association with functional loci within each context. We find most regions identified as enhancers are supported by only one method. The disagreement is sufficient to influence interpretation of GWAS SNPs and eQTL, and to lead to disparate conclusions about enhancer biology and disease mechanisms. We also find only limited evidence that regions identified by multiple enhancer identification methods are better candidates than those identified by a single method. Our results highlight the inherent complexity of enhancer biology and argue that current approaches have yet to adequately account for enhancer diversity. As a result, we cannot recommend the use of any single enhancer identification strategy in isolation. To facilitate assessment of enhancer diversity on studies conclusions, we developed creDB, a database of enhancer annotations designed to integrate into bioinformatics workflows. While our findings highlight a major challenge to mapping the genetic architecture of complex disease and interpreting regulatory variants found in patient genomes, a systematic understanding of similarities and differences in enhancer identification methodology will ultimately enable robust inferences about gene regulatory sequences.

genetics

Identification of genome-wide significant shared genomic segments in large extended Utah families at high risk for completed suicide

Suicide is the 10th leading cause of death in the US. While environment has undeniable impact, evidence suggests genetic factors play a significant role in completed suicide. We linked a resource of >4,500 DNA samples from completed suicides obtained from the Utah Medical Examiner to genealogical records and medical records data available on over 8 million individuals. This linking has resulted in the identification of high-risk extended families (7-9 generations) with significant familial risk of completed suicide. Familial aggregation across distant relatives minimizes effects of shared environment, provides more genetically homogeneous risk groups, and magnifies genetic risks through familial repetition. We analyzed Illumina PsychArray genotypes from suicide cases in 43 high-risk families, identifying 30 distinct shared genomic segments with genome-wide evidence (p=2.02E-07 to 1.30E-18) of segregation with completed suicide. The 207 genes implicated by the shared regions provide a focused set of genes for further study; 18 have been previously associated with suicide risk. While PsychArray variants do not represent exhaustive variation within the 207 genes, we investigated these for specific segregation within the high-risk families, and for association of variants with predicted functional impact in ~1300 additional Utah suicides unrelated to the discovery families. None of the limited PsychArray variants explained the high-risk family segregation; sequencing of these regions will be needed to discover segregating risk variants, which may be rarer or regulatory. However, additional association tests yielded four significant PsychArray variants (SP110, rs181058279; AGBL2, rs76215382; SUCLA2, rs121908538; APH1B, rs745918508), raising the likelihood that these genes confer risk of completed suicide.

genetics

GetOrganelle: a simple and fast pipeline for de novo assembly of a complete circular chloroplast genome using genome skimming data

GetOrganelle is a state-of-the-art toolkit to assemble accurate organelle genomes from NGS data. This toolkit recruit organelle-associated reads using a modified \"baiting and iterative mapping\" approach, conducts de novo assembly, filters and disentangles assembly graph, and produces all possible configurations of circular organelle genomes. For 50 published samples, we reassembled the circular plastome in 47 samples using GetOrganelle, but only in 12 samples using NOVOPlasty. In comparison with published/NOVOPlasty plastomes, we demonstrated that GetOrganelle assemblies are more accurate. Moreover, we assembled complete mitogenomes of fungi and animals using GetOrganelle. GetOrganelle is freely released under a GPL-3 license (https://github.com/Kinggerm/GetOrganelle).

bioinformatics

Tychus: a whole genome sequencing pipeline for assembly, annotation and phylogenetics of bacterial genomes

SummaryTychus is a tool that allows researchers to perform massively parallel whole genome sequence (WGS) analysis with the goal of producing a high confidence and comprehensive description of the bacterial genome. Key features of the Tychus pipeline include the assembly, annotation, alignment, variant discovery and phylogenetic inference of large numbers of WGS isolates in parallel using open-source bioinformatics tools and virtualization technology. All prerequisite tools and dependencies come packaged together in a single suite that can be easily downloaded and installed on Linux and Mac operating systems.\n\nAvailabilityTychus is freely available as an open-source package under the MIT license, and can be downloaded via GitHub (https://github.com/Abdo-Lab/Tychus).\n\nContactzaid.abdo@colostate.edu

bioinformatics

Inference of genomic spatial organization from a whole genome bisulfite sequencing sample.

Common approaches to characterize the structure of the DNA in the nucleus, such as the different Chromosome Conformation Capture methods, have not currently been widely applied to different tissue types due to several practical difficulties including the requirement for intact cells to start the sample preparation. In contrast, techniques based on sodium bisulfite conversion of DNA to assay DNA methylation, have been widely applied to many different tissue types in a variety of organisms. Recent work has shown the possibility of inferring some aspects of the three dimensional DNA structure from DNA methylation data, raising the possibility of three dimensional DNA structure prediction using the large collection of already generated DNA methylation datasets. We propose a simple method to predict the values of the first eigenvector of the Hi-C matrix of a sample (and hence the positions of the A and B compartments) using only the GC content of the sequence and a single whole genome bisulfite sequencing (WGBS) experiment which yields information on the methylation levels and their variability along the genome. We train and test our model on 10 samples for which we have data from both bisulfite sequencing and chromosome conformation experiments and our most relevant finding is that the variability of DNA methylation along the sequence is often a better predictor than methylation itself. We then run a prediction on 206 DNA methylation profiles produced by the Blueprint project and use ChIP-Seq and RNA-Seq data to confirm that the forecasted eigenvector delineates correctly the physical chromatin compartments observed with the Hi-C experiment.

bioinformatics