bioRxiv ScienceSearch

Biology subjects

von Haeseler, A.

Publications and source records attributed to von Haeseler, A..

8 recordsLinked to original sources

The evolutionary traceability of proteins

Orthologs document the evolution of genes and metabolic capacities encoded in extant and ancient genomes. Orthologous genes that are detected across the full diversity of contemporary life allow reconstructing the gene set of LUCA, the last universal common ancestor. These genes presumably represent the functional repertoire common to - and necessary for - all living organisms. Design of artificial life has the potential to test this. Recently, a minimal gene (MG) set for a self-replicating cell was determined experimentally, and a surprisingly high number of genes have unknown functions and are not represented in LUCA. However, as similarity between orthologs decays with time, it becomes insufficient to infer common ancestry, leaving ancient gene set reconstructions incomplete and distorted to an unknown extent. Here we introduce the evolutionary traceability, together with the software protTrace, that quantifies, for each protein, the evolutionary distance beyond which the sensitivity of the ortholog search becomes limiting. We show that the LUCA set comprises only high-traceable proteins most of which have catalytic functions. We further show that proteins in the MG set lacking orthologs outside bacteria mostly have low traceability, leaving open whether their eukaryotic orthologs have just been overlooked. On the example of REC8, a protein essential for chromosome cohesion, we demonstrate how a traceability-informed adjustment of the search sensitivity identifies hitherto missed orthologs in the fast-evolving microsporidia. Taken together, the evolutionary traceability helps to differentiate between true absence and non-detection of orthologs, and thus improves our understanding about the evolutionary conservation of functional protein networks.

evolutionary biology

A-to-I RNA editing uncovers hidden signals of adaptive genome evolution in animals

In animals, the most common type of RNA editing is the deamination of adenosines (A) into inosines (I). Because inosines base-pair with cytosines (C), they are interpreted as guanosines (G) by the cellular machinery and genomically encoded G alleles at edited sites mimic the function of edited RNAs. The contribution of this hardwiring effect on genome evolution remains obscure. We looked for population genomics signatures of adaptive evolution associated with A-to-I RNA edited sites in humans and Drosophila melanogaster. We found that single nucleotide polymorphisms at edited sites occur 3 (humans) to 15 times (Drosophila) more often than at unedited sites, the nucleotide G is virtually the unique alternative allele at edited sites and G alleles segregate at higher frequency at edited sites than at unedited sites. Our study reveals that coding synonymous and nonsynonymous as well as silent and intergenic A-to-I RNA editing sites are likely adaptive in the distantly related human and Drosophila lineages.

evolutionary biology

TRUmiCount: Correctly counting absolute numbers of molecules using unique molecular identifiers

Counting DNA or RNA molecules using next-generation sequencing (NGS) suffers from amplification biases. Counting unique molecular identifiers (UMIs) instead of reads is still prone to over-estimation due to amplification and sequencing artifacts and under-estimation due to lost molecules. We present an algorithm that corrects for these errors, based on a mechanistic model of the PCR and sequencing process whose parameters have an immediate physical interpretation and are easily estimated. We demonstrate that our algorithm outputs essentially unbiased counts with substantially improved accuracy.

bioinformatics

Complex models of sequence evolution require accurate estimators as exemplified with the invariable site plus Gamma model

The invariable site plus {Gamma} model is widely used to model rate heterogeneity among alignment sites in maximum likelihood and Bayesian phylogenetic analyses. The proof that the invariable site plus continuous {Gamma} model is identifiable (model parameters can be inferred correctly given enough data) has increased the creditability of its application to phylogeny reconstruction. However, most phylogenetic software implement the invariable site plus discrete {Gamma} model, whose identifiability is likely but unproven. How well the parameters of the invariable site plus discrete {Gamma} model are estimated is still disputed. Especially the correlation of the fraction of invariable sites with the fractions of sites with a slow evolutionary rate is discussed as being problematic. We show that optimization heuristics as implemented in frequently used phylogenetic software cannot always reliably estimate the shape parameter, the proportion of invariable sites and the tree length. Here, we propose an improved optimization heuristic that accurately estimates the three parameters. While research efforts mainly focus on tree search methods, our results signify the equal importance of verifying and developing effective estimation methods for complex models of sequence evolution.

evolutionary biology

Thiol-linked alkylation for the metabolic sequencing of RNA

Gene expression profiling by high-throughput sequencing reveals qualitative and quantitative changes in RNA species at steady-state but obscures the intracellular dynamics of RNA transcription, processing and decay. We developed thiol(SH)-linked alkylation for the metabolic sequencing of RNA (SLAM-seq), an orthogonal chemistry-based epitranscriptomics-sequencing technology that uncovers 4-thiouridine (s4U)-incorporation in RNA species at single-nucleotide resolution. In combination with well-established metabolic RNA labeling protocols and coupled to standard, low-input, high-throughput RNA sequencing methods, SLAM-seq enables rapid access to RNA polymerase II-dependent gene expression dynamics in the context of total RNA. When applied to mouse embryonic stem cells, SLAM-seq provides global and transcript-specific insights into pluripotency-associated gene expression. We validated the method by showing that the RNA-polymerase II-dependent transcriptional output scales with Oct4/Sox2/Nanog-defined enhancer activity; and we provide quantitative and mechanistic evidence for transcript-specific RNA turnover mediated by post-transcriptional gene regulatory pathways initiated by microRNAs and N6-methyladenosine. SLAM-seq facilitates the dissection of fundamental mechanisms that control gene expression in an accessible, cost-effective, and scalable manner.\n\nOne Sentence SummaryChemical nucleotide-analog derivatization provides global insights into transcriptional and post-transcriptional gene regulation

molecular biology

GHOST: Recovering Historical Signal from Heterotachously-evolved Sequence Alignments

Molecular sequence data that have evolved under the influence of heterotachous evolutionary processes are known to mislead phylogenetic inference. We introduce the General Heterogeneous evolution On a Single Topology (GHOST) model of sequence evolution, implemented under a maximum-likelihood framework in the phylogenetic program IQ-TREE (http://www.iqtree.org). Simulations show that using the GHOST model, IQ-TREE can accurately recover the tree topology, branch lengths and substitution model parameters from heterotachously-evolved sequences. We develop a model selection algorithm based on simulation results, and investigate the performance of the GHOST model on empirical data by sampling phylogenomic alignments of varying lengths from a plastome alignment. We then carry out inference under the GHOST model on a phylogenomic dataset composed of 248 genes from 16 taxa, where we find the GHOST model concurs with the currently accepted view, placing turtles as a sister lineage of archosaurs, in contrast to results obtained using traditional variable rates-across-sites models. Finally, we apply the model to a dataset composed of a sodium channel gene of 11 fish taxa, finding that the GHOST model is able to infer a subtle component of the historical signal, linked to the previously established convergent evolution of the electric organ in two geographically distinct lineages of electric fish. We compare inference under the GHOST model to partitioning by codon position and show that, owing to the minimization of model constraints, the GHOST model is able to offer unique biological insights when applied to empirical data.

evolutionary biology

Accurate detection of complex structural variations using single molecule sequencing

Structural variations (SVs) are the largest source of genetic variation, but remain poorly understood because of limited genomics technology. Single molecule long read sequencing from Pacific Biosciences and Oxford Nanopore has the potential to dramatically advance the field, although their high error rates challenge existing methods. Addressing this need, we introduce open-source methods for long read alignment (NGMLR, https://github.com/philres/ngmlr) and SV identification (Sniffles, https://github.com/fritzsedlazeck/Sniffles) that enable unprecedented SV sensitivity and precision, including within repeat-rich regions and of complex nested events that can have significant impact on human disorders. Examining several datasets, including healthy and cancerous human genomes, we discover thousands of novel variants using long reads and categorize systematic errors in short-read approaches. NGMLR and Sniffles are further able to automatically filter false events and operate on low amounts of coverage to address the cost factor that has hindered the application of long reads in clinical and research settings.

bioinformatics

UFBoot2: Improving the Ultrafast Bootstrap Approximation

The standard bootstrap (SBS), despite being computationally intensive, is widely used in maximum likelihood phylogenetic analyses. We recently proposed the ultrafast bootstrap approximation (UFBoot) to reduce computing time while achieving more unbiased branch supports than SBS under mild model violations. UFBoot has been steadily adopted as an efficient alternative to SBS and other bootstrap approaches.\n\nHere, we present UFBoot2, which substantially accelerates UFBoot and reduces the risk of overestimating branch supports due to polytomies or severe model violations. Additionally, UFBoot2 provides suitable bootstrap resampling strategies for phylogenomic data. UFBoot2 is 778 and 8.4 times (median) faster than SBS and RAxML rapid bootstrap on tested datasets, respectively. UFBoot2 is implemented in the IQ-TREE software package version 1.6 and freely available at http://www.iqtree.org.

evolutionary biology