bioRxiv ScienceSearch

Biology subjects

Tomislav Domazet-Loso

Publications and source records attributed to Tomislav Domazet-Loso.

3 recordsLinked to original sources

No evidence for phylostratigraphic bias impacting inferences on patterns of gene emergence and evolution

Phylostratigraphy is a computational framework for dating the emergence of sequences (usually genes) in a phylogeny. It has been extensively applied to make inferences on patterns of genome evolution, including patterns of disease gene evolution, ontogeny and de novo gene origination. Phylostratigraphy typically relies on BLAST searches along a species tree, but new simulation studies have raised concerns about the ability of BLAST to detect remote homologues and its impact on phylostratigraphic inferences. These simulations called into question some of our previously published work on patterns of gene emergence and evolution inferred from phylostratigraphy. Here, we re-assessed these simulations and found major problems including unrealistic parameter choices, irreproducibility, statistical flaws and partial representation of results. We found that, even with a possible overall BLAST false negative rate between 5-15%, the large majority (>74%) of sequences assigned to a recent evolutionary origin by phylostratigraphy is unaffected by technical concerns about BLAST. Where the results of the simulations did cast doubt on our previous findings, we repeated our analyses but now excluded all questionable sequences. The originally described patterns remained essentially unchanged. These new analyses strongly support our published inferences, including: genes that emerged after the origin of eukaryotes are more likely to be expressed in the ectoderm than in the endoderm or mesoderm in Drosophila, and the de novo emergence of protein-coding genes from non-genic sequences occurs through proto-gene intermediates in yeast. We conclude that BLAST is an appropriate and sufficiently sensitive tool in phylostratigraphic analysis.

Genomics

Thanatotranscriptome: genes actively expressed after organismal death

A continuing enigma in the study of biological systems is what happens to highly ordered structures, far from equilibrium, when their regulatory systems suddenly become disabled. In life, genetic and epigenetic networks precisely coordinate the expression of genes -- but in death, it is not known if gene expression diminishes gradually or abruptly stops or if specific genes are involved. We investigated the unwinding of the clock by identifying upregulated genes, assessing their functions, and comparing their transcriptional profiles through postmortem time in two species, mouse and zebrafish. We found transcriptional abundance profiles of 1,063 genes were significantly changed after death of healthy adult animals in a time series spanning from life to 48 or 96 h postmortem. Ordination plots revealed non-random patterns in profiles by time. While most thanatotranscriptome (thanatos-, Greek defn. death) transcript levels increased within 0.5 h postmortem, some increased only at 24 and 48 h. Functional characterization of the most abundant transcripts revealed the following categories: stress, immunity, inflammation, apoptosis, transport, development, epigenetic regulation, and cancer. The increase of transcript abundance was presumably due to thermodynamic and kinetic controls encountered such as the activation of epigenetic modification genes responsible for unraveling the nucleosomes, which enabled transcription of previously silenced genes (e.g., development genes). The fact that new molecules were synthesized at 48 to 96 h postmortem suggests sufficient energy and resources to maintain self-organizing processes. A step-wise shutdown occurs in organismal death that is manifested by the apparent upregulation of genes with various abundance maxima and durations. The results are of significance to transplantology and molecular biology.

Systems Biology

gmos: Rapid detection of genome mosaicism over short evolutionary distances

Prokaryotic and viral genomes are often altered by recombination and horizontal gene transfer. The existing methods for detecting recombination are primarily aimed at viral genomes or sets of loci, since the expensive computation of underlying statistical models often hinders the comparison of complete prokaryotic genomes. As an alternative, alignment-free solutions are more efficient, but cannot map (align) a query to subject genomes. To address this problem, we have developed gmos (Genome MOsaic Structure), a new program that determines the mosaic structure of query genomes when compared to a set of closely related subject genomes. The program first computes local alignments between query and subject genomes and then reconstructs the query mosaic structure by choosing the best local alignment for each query region. To accomplish the analysis quickly, the program mostly relies on pairwise alignments and constructs multiple sequence alignments over short overlapping subject regions only when necessary. This fine-tuned implementation achieves an efficiency comparable to an alignment-free tool. The program performs well for simulated and real data sets of closely related genomes and can be used for fast recombination detection; for instance, when a new prokaryotic pathogen is discovered. As an example, gmos was used to detect genome mosaicism in a pathogenic Enterococcus faecium strain compared to seven closely related genomes. The analysis took less than two minutes on a single 2.1 GHz processor. The output is available in fasta format and can be visualized using an accessory program, gmosDraw (freely available with gmos).

Bioinformatics