bioRxiv ScienceSearch

Biology subjects

Ridge, P. G.

Publications and source records attributed to Ridge, P. G..

3 recordsLinked to original sources

Codon Pairs are Phylogenetically Conserved: Codon pairing as a new class of phylogenetic characters

Identical codon pairing and co-tRNA codon pairing increase translational efficiency within genes when two codons that encode the same amino acid are located within a ribosomal window. By examining both identical and co-tRNA codon pairing across 23 423 species, we determined that both pairing techniques are phylogenetically informative across all domains of life using either an alignment-free or parsimony framework. We also determined that conserved codon pairing typically has a smaller window size than the length of a ribosome. We also analyzed frequencies of codon pairing for each codon to determine which codons are most likely to pair. The alignment-free method does not require orthologous gene annotations and recovers species relationships that are comparable to other alignment-free techniques. Parsimony generally recovers phylogenies that are more congruent with the established phylogenies than the alignment-free method. However, four of the ten taxonomic groups do not have sufficient ortholog annotations and are therefore recoverable using only the alignment-free methods. Since the recovered phylogenies using only codon pairing largely match established phylogenies and are comparable to other algorithms, we propose that codon pairing biases are phylogenetically conserved and should be considered in conjunction with current techniques in future phylogenomic studies. Furthermore, the phylogenetic conservation of codon pairing indicates that codon pairing plays a greater role in the speciation process than previously acknowledged.\n\nAvailabilityAll scripts used to recover and compare phylogenies, including documentation and test files, are freely available on GitHub at https://github.com/ridgelab/codon_pairing.

evolutionary biology

Codon Use and Aversion is Largely Phylogenetically Conserved Across the Tree of Life

Using parsimony, we analyzed codon usages across 12 337 species and 25 727 orthologous genes to rank specific genes and codons according to their phylogenetic signal. We examined each codon within each ortholog to determine the codon usage for each species. In total, 890 814 codons were parsimony informative. Next, we compared species that used a codon with species that did not use the codon. We assessed each codons congruence with species relationships provided in the Open Tree of Life (OTL) and determined the statistical probability of observing these results by random chance. We determined that 25 771 codons had no parallelisms or reversals when mapped to the OTL. Codon usages from orthologous genes spanning many species were 1 109x more likely to be congruent with species relationships in the OTL than would be expected by random chance. Using the OTL as a reference, we show that codon usage is phylogenetically conserved within orthologous genes in archaea, bacteria, plants, mammals, and other vertebrates. We also show how to use our provided framework to test different tree hypotheses by confirming the placement of turtles as sister taxa to archosaurs.\n\nAvailabilityAll scripts, a README, and necessary test files are freely available on GitHub at https://github.com/ridgelab/codon_congruence\n\nContactperry.ridge@byu.edu

evolutionary biology

Systematic analysis of dark and camouflaged genes: disease-relevant genes hiding in plain sight

BackgroundThe human genome contains dark gene regions that cannot be adequately assembled or aligned using standard short-read sequencing technologies, preventing researchers from identifying mutations within these gene regions that may be relevant to human disease. Here, we identify regions that are dark by depth (few mappable reads) and others that are camouflaged (ambiguous alignment), and we assess how well long-read technologies resolve these regions. We further present an algorithm to resolve most camouflaged regions (including in short-read data) and apply it to the Alzheimers Disease Sequencing Project (ADSP; 13142 samples), as a proof of principle.\n\nResultsBased on standard whole-genome lllumina sequencing data, we identified 37873 dark regions in 5857 gene bodies (3635 protein-coding) from pathways important to human health, development, and reproduction. Of the 5857 gene bodies, 494 (8.4%) were 100% dark (142 protein-coding) and 2046 (34.9%) were [≥]5% dark (628 protein-coding). Exactly 2757 dark regions were in protein-coding exons (CDS) across 744 genes. Long-read sequencing technologies from 10x Genomics, PacBio, and Oxford Nanopore Technologies reduced dark CDS regions to approximately 45.1%, 33.3%, and 18.2% respectively. Applying our algorithm to the ADSP, we rescued 4622 exonic variants from 501 camouflaged genes, including a rare, ten-nucleotide frameshift deletion in CR1, a top Alzheimers disease gene, found in only five ADSP cases and zero controls.\n\nConclusionsWhile we could not formally assess the CR1 frameshift mutation in Alzheimers disease (insufficient sample-size), we believe it merits investigating in a larger cohort. There remain thousands of potentially important genomic regions overlooked by short-read sequencing that are largely resolved by long-read technologies.

genomics