bioRxiv ScienceSearch

Biology subjects

Ruggiero, R. P.

Publications and source records attributed to Ruggiero, R. P..

2 recordsLinked to original sources

VARIATION IN BASE COMPOSITION UNDERLIES FUNCTIONAL AND EVOLUTIONARY DIVERGENCE IN NON-LTR RETROTRANSPOSONS

BackgroundNon-LTR retrotransposons often exhibit base composition that is markedly different from the nucleotide content of their hosts gene. For instance, the mammalian L1 element is AT-rich with a strong A bias on the positive strand, which results in a reduced transcription. It is plausible that the A-richness of mammalian L1 is a self-regulatory mechanism reflecting a trade-off between transposition efficiency and the deleterious effect of L1 on its host. We examined if the A-richness of L1 is a general feature of non-LTR retrotransposons or if different clades of elements have evolved different nucleotide content. We also investigated if elements belonging to the same clade evolved towards different base composition in different genomes or if elements from the same clades evolved towards similar base composition in the same genome.\n\nResultsWe found that non-LTR retrotransposons differ in base composition among clades within the same host but also that elements belonging to the same clade differ in base composition among hosts. We showed that nucleotide content remains constant within the same host over extended period of evolutionary time, despite mutational patterns that should drive nucleotide content away from the observed base composition.\n\nConclusionsOur results suggest that base composition is evolving under selection and may be reflective of the long-term co-evolution between non-LTR retrotransposons and their host. Finally, the coexistence of elements with drastically different base composition suggests that these elements may be using different strategies to persist and multiply in the genome of their host.

genomics

Finding and extending ancient simple sequence repeat-derived regions in the human genome

BackgroundPreviously, 3% of the human genome has been annotated as simple sequence repeats (SSRs), similar to the proportion annotated as protein coding. The origin of much of the genome is not well annotated, however, and some of the unidentified regions are likely to be ancient SSR-derived regions not identified by current methods. The identification of these regions is complicated because SSRs appear to evolve through complex cycles of expansion and contraction, often interrupted by mutations that alter both the repeated motif and mutation rate. We applied an empirical, kmer-based, approach to identify genome regions that are likely derived from SSRs.\n\nResultsThe sequences flanking annotated SSRs are enriched for similar sequences and for SSRs with similar motifs, suggesting that the evolutionary remains of SSR activity abound in regions near obvious SSRs. Using our previously described P-clouds approach, we identified SSR-clouds, groups of similar kmers (or oligos) that are enriched near a training set of unbroken SSR loci, and then used the SSR-clouds to detect likely SSR-derived regions throughout the genome.\n\nConclusionsOur analysis indicates that the amount of likely SSR-derived sequence in the human genome is 6.77%, over twice as much as previous estimates, including millions of newly identified ancient SSR-derived loci. SSR-clouds identified poly-A sequences adjacent to transposable element termini in over 74% of the oldest class of Alu (roughly, AluJ), validating the sensitivity of the approach. Poly-As annotated by SSR-clouds also had a length distribution that was more consistent with their poly-A origins, with mean about 35 bp even in older Alus. This work demonstrate that the high sensitivity provided by SSR-Clouds improves the detection of SSR-derived regions and will enable deeper analysis of how decaying repeats contribute to genome structure.

genomics