bioRxiv Science⌕ Search

Biology subjects

Zaravinos, A.

Publications and source records attributed to Zaravinos, A..

3 recordsLinked to original sources

Genome-wide mapping of persistent long-range correlation in the complete telomere-to-telomere human reference genome

Long-range correlations (LRCs) in DNA sequences have been reported for decades, but their interpretation has been limited by incomplete representation of repetitive and structurally complex regions in earlier human genome assemblies. Here, we use the complete telomere-to-telomere (T2T) human reference genome to systematically map LRC structure across all human chromosomes. Most chromosomes exhibited persistent long-range dependence (LRD), with chromosome-specific scaling intervals spanning kilobase to megabase scales and wavelet Hurst exponents consistently above the uncorrelated expectation of 0.5. Spatially resolved Hurst landscapes revealed that this signal is not uniformly distributed along chromosomes, but instead reflects a heterogeneous genomic mosaic. For the purine-pyrimidine encoding, chromosomes 9, 15, X, and Y showed complex fluctuation profiles not adequately summarized by a single scaling exponent; multifractal detrended fluctuation analysis (MF-DFA) confirmed broad, q-dependent multiscaling consistent with multifractal behavior in these chromosomes, consistent with heterogeneous contributions from satellite-rich and structurally distinct sequence compartments. Surrogate analyses further showed that the observed correlations exceed expectations from base composition alone and are not fully explained by local sequence structure preserved in block-shuffled controls. However, short-memory Autoregressive Moving-Average (ARMA) surrogates reproduced the observed H range in a subset of chromosomes, indicating that the strength of evidence for LRD is chromosome-dependent. Together, these results demonstrate that LRC is a widespread but spatially heterogeneous property of the complete human genome. The T2T assembly reveals that previously unresolved repetitive and satellite-rich regions are strongly associated with chromosome-scale and local scaling variation, providing a framework for future studies of genome evolution, chromatin structure, recombination, structural variation, and genomic instability.

genomics↗

ZSeekerDB: A database of Z-nucleic acid sequences across organismal genomes

Alternative (non-B) nucleic acid structures such as Z-nucleic acids are emerging as key regulators of genome function. The density and distribution of Z-nucleic acid sequences across organismal and viral genomes can provide insights into their biological roles and evolutionary trajectory. Using our recently developed ZSeeker algorithm, we systematically analyzed over 280,000 organismal genome assemblies, identifying more than 850 million putative Z-forming loci. We also incorporated genomic coordinates, Z-score, and taxonomic metadata, enabling cross-species comparative and functional analyses. We introduce ZSeekerDB, the first large-scale, multi-taxon database cataloging Z-nucleic acid sequences, across organisms representing all major branches of life. ZSeekerDB enables interactive searches, visualizations, and downloads of Z-nucleic acid sequence data for independent analysis. ZSeekerDB is implemented as a web-portal for browsing, analyzing and downloading Z-forming loci, publicly available at https://zseeker-db.com/.

bioinformatics↗

Quadrupia: Derivation of G-quadruplexes for organismal genomes across the tree of life

G-quadruplex DNA structures exhibit a profound influence on essential biological processes, including transcription, replication, telomere maintenance, and genomic stability. These structures have demonstrably shaped organismal evolution. However, a comprehensive, organism-wide G-quadruplex map encompassing the diversity of life has remained elusive. Here, we introduce Quadrupia, the most extensive and well-characterized G-quadruplex database to date, facilitating the exploration of G-quadruplex structures across the evolutionary spectrum. Quadrupia has identified G-quadruplex sequences in 108,449 reference genomes, with a total of 140,181,277 G-quadruplexes. The database also hosts a collection of 319,784 G-quadruplex clusters of 20 or more members, annotated by taxonomic distributions, multiple sequence alignments, profile Hidden Markov Models and cross-references to G-quadruplex 3D structures. Examination of G-quadruplexes across functional genomic elements in different taxa indicates preferential orientation and positioning, with significant differences between individual taxonomic groups. For example, we find that G-quadruplexes in bacteria with a single replication origin display profound preference for the leading orientation. Finally, we experimentally validate the most frequently observed G-quadruplexes using CD-spectroscopy, UV melting, and fluorescent-based approaches. Quadrupia is publicly available through https://www.pavlopoulos-lab.org/quadrupia.

genomics↗