bioRxiv Science⌕ Search

Biology subjects

Megalovasilis, G.

Publications and source records attributed to Megalovasilis, G..

3 recordsLinked to original sources

Genome-wide mapping of persistent long-range correlation in the complete telomere-to-telomere human reference genome

Long-range correlations (LRCs) in DNA sequences have been reported for decades, but their interpretation has been limited by incomplete representation of repetitive and structurally complex regions in earlier human genome assemblies. Here, we use the complete telomere-to-telomere (T2T) human reference genome to systematically map LRC structure across all human chromosomes. Most chromosomes exhibited persistent long-range dependence (LRD), with chromosome-specific scaling intervals spanning kilobase to megabase scales and wavelet Hurst exponents consistently above the uncorrelated expectation of 0.5. Spatially resolved Hurst landscapes revealed that this signal is not uniformly distributed along chromosomes, but instead reflects a heterogeneous genomic mosaic. For the purine-pyrimidine encoding, chromosomes 9, 15, X, and Y showed complex fluctuation profiles not adequately summarized by a single scaling exponent; multifractal detrended fluctuation analysis (MF-DFA) confirmed broad, q-dependent multiscaling consistent with multifractal behavior in these chromosomes, consistent with heterogeneous contributions from satellite-rich and structurally distinct sequence compartments. Surrogate analyses further showed that the observed correlations exceed expectations from base composition alone and are not fully explained by local sequence structure preserved in block-shuffled controls. However, short-memory Autoregressive Moving-Average (ARMA) surrogates reproduced the observed H range in a subset of chromosomes, indicating that the strength of evidence for LRD is chromosome-dependent. Together, these results demonstrate that LRC is a widespread but spatially heterogeneous property of the complete human genome. The T2T assembly reveals that previously unresolved repetitive and satellite-rich regions are strongly associated with chromosome-scale and local scaling variation, providing a framework for future studies of genome evolution, chromatin structure, recombination, structural variation, and genomic instability.

genomics↗

Characterization of Z-DNA dynamics across the tree of life

Z-DNA/Z-RNA is an alternative left-handed nucleic acid conformation with established and emerging roles in gene regulation, immunity, and genome instability. However, its occurrence dynamics and lineage specificity across the tree of life have not yet been fully characterized. Utilizing the recently developed and improved Z-DNA searching tool, ZSeeker, we analyzed 281,139 complete organismal genomes, including multiple Telomere-to-Telomere genome assemblies, and generated genome-wide Z-nucleic acid maps, examined their topography, and compared them to dinucleotide-preserving controls. Cellular genomes featured pervasive Z-DNA enrichment relative to expectation, with enrichments of [~]1.5 and [~]1.7-fold in Bacteria and Archaea and [~]3-fold in Eukaryota. In contrast, Viruses exhibited large differences between lineages, with modest enrichment in several DNA viral groups and pronounced depletion across RNA clades, most notably Influenza A/B strains. We built a LASSO regression model trained on non-Influenza viruses (cross-validated R{superscript 2} {approx} 0.73), which identified GC content, genome type, and host type as the leading predictors for Z-nucleic acid density, yet it significantly over-predicted Z-RNA density in Influenza A/B. More than 99% of assemblies exceeded the +2 SD threshold, and a "typical Influenza" genome was predicted at 2.76 bp/kb compared to [~]0.016 bp/kb observed (a [~]170-fold overestimation based on chance alone). Together, these results reveal domain- and lineage-specific regimes: cellular genomes are enriched for Z-DNA consistent with regulatory roles, whereas influenza viruses appear to have undergone strong, lineage-specific depletion of Z-RNA-forming sequences, likely reflecting evolutionary pressure tied to host sensing pathways.

evolutionary biology↗

ZSeekerDB: A database of Z-nucleic acid sequences across organismal genomes

Alternative (non-B) nucleic acid structures such as Z-nucleic acids are emerging as key regulators of genome function. The density and distribution of Z-nucleic acid sequences across organismal and viral genomes can provide insights into their biological roles and evolutionary trajectory. Using our recently developed ZSeeker algorithm, we systematically analyzed over 280,000 organismal genome assemblies, identifying more than 850 million putative Z-forming loci. We also incorporated genomic coordinates, Z-score, and taxonomic metadata, enabling cross-species comparative and functional analyses. We introduce ZSeekerDB, the first large-scale, multi-taxon database cataloging Z-nucleic acid sequences, across organisms representing all major branches of life. ZSeekerDB enables interactive searches, visualizations, and downloads of Z-nucleic acid sequence data for independent analysis. ZSeekerDB is implemented as a web-portal for browsing, analyzing and downloading Z-forming loci, publicly available at https://zseeker-db.com/.

bioinformatics↗