bioRxiv Science⌕ Search

Biology subjects

Vasquez, K. M.

Publications and source records attributed to Vasquez, K. M..

8 recordsLinked to original sources

HSeeker: an algorithm for systematic H-DNA sequence identification

H-DNA is a naturally occurring intramolecular DNA triplex structure formed by Hoogsteen hydrogen bonds at homopurine-homopyrimidine mirror repeats and has functional roles in gene regulation, genome instability, and human disease. The existing H-DNA detection tools capture only a subset of H-DNA sequences, often missing relevant sequence features or failing to assess key aspects of structural stability. To address this gap, we present "HSeeker", a state-of-the-art computational tool that compiles a three-part algorithm to identify and score potential H-DNA-forming sequences. Using a center-outward search algorithm approach, HSeeker evaluates candidate hinge positions and spacer lengths while allowing configurable mirror mismatches and purine-pyrimidine composition thresholds. The greedy overlap removal phase resolves overlapping candidates by retaining the longest and most compact motif within each overlapping region. Finally, the thermodynamic stability scoring algorithm evaluates the candidate motifs using an experimentally informed scoring model that incorporates Hoogsteen G-G and A-A bonds, consecutive-pair stacking, and imposes penalties for mismatches and disrupted stacks. The scoring procedure also optimizes motif boundaries by trimming weak terminal positions and reassigning unstable arm positions to the spacer. HSeeker reports genomic coordinates, sequence information, pairing and stacking components, and an overall stability score. HSeeker is also user-friendly, available as a Python package and as a web application, supporting configurable, high-throughput analysis and exportability of predicted H-DNA motifs. HSeeker provides an accessible and reproducible framework for investigating the distribution and potential stability of H-DNA-forming sequences across genomic datasets.

genomics↗

Characterization of Z-DNA dynamics across the tree of life

Z-DNA/Z-RNA is an alternative left-handed nucleic acid conformation with established and emerging roles in gene regulation, immunity, and genome instability. However, its occurrence dynamics and lineage specificity across the tree of life have not yet been fully characterized. Utilizing the recently developed and improved Z-DNA searching tool, ZSeeker, we analyzed 281,139 complete organismal genomes, including multiple Telomere-to-Telomere genome assemblies, and generated genome-wide Z-nucleic acid maps, examined their topography, and compared them to dinucleotide-preserving controls. Cellular genomes featured pervasive Z-DNA enrichment relative to expectation, with enrichments of [~]1.5 and [~]1.7-fold in Bacteria and Archaea and [~]3-fold in Eukaryota. In contrast, Viruses exhibited large differences between lineages, with modest enrichment in several DNA viral groups and pronounced depletion across RNA clades, most notably Influenza A/B strains. We built a LASSO regression model trained on non-Influenza viruses (cross-validated R{superscript 2} {approx} 0.73), which identified GC content, genome type, and host type as the leading predictors for Z-nucleic acid density, yet it significantly over-predicted Z-RNA density in Influenza A/B. More than 99% of assemblies exceeded the +2 SD threshold, and a "typical Influenza" genome was predicted at 2.76 bp/kb compared to [~]0.016 bp/kb observed (a [~]170-fold overestimation based on chance alone). Together, these results reveal domain- and lineage-specific regimes: cellular genomes are enriched for Z-DNA consistent with regulatory roles, whereas influenza viruses appear to have undergone strong, lineage-specific depletion of Z-RNA-forming sequences, likely reflecting evolutionary pressure tied to host sensing pathways.

evolutionary biology↗

ZSeekerDB: A database of Z-nucleic acid sequences across organismal genomes

Alternative (non-B) nucleic acid structures such as Z-nucleic acids are emerging as key regulators of genome function. The density and distribution of Z-nucleic acid sequences across organismal and viral genomes can provide insights into their biological roles and evolutionary trajectory. Using our recently developed ZSeeker algorithm, we systematically analyzed over 280,000 organismal genome assemblies, identifying more than 850 million putative Z-forming loci. We also incorporated genomic coordinates, Z-score, and taxonomic metadata, enabling cross-species comparative and functional analyses. We introduce ZSeekerDB, the first large-scale, multi-taxon database cataloging Z-nucleic acid sequences, across organisms representing all major branches of life. ZSeekerDB enables interactive searches, visualizations, and downloads of Z-nucleic acid sequence data for independent analysis. ZSeekerDB is implemented as a web-portal for browsing, analyzing and downloading Z-forming loci, publicly available at https://zseeker-db.com/.

bioinformatics↗

Landscape and mutational dynamics of G-quadruplexes in the complete human genome and in haplotypes of diverse ancestry

G-quadruplexes (G4s) are alternative DNA structures with diverse biological roles, but their examination in highly repetitive parts of the human genome has been hindered by the lack of reliable sequencing technologies. Recent long-read based genome assemblies have enabled their characterization in previously inaccessible parts of the human genome. Here, we examine the topography and genomic instability of potential G4-forming sequences in the gap-less, reference human genome assembly and in 88 haplotypes of diverse ancestry. We report that G4s are highly enriched in specific repetitive regions, including in certain centromeric and pericentromeric repeat types, and in ribosomal DNA arrays, and experimentally validate the most prevalent G4s detected. G4s tend to have lower methylation than expected throughout the human genome and are genomically unstable, showing an excess of all mutation types, including substitutions, insertions and deletions and most prominently structural variants. Finally, we show that G4s are consistently enriched at PRDM9 binding sites, a protein involved in meiotic recombination. Together, our findings establish G4s as dynamic and functionally significant elements of the human genome and highlight new avenues for investigating their contributions to human disease and evolution.

genomics↗

ZSeeker: An optimized algorithm for Z-DNA detection in genomic sequences

Z-DNA is an alternative left-handed helical form of DNA with a zigzag-shaped backbone that differs from the right-handed canonical B-DNA helix. Z-DNA has been implicated in various biological processes, including transcription, replication, and DNA repair, and can induce genetic instability. Repetitive sequences of alternating purines and pyrimidines have the potential to adopt Z-DNA structures. ZSeeker is a novel computational tool developed for the accurate detection of potential Z-DNA-forming sequences in genomes, addressing limitations of prior methods. By introducing a novel methodology informed and validated by experimental data, ZSeeker enables the refined detection of potential Z-DNA-forming sequences. Built both as a standalone Python package and as an accessible web interface, ZSeeker allows users to input genomic sequences, adjust detection parameters, and view potential Z-DNA sequence distributions and Z-scores via downloadable visualizations. Our Web Platform provides a no-code solution for Z-DNA identification, with a focus on accessibility, user-friendliness, speed and customizability. By providing efficient, high-throughput analysis and enhanced detection accuracy, ZSeeker has the potential to support significant advancements in understanding the roles of Z-DNA in normal cellular functions, genetic instability, and its implications in human diseases. AvailabilityZSeeker is released as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/ZSeeker. A web-interface of ZSeeker is publicly available at https://zseeker.netlify.app/.

bioinformatics↗

Characterization of hairpin loops and cruciforms across 118,065 genomes spanning the tree of life

Inverted repeats (IRs) can form alternative DNA secondary structures called hairpins and cruciforms, which have a multitude of functional roles and have been associated with genomic instability. However, their prevalence across diverse organismal genomes remains only partially understood. Here, we examine the prevalence of IRs across 118,065 complete organismal genomes. Our comprehensive analysis across taxonomic subdivisions reveals significant differences in the distribution, frequency, and biophysical properties of perfect IRs among these genomes. We identify a total of 29,589,132 perfect IRs and show a highly variable density across different organisms, with strikingly distinct patterns observed in Viruses, Bacteria, Archaea, and Eukaryota. We report IRs with perfect arms of extreme lengths, which can extend to hundreds of thousands of base pairs. Our findings demonstrate a strong correlation between IR density and genome size, revealing that Viruses and Bacteria possess the highest density, whereas Eukaryota and Archaea exhibit the lowest relative to their genome size. Additionally, the study reveals the enrichment of IRs at transcription start and termination end sites in prokaryotes and Viruses and underscores their potential roles in gene regulation and genome organization. Through a comprehensive overview of the distribution and characteristics of IRs in a wide array of organisms, this largest-scale analysis to date sheds light on the functional significance of inverted repeats, their contribution to genomic instability, and their evolutionary impact across the tree of life.

genomics↗

Quadrupia: Derivation of G-quadruplexes for organismal genomes across the tree of life

G-quadruplex DNA structures exhibit a profound influence on essential biological processes, including transcription, replication, telomere maintenance, and genomic stability. These structures have demonstrably shaped organismal evolution. However, a comprehensive, organism-wide G-quadruplex map encompassing the diversity of life has remained elusive. Here, we introduce Quadrupia, the most extensive and well-characterized G-quadruplex database to date, facilitating the exploration of G-quadruplex structures across the evolutionary spectrum. Quadrupia has identified G-quadruplex sequences in 108,449 reference genomes, with a total of 140,181,277 G-quadruplexes. The database also hosts a collection of 319,784 G-quadruplex clusters of 20 or more members, annotated by taxonomic distributions, multiple sequence alignments, profile Hidden Markov Models and cross-references to G-quadruplex 3D structures. Examination of G-quadruplexes across functional genomic elements in different taxa indicates preferential orientation and positioning, with significant differences between individual taxonomic groups. For example, we find that G-quadruplexes in bacteria with a single replication origin display profound preference for the leading orientation. Finally, we experimentally validate the most frequently observed G-quadruplexes using CD-spectroscopy, UV melting, and fluorescent-based approaches. Quadrupia is publicly available through https://www.pavlopoulos-lab.org/quadrupia.

genomics↗

MoCoLo: a testing framework for motif co-localization

Sequence-level data offers insights into biological processes through the interaction of two or more genomic features from the same or different molecular data types. Within motifs, this interaction is often explored via the co-occurrence of feature genomic tracks using fixed-segments or analytical tests that respectively require window size determination and risk of false positives from over-simplified models. Moreover, methods for robustly examining the co-localization of genomic features, and thereby understanding their spatial interaction, have been elusive. We present a new analytical method for examining feature interaction by introducing the notion of reciprocal co-occurrence, define statistics to estimate it, and hypotheses to test for it. Our approach leverages conditional motif co-occurrence events between features to infer their co-localization. Using reverse conditional probabilities and introducing a novel simulation approach that retains motif properties (e.g., length, guanine-content), our method further accounts for potential confounders in testing. As a proof-of-concept, MoCoLo confirmed the co-occurrence of histone markers in a breast cancer cell line. As a novel analysis, MoCoLo identified significant co-localization of oxidative DNA damage within non-B DNA forming regions that significantly differed between non-B DNA structures. Altogether, these findings demonstrate the potential utility of MoCoLo for testing spatial interactions between genomic features via their co-localization.

bioinformatics↗