bioRxiv Science⌕ Search

Biology subjects

Chantzi, N.

Publications and source records attributed to Chantzi, N..

10 recordsLinked to original sources

ZSeeker: An optimized algorithm for Z-DNA detection in genomic sequences

Z-DNA is an alternative left-handed helical form of DNA with a zigzag-shaped backbone that differs from the right-handed canonical B-DNA helix. Z-DNA has been implicated in various biological processes, including transcription, replication, and DNA repair, and can induce genetic instability. Repetitive sequences of alternating purines and pyrimidines have the potential to adopt Z-DNA structures. ZSeeker is a novel computational tool developed for the accurate detection of potential Z-DNA-forming sequences in genomes, addressing limitations of prior methods. By introducing a novel methodology informed and validated by experimental data, ZSeeker enables the refined detection of potential Z-DNA-forming sequences. Built both as a standalone Python package and as an accessible web interface, ZSeeker allows users to input genomic sequences, adjust detection parameters, and view potential Z-DNA sequence distributions and Z-scores via downloadable visualizations. Our Web Platform provides a no-code solution for Z-DNA identification, with a focus on accessibility, user-friendliness, speed and customizability. By providing efficient, high-throughput analysis and enhanced detection accuracy, ZSeeker has the potential to support significant advancements in understanding the roles of Z-DNA in normal cellular functions, genetic instability, and its implications in human diseases. AvailabilityZSeeker is released as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/ZSeeker. A web-interface of ZSeeker is publicly available at https://zseeker.netlify.app/.

bioinformatics↗

Unraveling diversity by isolating peptide sequences specific to distinct taxonomic groups

The identification of succinct, universal fingerprints that enable the characterization of individual taxonomies can reveal insights into trait development and can have widespread applications in pathogen diagnostics, human healthcare, ecology and the characterization of biomes. Here, we investigated the existence of peptide k-mer sequences that are exclusively present in a specific taxonomy and absent in every other taxonomic level, termed taxonomic quasi-primes. By analyzing proteomes across 24,073 species, we identified quasi-prime peptides specific to superkingdoms, kingdoms, and phyla, uncovering their taxonomic distributions and functional relevance. These peptides exhibit remarkable sequence uniqueness at six- and seven-amino- acid lengths, offering insights into evolutionary divergence and lineage-specific adaptations. Moreover, we show that human quasi-prime loci are more prone to harboring pathogenic variants, underscoring their functional significance. This study introduces taxonomic quasi-primes and offers insights into their contributions to proteomic diversity, evolutionary pathways, and functional adaptations across the tree of life, while emphasizing their potential impact on human health and disease.

bioinformatics↗

invertiaDB: A Database of Inverted Repeats Across Organismal Genomes

Inverted repeats are repetitive elements that can form hairpin and cruciform structures. They are linked to genomic instability, however they also have various biological functions. Their distribution differs markedly across taxonomic groups in the tree of life, and they exhibit high polymorphism due to their inherent genomic instability. Advances in sequencing technologies and declined costs have enabled the generation of an ever-growing number of complete genomes for organisms across taxonomic groups in the tree of life. However, a comprehensive database encompassing inverted repeats across diverse organismal genomes has been lacking. We present InvertiaDB, the first comprehensive database of inverted repeats spanning multiple taxa, featuring repeats identified in the genomes of 118,070 organisms across all major taxonomic groups. The database currently hosts 30,067,666 inverted repeat sequences, serving as a centralized, user-friendly repository to perform searches, interactive visualization, and download existing inverted repeat data for independent analysis. invertiaDB is implemented as a web portal for browsing, analyzing and downloading inverted repeat data. invertiaDB is publicly available at https://invertiadb.netlify.app/homepage.html.

genomics↗

Characterization of hairpin loops and cruciforms across 118,065 genomes spanning the tree of life

Inverted repeats (IRs) can form alternative DNA secondary structures called hairpins and cruciforms, which have a multitude of functional roles and have been associated with genomic instability. However, their prevalence across diverse organismal genomes remains only partially understood. Here, we examine the prevalence of IRs across 118,065 complete organismal genomes. Our comprehensive analysis across taxonomic subdivisions reveals significant differences in the distribution, frequency, and biophysical properties of perfect IRs among these genomes. We identify a total of 29,589,132 perfect IRs and show a highly variable density across different organisms, with strikingly distinct patterns observed in Viruses, Bacteria, Archaea, and Eukaryota. We report IRs with perfect arms of extreme lengths, which can extend to hundreds of thousands of base pairs. Our findings demonstrate a strong correlation between IR density and genome size, revealing that Viruses and Bacteria possess the highest density, whereas Eukaryota and Archaea exhibit the lowest relative to their genome size. Additionally, the study reveals the enrichment of IRs at transcription start and termination end sites in prokaryotes and Viruses and underscores their potential roles in gene regulation and genome organization. Through a comprehensive overview of the distribution and characteristics of IRs in a wide array of organisms, this largest-scale analysis to date sheds light on the functional significance of inverted repeats, their contribution to genomic instability, and their evolutionary impact across the tree of life.

genomics↗

The repertoire of short tandem repeats across the tree of life

Short tandem repeats (STRs) are widespread, dynamic repetitive elements with a number of biological functions and relevance to human diseases. However, their prevalence across taxa remains poorly characterized. Here we examined the impact of STRs in the genomes of 117,253 organisms spanning the tree of life. We find that there are large differences in the frequencies of STRs between organismal genomes and these differences are largely driven by the taxonomic group an organism belongs to. Using simulated genomes, we find that on average there is no enrichment of STRs in bacterial and archaeal genomes, suggesting that these genomes are not particularly repetitive. In contrast, we find that eukaryotic genomes are orders of magnitude more repetitive than expected. STRs are preferentially located at functional loci at specific taxa. Finally, we utilize the recently completed Telomere-to-Telomere genomes of human and other great apes, and find that STRs are highly abundant and variable between primate species, particularly in peri/centromeric regions. We conclude that STRs have expanded in eukaryotic and viral lineages and not in archaea or bacteria, resulting in large discrepancies in genomic composition.

genomics↗

Ribosomal DNA arrays are the most H-DNA rich element in the human genome

Repetitive DNA sequences can form non-canonical structures such as H-DNA which is an intramolecular triplex DNA structure. The new Telomere-to-Telomere (T2T) genome assembly for the human genome has eliminated gaps, enabling the examination of highly repetitive regions including centromeric and pericentromeric repeats and ribosomal DNA arrays. This gapless assembly allows for the examination of the distribution of H-DNA sequences in parts of the human genome that were not previously annotated. We find that H-DNA appears once every 30,000 bps in the human genome. Its distribution is highly inhomogeneous with H-DNA motif hotspots being detectable in acrocentric chromosomes. Ribosomal DNA arrays in acrocentric chromosomes are the genomic element with the highest H-DNA enrichment, with 13.22% of total H-DNA motifs being found in ribosomal DNA arrays, representing a 42.65-fold enrichment over what would be expected by chance. Across the acrocentric chromosomes we report that 55.87% of all H-DNA motifs found in these chromosomes are in rDNA array loci. The H-DNA motifs are primarily found in the intergenic spacer regions of the ribosomal DNA arrays, generating repeated clusters. We also discover that binding sites for PRDM9, a protein that regulates the formation of double-strand breaks and determines the meiotic recombination hotspots in humans and most mammals, are over 5-fold enriched for H-DNA motifs. Finally, we provide evidence that our findings are consistent in other non-human great ape genomes. We conclude that ribosomal DNA arrays are the most enriched genomic loci for H-DNA sequences in human and other great ape genomes.

genomics↗

Quadrupia: Derivation of G-quadruplexes for organismal genomes across the tree of life

G-quadruplex DNA structures exhibit a profound influence on essential biological processes, including transcription, replication, telomere maintenance, and genomic stability. These structures have demonstrably shaped organismal evolution. However, a comprehensive, organism-wide G-quadruplex map encompassing the diversity of life has remained elusive. Here, we introduce Quadrupia, the most extensive and well-characterized G-quadruplex database to date, facilitating the exploration of G-quadruplex structures across the evolutionary spectrum. Quadrupia has identified G-quadruplex sequences in 108,449 reference genomes, with a total of 140,181,277 G-quadruplexes. The database also hosts a collection of 319,784 G-quadruplex clusters of 20 or more members, annotated by taxonomic distributions, multiple sequence alignments, profile Hidden Markov Models and cross-references to G-quadruplex 3D structures. Examination of G-quadruplexes across functional genomic elements in different taxa indicates preferential orientation and positioning, with significant differences between individual taxonomic groups. For example, we find that G-quadruplexes in bacteria with a single replication origin display profound preference for the leading orientation. Finally, we experimentally validate the most frequently observed G-quadruplexes using CD-spectroscopy, UV melting, and fluorescent-based approaches. Quadrupia is publicly available through https://www.pavlopoulos-lab.org/quadrupia.

genomics↗

Nucleic Quasi-Primes: Identification of the Shortest Unique Oligonucleotide Sequences in a Species

Despite the exponential increase in sequencing information driven by massively parallel DNA sequencing technologies, universal and succinct genomic fingerprints for each organism are still missing. Identifying the shortest species-specific nucleic sequences offers insights into species evolution and holds potential practical applications in agriculture, wildlife conservation, and healthcare. We propose a new method for sequence analysis termed nucleic "quasi-primes", the shortest occurring sequences in each of 45,785 organismal reference genomes, present in one genome and absent from every other examined genome. In the human genome, we find that the genomic loci of nucleic quasi-primes are most enriched for genes associated with brain development and cognitive function. In a single-cell case study focusing on the human primary motor cortex, nucleic quasi-prime genes account for a significantly larger proportion of the variation based on average gene expression. Non-neuronal cell-types, including astrocytes, endothelial cells, microglia perivascular-macrophages, oligodendrocytes, and vascular and leptomeningeal cells, exhibited significant activation of quasi-prime containing gene associations related to cancer, while simultaneously suppressing quasi-prime containing genes were associated with cognitive, mental, and developmental disorders. We also show that human disease-causing variants, eQTLs, mQTLs and sQTLs are 4.43-fold, 4.34-fold, 4.29-fold and 4.21-fold enriched at human quasi-prime loci, respectively. These findings indicate that nucleic quasi-primes are genomic loci linked to the evolution of species-specific traits and in humans they provide insights in the development of cognitive traits and human diseases, including neurodevelopmental disorders.

genomics↗

kmerDB: A Database Encompassing the Set of Genomic and Proteomic Sequence Information for Each Species

The rapid decline in sequencing cost has enabled the generation of reference genomes and proteomes for a growing number of organisms. However, at the present time, there is no established repository that provides information about organism-specific genomic and proteomic sequences of certain lengths, also known as kmers, that are either present or absent in each genome or proteome. In this article, we present kmerDB, a database accessible through an interactive web interface that provides kmer based information from genomic and proteomic sequences in a systematic way. kmerDB currently contains 202,340,859,107 base pairs and 19,304,903,356 amino acids, spanning 45,785 and 22,386 reference genomes and proteomes, respectively, as well as 14,658,776 and 149,264,442 genomic and proteomic species-specific sequences, termed quasi-primes. Additionally, we provide access to 5,186,757 nucleic and 214,904,089 peptide sequences that are absent from every genome and proteome, termed primes. kmerDB features a user-friendly interface offering various search options and filters for easy parsing and searching. The service is available at: www.kmerdb.com.

genomics↗

The determinants of the rarity of nucleic and peptide short sequences in nature

The prevalence of nucleic and peptide short sequences across organismal genomes and proteomes has not been thoroughly investigated. Here we examined 45,785 reference genomes and 21,871 reference proteomes, spanning archaea, bacteria, viruses and eukaryotes to calculate the rarity of short sequences in them. To capture this, we developed a metric of the rarity of each sequence in nature, the Anti-Kardashian index. We find that the frequency of certain dipeptides in rare oligopeptide sequences is hundreds of times lower than expected, which is not the case for any dinucleotides. We also generate predictive regression models that infer the rarity of nucleic and proteomic sequences in nature. For six-mer peptide kmers the R2 performance of the regression models based on amino acid and dipeptide content is 0.816, whereas models based on physicochemical features achieve an R2 of 0.788. For twelve-mer nucleic kmers the R2 performance of our models based on mono and dinucleotides is 0.481. Our results indicate that the mono and dinucleotide composition of nucleic sequences and the amino acids, dipeptides and physicochemical properties of peptide sequences can explain a significant proportion of the variance in their frequencies between organisms in nature.

genomics↗