bioRxiv Science⌕ Search

Biology subjects

Chartoumpekis, D.

Publications and source records attributed to Chartoumpekis, D..

4 recordsLinked to original sources

Unraveling diversity by isolating peptide sequences specific to distinct taxonomic groups

The identification of succinct, universal fingerprints that enable the characterization of individual taxonomies can reveal insights into trait development and can have widespread applications in pathogen diagnostics, human healthcare, ecology and the characterization of biomes. Here, we investigated the existence of peptide k-mer sequences that are exclusively present in a specific taxonomy and absent in every other taxonomic level, termed taxonomic quasi-primes. By analyzing proteomes across 24,073 species, we identified quasi-prime peptides specific to superkingdoms, kingdoms, and phyla, uncovering their taxonomic distributions and functional relevance. These peptides exhibit remarkable sequence uniqueness at six- and seven-amino- acid lengths, offering insights into evolutionary divergence and lineage-specific adaptations. Moreover, we show that human quasi-prime loci are more prone to harboring pathogenic variants, underscoring their functional significance. This study introduces taxonomic quasi-primes and offers insights into their contributions to proteomic diversity, evolutionary pathways, and functional adaptations across the tree of life, while emphasizing their potential impact on human health and disease.

bioinformatics↗

Characterization of hairpin loops and cruciforms across 118,065 genomes spanning the tree of life

Inverted repeats (IRs) can form alternative DNA secondary structures called hairpins and cruciforms, which have a multitude of functional roles and have been associated with genomic instability. However, their prevalence across diverse organismal genomes remains only partially understood. Here, we examine the prevalence of IRs across 118,065 complete organismal genomes. Our comprehensive analysis across taxonomic subdivisions reveals significant differences in the distribution, frequency, and biophysical properties of perfect IRs among these genomes. We identify a total of 29,589,132 perfect IRs and show a highly variable density across different organisms, with strikingly distinct patterns observed in Viruses, Bacteria, Archaea, and Eukaryota. We report IRs with perfect arms of extreme lengths, which can extend to hundreds of thousands of base pairs. Our findings demonstrate a strong correlation between IR density and genome size, revealing that Viruses and Bacteria possess the highest density, whereas Eukaryota and Archaea exhibit the lowest relative to their genome size. Additionally, the study reveals the enrichment of IRs at transcription start and termination end sites in prokaryotes and Viruses and underscores their potential roles in gene regulation and genome organization. Through a comprehensive overview of the distribution and characteristics of IRs in a wide array of organisms, this largest-scale analysis to date sheds light on the functional significance of inverted repeats, their contribution to genomic instability, and their evolutionary impact across the tree of life.

genomics↗

Human Recombinant Thyrotropin Fails to Induce Thyroid Cell Proliferation

Thyrotropin (TSH) suppression is required in the management of patients with papillary thyroid carcinoma (PTC) to improve their outcomes, inevitably causing iatrogenic thyrotoxicosis. Nevertheless, the evidence supporting this practice remains limited and weak, and in vitro studies examining the mitogenic effects of TSH in cancerous cells used supraphysiological doses of bovine TSH, which produced conflicting results. Our study explores for the first time the impact of human recombinant thyrotropin (rh-TSH) on human PTC cell lines (K1 and TPC-1) that were transformed to overexpress the thyrotropin receptor (TSHR). The cells were treated with escalating doses of rh-TSH under various conditions, such as the presence or absence of insulin. The expression levels of TSHR and thyroglobulin (Tg) were determined, and subsequently, the proliferation and migration of both transformed and non-transformed cells were assessed. Under the conditions employed, rh-TSH was not adequate to induce either the proliferation or the migration rate of the cells, while Tg expression was increased. Our experiments indicate that clinically relevant concentrations of rh-TSH cannot induce proliferation and migration in PTC cell lines, even after overexpression of TSHR. Further research is warranted to dissect the underlying molecular mechanisms, and these results could translate into better management of PTC patients.

cancer biology↗

kmerDB: A Database Encompassing the Set of Genomic and Proteomic Sequence Information for Each Species

The rapid decline in sequencing cost has enabled the generation of reference genomes and proteomes for a growing number of organisms. However, at the present time, there is no established repository that provides information about organism-specific genomic and proteomic sequences of certain lengths, also known as kmers, that are either present or absent in each genome or proteome. In this article, we present kmerDB, a database accessible through an interactive web interface that provides kmer based information from genomic and proteomic sequences in a systematic way. kmerDB currently contains 202,340,859,107 base pairs and 19,304,903,356 amino acids, spanning 45,785 and 22,386 reference genomes and proteomes, respectively, as well as 14,658,776 and 149,264,442 genomic and proteomic species-specific sequences, termed quasi-primes. Additionally, we provide access to 5,186,757 nucleic and 214,904,089 peptide sequences that are absent from every genome and proteome, termed primes. kmerDB features a user-friendly interface offering various search options and filters for easy parsing and searching. The service is available at: www.kmerdb.com.

genomics↗