bioRxiv Science⌕ Search

Biology subjects

Karakostis, K.

Publications and source records attributed to Karakostis, K..

2 recordsLinked to original sources

Tracing the expansion of p53 retrogenes in elephant species: A foundation for functional insights.

Elephants have evolved multiple TP53 copies through a retrotransposition event followed by successive duplications. Some of these TP53 retrogenes (RTGs) are expressed and hypothesized to have functional roles in cellular regulation. However, comparative genomic studies on TP53 evolution and function are limited due to scarce genomic data for elephants and other afrotherians. Most existing research relies on scaffold assemblies of Loxodonta africana (LoxAfr3 and LoxAfr4), with some focus on Elephas maximus chromosomal assembly. In this in silico study, we analyzed three elephant genomes to validate TP53 RTGs, assess their copy variation, and trace their evolution. For the first time we describe 29 TP53 RTGs in E. maximus versus 18-19 in L. africana. These copies show sequence variation, especially in the duplicated regions and their flanking repetitive elements. Chromosomal mapping in E. maximus revealed that two major classes of TP53 RTGs are consistently arranged in pairs on chromosome 27, which harbours 27 of the 29 identified copies. The observed distribution strongly supports an evolutionary model in which large-scale genomic segments, each encompassing at least two retrogenes of different groups, were duplicated early in the elephant lineage, driving the extensive amplification of TP53 RTGs, as suggested also by the patterns of the flanking repetitive elements. The TP53 RTGs expansion was followed by a unique inversion on chromosome 27 that separates the duplication clusters. Thus, this study enhances our understanding of the elephants multi-p53 system, linked to cancer resistance, body size, and Petos paradox, and supports ongoing research into functional aspects. Significance StatementThis study uncovers the complex genomic architecture of elephant TP53 RTGs, which have diversified into two phylogenetically distinct types present in both African and Asian elephants. While both species share these 2 types, they differ substantially in copy number (18 in African versus 28 in Asian elephants), and in sequence variation, highlighting lineage-specific evolutionary trajectories. Detailed sequence analyses of the chromosomal organization of these copies in the Elephas maximus high-quality genome assembly, particularly the cluster on chromosome 27, indicates that they arose from stepwise duplications of extended genomic segments, which likely facilitated both the expansion and regulatory diversification of the TP53 RTGs repertoire. Comparative analysis with other mammals reveals an elephant-specific inversion that reorganized the expanded TP53 RTGs copies, creating a new genomic configuration that could have influenced the regulation and expression of some of the p53 RTGs. Therefore, this work advances our understanding of how evolutionary pressures shaped the landscape of TP53 RTGs and their flanking genetics elements in elephants. By identifying accurately the retrogene copy numbers, sequences and putative functional domains, it establishes a critical foundation for future studies investigating their functional roles in DNA damage and genome stability.

evolutionary biology↗

Resolving the full set of human polymorphic inversions and other complex variants from ultra-long read data

Inversions are a unique type of balanced structural variants (SVs) with important consequences in multiple organisms. However, despite considerable effort, this and other complex SVs remain poorly characterized due to the presence of large repeats. New techniques are finally allowing us to identify the full spectrum of human inversions, but the number of individuals analyzed is still quite limited. Here, we take advantage of Oxford Nanopore Technologies (ONT) long reads to characterize an exhaustive catalogue of 612 candidate inversions between 197 bp and 4.4 Mb of length and flanked by <190-kb long inverted repeats (IRs). For that, we developed a bioinformatic package to identify inversion alleles reliably from long read data. Next, using a combination of different DNA extraction, library preparation, and ONT sequencing protocols, we showed that ultra-long reads (50-100 kb) and adaptive sampling are an efficient method to detect most human inversions. Lastly, by analyzing ONT data from 54 diverse individuals, 87-99% of the inversions could be genotyped in each sample, depending mainly on read and IR length and genome coverage. Both orientations were observed for 155 of the analyzed regions (frequency 0.01-0.49), which multiplies by three the polymorphic IR-mediated inversions studied in detail so far. Moreover, we found more than 300 additional independent SVs in the studied regions and resolved several complex rearrangements. Our work therefore provides an accurate benchmark of those inversions that typically escape most analyses, improving existing resources, such as the Pangenome. In addition, it demonstrates the potential of nanopore sequencing to determine the functional impact of missing human genomic variation.

genomics↗