bioRxiv Science⌕ Search

Biology subjects

Tico, M.

Publications and source records attributed to Tico, M..

3 recordsLinked to original sources

Pyranges v1: a Python framework for ultrafast sequence interval operations

Sequence interval algebra is key to modern bioinformatics. Pyranges v1 offers a Python Pandas-based interface to a comprehensive palette of Rust-powered operations (e.g., overlap, count, slice intervals), enabling the intuitive development of efficient pipelines for diverse sequence data, including gene annotations, mapped reads, and protein domains. Pyranges is faster, consumes less memory, and offers more functionalities than alternative tools including BEDTools, emerging as an innovative one-stop shop for omics analysis.

bioinformatics↗

An updated view of the vertebrate selenoproteome reveals convergent depletions in tetrapods and expansions in ray-finned fishes

Selenoproteins incorporate the rare selenium-containing amino acid selenocysteine (Sec) and play crucial roles for redox homeostasis, stress response, and hormone regulation. Sec is inserted by co-translational recoding of the UGA codon, normally a stop. As a consequence, selenoproteins are often misannotated in public databases and require specialized bioinformatic methods and resources. Here, we present a refined characterization of the composition and evolution of the vertebrate selenoproteome. Based on analyses of 19 gene families across hundreds of genomes, we show that extant selenoproteomes were shaped by extensive gene duplications (56 selenoproteins), losses (50), and Sec-to-cysteine (Cys) conversions (21). Tetrapods including mammals encode 24-25 selenoproteins, with variations in 6 families. Notably, the same genes underwent convergent evolutionary events in multiple tetrapods, namely Sec-to-Cys substitutions (SELENOU1, GPX6) and gene losses (SELENOV). In contrast, ray-finned fish exhibit larger and more dynamic selenoproteomes, reinforcing the hypothesis that the selective advantage of Sec is stronger in aquatic environments. We detected selenoprotein duplications spread across the actinopterygian clade involving 13 families, mainly involved in antioxidant defense. The richest selenoproteomes were found in Salmonidae and Cyprinoidei fish with 56 and 44 selenoproteins, respectively, owing to whole genome duplications. Among our findings, the SELENOP family stands out in lampreys, carrying up to an unprecedented 162 UGAs putatively recoded to Sec. Our study presents the most comprehensive evolutionary map of vertebrate selenoproteins to date and delineates the specific selenoproteome of each lineage, establishing a foundational framework for selenium biology research in the era of biodiversity genomics.

evolutionary biology↗

Overcoming the widespread flaws in the annotation of vertebrate selenoprotein genes in public databases

Selenocysteine (Sec) is a non-canonical amino acid incorporated into selenoproteins, oxidoreductase enzymes carrying essential roles in redox homeostasis. Sec insertion occurs in response to UGA, normally interpreted as stop codon, but recoded in selenoprotein mRNAs. Owing to the dual function of UGA, the identification of selenoprotein genes poses a challenge. We show that the vertebrate selenoprotein genes are widely misannotated in major public databases. Only 12% and 6% of selenoprotein genes are well annotated in Ensembl and NCBI GenBank, respectively, due to the lack of dedicated selenoprotein annotation pipelines. In most cases (81% and 84%), overlapping flawed annotations are present which lack the Sec-encoding UGA. In contrast, NCBI RefSeq employs a dedicated selenoprotein pipeline, yet with some shortcomings: its selenoprotein annotations are correct in 76% of cases, and most errors affect families with a C-terminal Sec residue. We argue that selenoproteins must be correctly annotated in public databases and that must occur via automated pipelines, to keep the pace with genome sequencing. To facilitate this task, we present a new version of Selenoprofiles, an homology based tool for selenoprotein prediction that produces predictions with accuracy comparable to manual curation, and can be easily deployed and integrated in existing annotation pipelines.

bioinformatics↗