bioRxiv Science⌕ Search

Biology subjects

Caputo, V.

Publications and source records attributed to Caputo, V..

6 recordsLinked to original sources

Dynamic consensus pocket detection across molecular dynamics ensembles reveals persistent and transient druggable sites

The traditional "one drug, one target" paradigm assumes that drugs interact with a single specific binding site. Modern pharmacology has proven this definition overly simplistic and, instead, recognizes that drugs operate within complex biological systems and often interact with multiple targets. In this context, proteins cannot be viewed as possessing a single functional binding site, but rather as dynamic entities capable of accommodating ligands at multiple regions, including transient and cryptic pockets. Here, we review and repurpose representative pocket detection tools across geometry-based, energy-based, and machine/deep learning approaches, originally designed to work on static conformations, to evaluate their agreement on molecular dynamics-derived conformational ensembles. Using GLUT1 protein as a dynamic transporter model and Aldose reductase as a cryptic-pocket reference system, we combine inter-tool concordance, HDBSCAN-based spatial clustering, volumetric IoU analysis, and temporal persistence scoring. Our results show that different algorithmic classes capture complementary aspects of pocket dynamics, with energy-based methods showing stronger sensitivity to transient cryptic regions and geometry-based approaches depending more strongly on pre-formed cavities. This work proposes a consensus-oriented framework for identifying conserved and transient druggable pockets in dynamic protein systems.

bioinformatics↗

STRmie-HD enables interruption-aware HTT repeat genotyping and somatic mosaicism profiling across sequencing platforms

Short tandem repeat expansions in exon 1 of the HTT gene drive Huntingtons disease (HD) pathogenesis, with disease onset and progression heavily influenced by somatic mosaicism and sequence interruptions. While sequencing technologies enable repeat sizing, many computational tools lack the resolution to capture subtle interruption motifs and allele-specific somatic variation. We present STRmie-HD, an alignment-free, de novo framework for interruption-aware genotyping and quantitative profiling of somatic mosaicism at single-read resolution. The tool parses individual reads to quantify uninterrupted CAG tract length, CCG repeat content, and critical interruption variants, including Loss of Interruption (LOI) and Duplication of Interruption (DOI). Validated across Illumina, PacBio SMRT, and Oxford Nanopore platforms, STRmie-HD demonstrates high concordance with reference genotypes and superior sensitivity in identifying rare interruption patterns that conventional tools often overlook. Furthermore, it implements somatic mosaicism metrics to characterize repeat dynamics, successfully distinguishing the higher somatic expansion burden in brain tissues compared to peripheral blood. STRmie-HD offers a comprehensive and extensible solution for high-resolution molecular characterization of HTT variation, providing a robust framework for patient stratification and genetic research in HD. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=72 SRC="FIGDIR/small/713334v1_ufig1.gif" ALT="Figure 1"> View larger version (27K): org.highwire.dtl.DTLVardef@17a54aforg.highwire.dtl.DTLVardef@4dcfc5org.highwire.dtl.DTLVardef@8398edorg.highwire.dtl.DTLVardef@1acefde_HPS_FORMAT_FIGEXP M_FIG Graphical Abstract: STRmie-HD flowchart. STRmie-HD is a comprehensive analytical framework that processes sequencing reads to analyze CAG/CCG trinucleotide repeats, interruption variants, and somatic mosaicism in the HTT gene. The workflow begins with sequencing reads (FASTA/FASTQ) that can undergo optional custom processing eq]based on the sequencing design. These reads are then fed into a regular expression-based engine (STRmie-HD) to identify CAG and CCG motifs. The identified motifs lead to the estimation of CAG/CCG alleles, visualized as distinct peaks representing different allele sizes, interruption variant assessment, and somatic mosaicism quantification. STRmie-HD produces an HTML output that wraps this information into a report. C_FIG

bioinformatics↗

nAPOGEE: A machine-learning platform for clinically actionable pathogenicity assessment of all mitochondrial noncoding variants

Mitochondrial noncoding variants, particularly those in tRNA and rRNA genes, pose significant challenges for clinical interpretation due to heteroplasmy, broad phenotypic heterogeneity where symptoms can overlap with other conditions, and the limited availability of well-established genotype-phenotype correlations. Despite their central role in mitochondrial translation, these variants have remained largely unexplored by the existing variant-effect predictors. Here, we present nAPOGEE, a novel machine-learning framework specifically designed to assess the pathogenicity of all possible single-nucleotide variants in human mitochondrial noncoding RNAs. nAPOGEE integrates two specialized predictors: tAPOGEE, which outperforms existing tools for tRNAs, and rAPOGEE, the first dedicated classifier for mitochondrial rRNA variants. Using curated training datasets, phylogenetic conservation metrics, secondary structure modeling, RNA-specific embeddings, and thermodynamic features, nAPOGEE provides biologically interpretable predictions and posterior probabilities aligned with the ACMG/AMP guidelines. Applied to both curated variant sets and population-scale data, nAPOGEE revealed consistent spatial correlation of the predicted pathogenicity, reflecting underlying structural and evolutionary constraints. This study addresses a longstanding gap in mitochondrial genomics and offers a clinically applicable tool for variant prioritization, reclassification, and research into mitochondrial disease mechanisms.

bioinformatics↗

NetMD: Unsupervised Synchronization of Molecular Dynamics Trajectories via Graph Embedding and Time Warping

Molecular dynamics (MD) simulations yield detailed atomistic views of biomolecular processes, yet comparing independent trajectories is hindered by stochastic divergence. Here, we introduce NetMD, a computational approach that synchronizes and analyzes MD trajectories by combining graph-based representations with dynamic time warping. Frames are transformed into residue-contact graphs, entropy-filtered to retain variable interactions, and embedded as low-dimensional vectors. NetMD then uses time-warping barycenter averaging to align these vector trajectories, yielding a consensus "average" trajectory while pruning the outlier simulations. Applied to diverse systems, such as transporters, demethylases, and protein complexes, NetMD revealed shared multiphase dynamics and pinpointed mutation- or ligand-specific deviations. Thus, this method enables an unsupervised, time-resolved comparison of MD ensembles across conditions. It is robust, broadly applicable, and available as an open-source software, offering a powerful tool for uncovering common patterns and critical divergences in biomolecular dynamics.

bioinformatics↗

APOGEE 2: multi-layer machine-learning model for the interpretable prediction of mitochondrial missense variants

APOGEE 2 is a mitochondrially-centered ensemble method designed to improve the accuracy of pathogenicity predictions for interpreting missense mitochondrial variants. Built on the joint consensus recommendations by the American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP), APOGEE 2 features an improved machine learning method and a curated training set for enhanced performance metrics. It offers region-wise assessments of genome fragility and mechanistic analyses of specific amino acids that cause perceptible long-range effects on protein structure. With clinical and research use in mind, APOGEE 2 scores and pathogenicity probabilities are precompiled and available in MitImpact. APOGEE 2s ability to address challenges in interpreting mitochondrial missense variants makes it an essential tool in the field of mitochondrial genetics.

bioinformatics↗

Investigating mitochondrial gene expression patterns in Drosophila melanogaster using network analysis to understand aging mechanisms

The process of aging is a complex phenomenon that involves a progressive decline in physiological functions required for survival and fertility. To better understand the mechanisms underlying this process, the scientific community has utilized several tools. Among them, mitochondrial DNA has emerged as a crucial factor in biological aging, given that mitochondrial dysfunction is thought to significantly contribute to this phenomenon. Additionally, Drosophila melanogaster has proven to be a valuable model organism for studying aging due to its low cost, capacity to generate large populations, and ease of genetic manipulation and tissue dissection. Moreover, graph theory has been employed to understand the dynamic changes in gene expression patterns associated with aging and to investigate the interactions between aging and aging-related diseases. In this study, we have integrated these approaches to examine the patterns of gene co-expression in Drosophila melanogaster at various stages of development. By applying graph-theory techniques, we have identified modules of co-expressing genes, highlighting those that contain a significantly high number of mitochondrial genes. We found important mitochondrial genes involved in aging and age-related diseases in Drosophila melanogaster, including UQCR-C1, ND-B17.2, ND-20, and Pdhb. Our findings shed light on the role of mitochondrial genes in the aging process and demonstrate the utility of Drosophila melanogaster as a model organism and graph theory in aging research.

bioinformatics↗