bioRxiv Science⌕ Search

Biology subjects

Vilicich, F.

Publications and source records attributed to Vilicich, F..

3 recordsLinked to original sources

Structure-based Predictions of Conformational B Cell Epitopes by Protein Language Model and Deep Learning

Mapping conformational B-cell epitopes remains a central challenge for antibody discovery: experiments are costly and most computational tools trained on generic protein-protein interfaces transfer poorly to antibody-antigen recognition. We introduce a patch-centric framework that predicts epitopes directly on antigen structures. Each surface "patch" is defined as a triad of neighboring residues, capturing the smallest local unit that encodes both shape and chemistry. We evaluate two classifiers: (i) a protein language model (PLM) approach that averages ESM-2 embeddings over each triad and scores them with a small multilayer perceptron [1], and (ii) a convolutional baseline that consumes a hand-crafted 15x20 feature matrix summarizing amino-acid identity, secondary structure, solvent accessibility, and shape index. Trained with five-fold cross-validation on 1,151 AbDb antibody-antigen complexes, the PLM model markedly outperforms the CNN at the patch level (e.g., F1{approx} 0.986, ROC-AUC{approx} 0.998). Aggregating patch scores to residues with an ensemble over all folds yields robust residue-wise performance, surpassing the CNN (ROC-AUC 0.689{+/-}0.072 vs. 0.548{+/-}0.018). Against widely used sequence- and structure-based tools on AbDb, our PLM achieves the best summary metrics (ROC-AUC 0.67, PR- AUC 0.56) with full coverage of all antigens. On five external complexes unseen during development, the model generalizes well (ROC-AUC 0.663) and accurately localizes binding regions qualitatively. The method converts PLM representations into interpretable epitope likelihood maps, offering a practical aid for antigen prioritization, antibody engineering, and vaccine design.

bioinformatics↗

Inferring Dynamic Information from Protein Structures by Gaussian Integrals and Deep Learning

Protein conformational flexibility underlies a wide range of biological functions, yet experimentally probing dynamics at atomic resolution remains costly and low-throughput. Here, we present a deep learning framework that predicts protein flexibility directly from static structural descriptors, bypassing the need for molecular dynamics (MD) simulations. Using the ATLAS database of standardized all-atom MD trajectories, we encoded 1,374 protein chains as 30-dimensional Gaussian integral (GI) vectors--global shape and topology invariants of the protein backbone. Principal component analysis of GI profiles revealed four structural clusters with distinct secondary structure compositions and flexibility distributions. We trained an attention-based one-dimensional convolutional neural network (1D-CNN) to classify proteins as flexible or non-flexible based on their root-mean-square fluctuation (RMSF) relative to the dataset-wide mean. The classifier achieved an AUC of 0.772 (95% CI: 0.712-0.826) on an independent test set, with balanced sensitivity and specificity, and identified a small subset of GI components as the most predictive. In a regression setting, a recurrent neural network outperformed other architectures, attaining an R2 of 0.537, though high-flexibility values were systematically underestimated. Cluster-specific analyses indicated that coil-rich and {beta}-sheet-dominated proteins were more amenable to flexibility prediction than -helical proteins, likely due to greater structural heterogeneity. Our results demonstrate that compact GI descriptors preserve sufficient information to recover MD-derived flexibility trends, offering a computationally efficient complement to simulation-based approaches. This framework enables large-scale screening of protein dynamics from structural data alone, with potential applications in structural bioinformatics, drug design, and functional annotation.

bioinformatics↗

Multi-omics identification of extracellular components of the fetal monkey and human neocortex.

During development, precursor cells are continuously and intimately interacting with their extracellular environment, which guides their ability to generate functional tissues and organs. Much is known about the development of the neocortex in mammals. This information has largely been derived from histological analyses, heterochronic cell transplants, and genetic manipulations in mice, and to a lesser extent from transcriptomic and histological analyses in humans. However, these approaches have not led to a characterization of the extracellular composition of the developing neocortex in any species. Here, using a combination of single-cell transcriptomic analyses from published datasets, and our proteomics and immunohistofluorescence analyses, we provide a more comprehensive and unbiased picture of the early developing fetal neocortex in humans and non-human primates. Our findings provide a starting point for further hypothesis-driven studies on structural and signaling components in the developing cortex that had previously not been identified.

neuroscience↗