bioRxiv Science⌕ Search

Biology subjects

Canzani, D.

Publications and source records attributed to Canzani, D..

4 recordsLinked to original sources

Structure-free, site-resolved contrastive learningextends small-molecule discovery beyond the reachof structure-based modeling

Virtual screening asks which molecules, among an enormous space of drug-like chemistry, are worth synthesizing and testing against a protein target. Most modern methods answer by building and scoring an explicit three-dimensional pose through molecular docking, or the co-folding models that now approach experimental accuracy. Building these poses presumes a well-defined pocket, and the non-orthosteric, cryptic, and intrinsically disordered sites where unexplored ligandability lies offer none. There, these methods fail to generalize. Here we present Ptarmigan-1, a contrastive model that co-embeds the residues of a protein with candidate small molecules in a shared latent space, from sequence and two-dimensional chemistry alone, and without ever constructing a pose. Engagement reduces to the proximity of precomputed embeddings. Freed from the pose, Ptarmigan-1 trains directly on chemoproteomic and bioactivity data of mixed resolution, scores a compound in ten milliseconds rather than the tens of seconds a co-folding model demands, and resolves each prediction to the residues a compound engages. On well-folded, orthosteric targets it performs comparably to a collection of co-folding and docking models, and on covalent, cryptic, and disordered sites it matches or exceeds them. It localizes reversible and covalent inhibitors to the pockets they engage, even for targets withheld from training, and screens the entire human proteome against a library of 3.4 billion compounds in under a day. By decoupling molecular recognition from structure, Ptarmigan-1 recasts virtual screening as a reusable index that continuously improves as data accumulate.

molecular biology↗

Carafe2 enables high quality in silico spectral library generation for timsTOF data-independent acquisition proteomics

Data-independent acquisition (DIA) proteomics enables reproducible and systematic peptide detection and quantification, and trapped ion mobility spectrometry (TIMS) on the timsTOF platform further improves DIA by synchronizing ion mobility separation with quadrupole precursor sampling. Analyzing the highly multiplexed spectra generated by DIA typically relies on spectral libraries, and fully leveraging the additional ion mobility dimension requires these libraries to include accurate retention time, fragment ion intensity, and ion mobility annotations. Existing in silico spectral library generation tools either lack ion mobility support entirely or rely on models trained on data-dependent acquisition (DDA) data, that can introduce a mismatch that may not capture unique experiment-specific biases when applied to each respective timsTOF dataset. Carafe is a software tool that uses deep learning models to generate high-quality, experiment-specific in silico libraries by training directly on DIA data. In this study, we extend Carafe to generate libraries for timsTOF DIA data, which involves fine-tuning retention time (RT), fragment ion intensity, and ion mobility prediction models using timsTOF DIA data. Carafe2 operates directly on native timsTOF raw data (Bruker .d directories) without the need for data conversion. We demonstrate the performance of Carafe2 across a wide range of DIA applications, including global proteome, phosphoproteome, and plasma proteome datasets. Comparing Carafe2 fine-tuned RT, fragment ion intensity, and ion mobility prediction models with pretrained DDA models, we find that Carafe2 models outperform pretrained models on a variety of DIA datasets. We then demonstrate the utility of in silico libraries generated by Carafe2 for peptide detection on several different types of timsTOF DIA datasets by comparing with the libraries generated with DDA-trained AlphaPeptDeep models, DIA-NN built-in models, and empirical spectral libraries generated from DDA experiments.

bioinformatics↗

Native, Spatiotemporal Profiling of the Global Human Regulome

The regulome, comprising transcription factors, cofactors, chromatin remodelers, and other regulatory proteins, forms the core machinery by which cells interpret signals and execute gene expression programs. Despite its central role in development, disease, and drug response, the regulome remains largely uncharted at scale due to its dynamic, low-abundance, and chromatin-associated nature. Here, we present a method for scalable, regulome profiling for global, compartment-resolved quantification of native regulome proteins. By enriching DNA- and chromatin-associated proteins and profiling them using high-throughput, label-free DIA mass spectrometry, regulome profiling captures chromatin-associated proteins across 36 human cell lines and thousands of perturbations. The resulting Regulome Atlas recovers nearly 60% of known human transcription factors, reveals lineage-specific TF localization, and distinguishes active nuclear engagement from latent, unbound states. We demonstrate that regulome profiles resolve acute immune pathway activation prior to transcriptional changes, identify previously unrecognized drug-induced regulome responses, and enable proteome-scale readouts of compound target engagement and complex remodeling. This work establishes a foundational resource for decoding the regulatory proteome and provides a blueprint for integrating regulome data into next-generation models of cellular behavior.

cell biology↗

An artificial intelligence accelerated virtual screening platform for drug discovery

Structure-based virtual screening is a key tool in early drug discovery, with growing interest in the screening of multi-billion chemical compound libraries. However, the success of virtual screening crucially depends on the accuracy of the binding pose and binding affinity predicted by computational docking. Here we developed a highly accurate structure-based virtual screen method, RosettaVS, for predicting docking poses and binding affinities. Our approach outperforms other state-of-the-art methods on a wide range of benchmarks, partially due to our ability to model receptor flexibility. We incorporate this into a new open-source artificial intelligence accelerated virtual screening platform for drug discovery. Using this platform, we screened multi-billion compound libraries against two unrelated targets, a novel ubiquitin ligase target KLHDC2 and the human voltage-gated sodium channel NaV1.7. On both targets, we discover hits, including seven novel hits (14% hit rate) to KLHDC2 and four novel hits (44% hit rate) to NaV1.7 with single digit micromolar binding affinities. Screening in both cases was completed in less than seven days. Finally, a high resolution X-ray crystallographic structure validates the predicted docking pose for the KLHDC2 ligand complex, demonstrating the effectiveness of our method in lead discovery.

biochemistry↗