bioRxiv Science⌕ Search

Biology subjects

Hummer, A. M.

Publications and source records attributed to Hummer, A. M..

3 recordsLinked to original sources

Assessment of nucleic acid structure prediction in CASP16

Consistently accurate 3D nucleic acid structure prediction would facilitate studies of the diverse RNA and DNA molecules underlying life. In CASP16, blind predictions for 42 targets canvassing a full array of nucleic acid functions, from dopamine binding by DNA to formation of elaborate RNA nanocages, were submitted by 65 groups from 46 different labs worldwide. In contrast to concurrent protein structure predictions, performance on nucleic acids was generally poor, with no predictions of previously unseen natural RNA structures achieving TM-scores above 0.8. Even though automated server performance has improved, all top-performing groups were human expert predictors: Vfold, GuangzhouRNA-human, and KiharaLab. Good performance on one template-free modeling target (OLE RNA) and accurate global secondary structure prediction suggested that structural information can be extracted from multiple sequence alignments. However, 3D accuracy generally appeared to depend on the availability of closely related 3D structure templates, and predictions still did not achieve consistent recovery of pseudoknots, singlet Watson-Crick-Franklin pairs, non-canonical pairs, or tertiary motifs like A-minor interactions. For the first time, blind predictions of nucleic acid interactions with small molecules, proteins, and other nucleic acids could be assessed in CASP16. As with nucleic acid monomers, prediction accuracy for nucleic acid complexes was generally poor unless 3D templates were available. Accounting for template availability, there has not been a notable increase in nucleic acid modeling accuracy between previous blind challenges and CASP16.

biophysics↗

Baselining the Buzz. Trastuzumab-HER2 Affinity, and Beyond!

Strong antibody-antigen binding is the primary consideration when developing an efficacious therapeutic antibody. In recent years, much work has been devoted to applying complex machine learning models to this cause, yet simple baselines are often lacking. Here, we show that the widely used sequence alignment method, BLOSUM, can yield diverse, binder-enriched libraries from a single starting antibody. Using Trastuzumab-HER2 as a model system, we experimentally validated 720 novel designs generated with five different computational methods using surface plasmon resonance. The BLOSUM substitution matrix outperformed all four deep learning design approaches tested, achieving an estimated minimum binder enrichment of 12.5% and producing nine sub-nanomolar binders. These results underscore the importance of comparing against simple baselines and set a benchmark to guide future computational antibody library design. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/586756v2_ufig1.gif" ALT="Figure 1"> View larger version (32K): org.highwire.dtl.DTLVardef@1597ee9org.highwire.dtl.DTLVardef@9af4b6org.highwire.dtl.DTLVardef@1380e61org.highwire.dtl.DTLVardef@1380d29_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Investigating the Volume and Diversity of Data Needed for Generalizable Antibody-Antigen ΔΔG Prediction

Antibody-antigen binding affinity lies at the heart of therapeutic antibody development: efficacy is guided by specific binding and control of affinity. Here we present Graphinity, an equivariant graph neural network architecture built directly from antibody-antigen structures that achieves state-of-the-art performance on experimental {triangleup}{triangleup}G prediction. However, our model, like previous methods, appears to be overtraining on the few hundred experimental data points available. To test if we could overcome this problem, we built a synthetic dataset of nearly 1 million FoldX-generated {triangleup}{triangleup}G values. Graphinity achieved Pearsons correlations nearing 0.9 and was robust to train-test cutoffs and noise on this dataset. The synthetic dataset also allowed us to investigate the role of dataset size and diversity in model performance. Our results indicate there is currently insufficient experimental data to accurately and robustly predict {triangleup}{triangleup}G, with orders of magnitude more likely needed. Dataset size is not the only consideration - our tests demonstrate the importance of diversity. We also confirm that Graphinity can be used for experimental binding prediction by applying it to a dataset of >36,000 Trastuzumab variants.

bioinformatics↗