bioRxiv Science⌕ Search

Biology subjects

Kothiwal, D.

Publications and source records attributed to Kothiwal, D..

5 recordsLinked to original sources

An interpretable open platform for sequence-based antibody developability prediction

Antibody developability is increasingly predictable from sequence, yet software and trained models are rarely made available. We present DELPHI, open software for training developability predictors from labelled antibody assay data, together with ready-to-run, retrainable models. DELPHI compares 25 language-model and classifier combinations under CDR H3-cluster cross-validation that reduces sequence-similarity leakage, measures how performance changes with labelled training-set size, and reports residue-level model attributions. Applied to in-house polyreactivity and size-exclusion (SEC) data, it reaches mean AUC 0.959 and 0.933. Trained on those data alone, it transfers to a 246,293-antibody public library (AUC 0.950 with our deployed model) and ranks polyreactivity at a level similar to the best reported Ginkgo competition point estimate, without training on its data. Any laboratory can screen candidates before running assays, generate residue-level engineering hypotheses, and retrain DELPHI for a new assay.

bioinformatics↗

OpenAntigens: a structure-aware database for antigen construct design across the human cell-surface and secreted proteome

Choosing which region of a protein to express remains poorly standardized in antibody discovery, recombinant reagent generation, structural biology and computational binder design. For human cell-surface and secreted proteins, this requires reconciling topology, processing, predicted and experimental structure, modifications, interaction partners, orthologs, paralogs and cross-reactivity risk before ordering DNA. OpenAntigens is a free, no-login database of construct-design reports for 5328 human secreted, GPI-anchored, single-pass and multipass proteins. It integrates UniProt topology, AlphaFold pLDDT and PAE, PDB precedent, InterPro and Pfam domains, mouse and cynomolgus orthologs, paralog and family context, Open Targets disease associations, partner and assembly context, and BLAST searches. It provides 55 305 construct suggestions spanning full design regions, PDB-backed boundaries, annotated domains, pLDDT/PAE-derived regions and membrane-expression options, plus 148 722 sequence-similarity hits to help choose constructs and assess cross-reactivity. For targets with compatible AlphaFold models, the interactive designer links sequence, structure, pLDDT and PAE, allowing users to revise boundaries and export species-equivalent sequences with real-time cysteine and modification warnings. OpenAntigens places reproducible construct suggestions, comparative context and browser editing in one workflow, reducing manual reconciliation across resources. OpenAntigens is available at openantigens.org.

bioinformatics↗

Cohesin sumoylation is required for repression of subtelomeric gene expression in Saccharomyces cerevisiae.

Cohesin is an evolutionary conserved protein complex first described for its role in sister chromatid cohesion, that impacts several chromosomal processes. The functions of budding yeast cohesin in chromosome segregation, replication and repair are regulated by its post-translational processing and modifications. Cohesin associates with centromeres, pericentric regions and discrete sites along chromosome arms till the end. Near chromosome ends, telomeres exist in a heterochromatin-like configuration and exert SIR-complex mediated telomere position effect resulting in repression of sub-telomeric gene transcription. Previously, we reported that cohesin has a SIR-independent role in subtelomeric gene silencing (Kothiwal and Laloraya, 2019). Here, we investigated the requirement of cohesin sumoylation in subtelomeric repression. We created a sumoylation-deficient cohesin complex by fusing the catalytic domain of a SUMO protease, ULP1, to the C-terminus of Mcd1/Scc1 (Mcd1-UD), a cohesin subunit. We show that cohesin sumoylation is required for repression of sub-telomeric genes. In agreement with our earlier observations of a SIR-independent role of cohesin in telomere silencing, SIR-proteins remained bound to a de-repressed subtelomeric gene in MCD1-UD and its expression further increased upon deletion of SIR2. Interestingly, we did not observe a cohesion defect in this mutant suggesting that sister chromatid cohesion and regulation of sub-telomere gene silencing are separable functions of cohesin. Telomere tethering to the nuclear envelope and telomere compaction are defective in MCD1-UD, indicating that sumoylation contributes to cohesins role in subtelomeric chromosome organization. Our data establish the relevance of cohesin sumoylation in subtelomeric repression, a function independent of cohesins role in sister chromatid cohesion.

genetics↗

High-Throughput Machine Learning-Aided Antibody Discovery for Cell Surface Antigens

Machine learning (ML) has the potential to revolutionize antibody design and selection, but its success depends on access to extensive, well-curated datasets of antibody-antigen interactions. To address this need, we developed a synthetic Fab yeast display library optimized for seamless ML integration, focusing on sequence diversity within the CDRH3 loop. The library incorporates key sequence features derived from human B cell repertoires essential for efficient antibody generation captured in a compact antigen recognition module (ARM) format. Built using the VH1-69 heavy chain and four light chains, the library was evaluated against ten human and murine cell surface antigens, including PD-L1, TIGIT, and ROBO1. This approach yielded hundreds of antibodies with robust biophysical properties, validated for functional performance in flow cytometry and immunohistochemistry. Furthermore, ML analysis identified additional antibodies for ROBO2 and PD-L2 from the aggregate sequencing data, demonstrating utility for hybrid in silico and experimental workflows. We provide a publicly accessible dataset comprising more than 68,000 Fab sequences and 486 characterized antibodies. This study establishes an ML-compatible framework designed to accelerate and streamline antibody discovery and development.

biophysics↗

Leveraging HILIC/ERLIC Separations for Online Nanoscale LC- MS/MS Analysis of Phosphopeptide Isoforms from RNA Polymerase II C-terminal Domain

The eukaryotic RNA polymerase II (Pol II) multi-protein complex transcribes mRNA and coordinates several steps of co-transcriptional mRNA processing and chromatin modification. The largest Pol II subunit, Rpb1, has a C-terminal domain (CTD) comprising dozens of repeated heptad sequences (Tyr1-Ser2-Pro3-Thr4-Ser5-Pro6-Ser7), each containing five phospho-accepting amino acids. The CTD heptads are dynamically phosphorylated, creating specific patterns correlated with steps of transcription initiation, elongation, and termination. This CTD phosphorylation code choreographs dynamic recruitment of important co-regulatory proteins during gene transcription. Genetic tools were used to engineer protease cleavage sites across the CTD (msCTD), creating tryptic peptides with unique sequences amenable to mass spectrometry analysis. However, phosphorylation isoforms within each msCTD sequence are difficult to resolve by standard reversed phase chromatography typically used for LC-MS/MS applications. Here, we use a panel of synthetic CTD phosphopeptides to explore the potential of hydrophilic interaction and electrostatic repulsion hydrophilic interaction (HILIC and ERLIC) chromatography as alternatives to reversed phase separation for CTD phosphopeptide analysis. Our results demonstrate that ERLIC provides improved performance for separation of singly- and doubly-phosphorylated CTD peptides for sequence analysis by LC-MS/MS. Analysis of native yeast msCTD confirms that phosphorylation on Ser5 and Ser2 represents the major endogenous phosphoisoforms. We expect this methodology will be especially useful in the investigation of pathways where multiple protein phosphorylation events converge in close proximity.

biochemistry↗