bioRxiv Science⌕ Search

Biology subjects

Bielska, W.

Publications and source records attributed to Bielska, W..

2 recordsLinked to original sources

ASD: Antigen-Specific Antibody Database

The development of computational models addressing therapeutic antibodies faces significant challenges due to the scarcity of data. A critical data element is the set of antibody-antigen interaction pairs associated with sequences. To address this issue, we developed the Antigen Specific Antibody Database (ASD, https://naturalantibody.com/agab/), a database aggregating antibody-antigen interaction data from multiple studies with standardized formatting and annotations. Our dataset compilation strategy resulted in data from 15 distinct sources, resulting in 1,097,946 unique antibody-antigen interactions (with 9,575 unique antigens). The ASD database captures diverse affinity measures and qualitative binding assessment, along with metadata including UniProt and PDB identifiers, target protein names, confidence levels, and experimental conditions. Through this integration drive, we make available an ample resource of interaction data gathered from the public domain to act as a foundation for model development and further data generation.

bioinformatics↗

Conserved heavy/light contacts and germline preferences revealed by a large-scale analysis of natively paired human antibody sequences and structural data.

Antibody next-generation sequencing (NGS) datasets have become crucial to develop computational models addressing this successful class of therapeutics. Although antibodies are composed of both heavy and light chains, most NGS sequencing depositions provide them in unpaired form, reducing their utility. Here we introduce PairedAbNGS, a novel database with paired heavy/light antibody chains. To the best of our knowledge, this is the largest resource for paired natural antibody sequences with 58 bioprojects and over 14 million assembled productive sequences. We make the database accessible at http://naturalantibody.com/paired-ngs as a valuable tool for biological and machine-learning applications. Using this dataset, we investigated heavy and light chain variable (V) gene pairing preferences and found significant biases beyond gene usage frequencies, possibly due to receptor editing favoring less autoreactive combinations. Analyzing the available antibody structures from the Protein Data Bank, we studied conserved contact residues between heavy and light chains, particularly interactions between the CDR3 region of one chain and the FWR2 region of the opposite chain. Examination of amino acid pairs at key contact sites revealed significant deviations of amino acids distributions compared to random pairings, in the heavy chains CDR3 region contacting the opposite chain, indicating specific interactions might be crucial for proper chain pairing. This observation is further reinforced by preferential IGHV-IGLJ and IGLV-IGHJ pairing preferences. We hope that both our resources and the findings would contribute to improving the engineering of biological drugs.

immunology↗