bioRxiv Science⌕ Search

Biology subjects

Williamson, P. C.

Publications and source records attributed to Williamson, P. C..

2 recordsLinked to original sources

Modeling the Sequence Dependence of Differential Antibody Binding in the Immune Response to Infectious Disease

Past studies have shown that incubation of human serum samples on high density peptide arrays followed by measurement of total antibody bound to each peptide sequence allows detection and discrimination of humoral immune responses to a wide variety of infectious disease agents. This is true even though these arrays consist of peptides with near-random amino acid sequences that were not designed to mimic biological antigens. Previously, this immune profiling approach or "immunosignature" has been implemented using a purely statistical evaluation of pattern binding, with no regard for information contained in the amino acid sequences themselves. Here, a neural network is trained on immunoglobulin G binding to 122,926 amino acid sequences selected quasi-randomly to represent a sparse sample of the entire combinatorial binding space in a peptide array using human serum samples from uninfected controls and 5 different infectious disease cohorts infected by either dengue virus, West Nile virus, hepatitis C virus, hepatitis B virus or Trypanosoma cruzi. This results in a sequence-binding relationship for each sample that contains the differential disease information. Processing array data using the neural network effectively aggregates the sequence-binding information, removing sequence-independent noise and improving the accuracy of array-based classification of disease compared to the raw binding data. Because the neural network model is trained on all samples simultaneously, the information common to all samples resides in the hidden layers of the model and the differential information between samples resides in the output layer of the model, one column of a few hundred values per sample. These column vectors themselves can be used to represent each sample for classification or unsupervised clustering applications such as human disease surveillance. Author SummaryPrevious work from Stephen Johnstons lab has shown that it is possible to use high density arrays of near-random peptide sequences as a general, disease agnostic approach to diagnosis by analyzing the pattern of antibody binding in serum to the array. The current approach replaces the purely statistical pattern recognition approach with a machine learning-based approach that substantially enhances the diagnostic power of these peptide array-based antibody profiles by incorporating the sequence information from each peptide with the measured antibody binding, in this case with regard to infectious diseases. This makes the array analysis much more robust to noise and provides a means of condensing the disease differentiating information from the array into a compact form that can be readily used for disease classification or population health monitoring.

immunology↗

Phylogenetic diversity of two common Trypanosoma cruzi lineages in the Southwestern United States

Trypanosoma cruzi is the causative agent of Chagas disease, a devastating parasitic disease endemic to Central and South America, Mexico, and the USA. We characterized the genetic diversity of T. cruzi circulating in five triatomine species (Triatoma gerstaeckeri, T. lecticularia, T. indictiva, T. sanguisuga and T. recurva) collected in Texas and Southern Arizona using nucleotide sequences from four single-copy loci (COII-ND1, MSH2, DHFR-TS, TcCLB.506529.310). All T. cruzi variants fall in two main genetic lineages: 75% of the samples corresponded to T. cruzi Discrete Typing Unit (DTU) I (TcI), and 25% to a North American specific lineage previously labelled TcIV-USA. Phylogenetic and sequence divergence analyses of our new data plus all previously published sequence data from those 4 genes collected in the USA, show that TcIV-USA is significantly different from any other previously defined T. cruzi DTUs. The significant level of genetic divergence between TcIV-USA and other T. cruzi lineages should lead to an increased focus on understanding the epidemiological importance of this lineage, as well as its geographical range and pathogenicity in humans and domestic animals. Our findings further corroborate the fact that there is a high genetic diversity of the parasite in North America and emphasize the need for appropriate surveillance and vector control programs for Chagas disease in southern USA and Mexico.

evolutionary biology↗