bioRxiv Science⌕ Search

Biology subjects

Wade, K. J.

Publications and source records attributed to Wade, K. J..

2 recordsLinked to original sources

High-throughput complement component 4 genomic sequence analysis with C4Investigator

The complement component 4 gene locus, composed of the C4A and C4B genes and located on chromosome 6, encodes for C4 protein, a key intermediate in the classical and lectin pathways of the complement system. The complement system is an important modulator of immune system activity and is also involved in the clearance of immune complexes and cellular debris. The C4 gene locus exhibits copy number variation, with each composite gene varying between 0-5 copies per haplotype, C4 genes also vary in size depending on the presence of the HERV retrovirus in intron 9, denoted by C4(L) for long-form and C4(S) for short-form, which modulates expression and is found in both C4A and C4B. Additionally, human blood group antigens Rodgers and Chido are located on the C4 protein, with the Rodger epitope generally found on C4A protein, and the Chido epitope generally found on C4B protein. C4 copy number variation has been implicated in numerous autoimmune and pathogenic diseases. Despite the central role of C4 in immune function and regulation, high-throughput genomic sequence analysis of C4 variants has been impeded by the high degree of sequence similarity and complex genetic variation exhibited by these genes. To investigate C4 variation using genomic sequencing data, we have developed a novel bioinformatic pipeline for comprehensive, high-throughput characterization of human C4 sequence from short-read sequencing data, named C4Investigator. Using paired-end targeted or whole genome sequence data as input, C4Investigator determines gene copy number for overall C4, C4A, C4B, C4(Rodger), C4(Ch), C4(L), and C4(S), additionally, C4Ivestigator reports the full overall C4 aligned sequence, enabling nucleotide level analysis of C4. To demonstrate the utility of this workflow we have analyzed C4 variation in the 1000 Genomes Project Dataset, showing that the C4 genes are highly poly-allelic with many variants that have the potential to impact C4 protein function.

bioinformatics↗

A fast, general synteny detection engine

The increasingly widespread availability of genomic data has created a growing need for fast, sensitive and scalable comparative analysis methods. A key aspect of comparative genomic analysis is the study of synteny, co-localized gene clusters shared among genomes due to descent from common ancestors. Synteny can provide unique insight into the origin, function, and evolution of genome architectures, but methods to identify syntenic patterns in genomic datasets are often inflexible and slow, and use diverse definitions of what counts as likely synteny. Moreover, the reliable identification of putatively syntenic regions (i.e., whether they are truly indicative of homology) with different lengths and signal to noise ratios can be difficult to quantify. Here, we present Mology, a fast, flexible, alignment-free, nonparametric method to detect regions of syntenic elements among genomes or other datasets. The core algorithm operates on consecutive, rank-ordered elements, which could be genes, operons, motifs, sequence fragments, or any other orderable element. It is agnostic to the physical distance between distinct elements and also to directionality and order within syntenic regions, although such considerations can be addressed post hoc. We describe the underlying statistical theory behind our analysis method, and employ a Monte Carlo approach to estimate the false positive rate and positive predictive values for putative syntenic regions. We also evaluate how varying amounts of noise affect recovery of true syntenic regions among Saccharomycetaceae yeast genomes with up to ~100 million years of divergence. We discuss different strategies for recursive application of our method on syntenic regions with sparser signal than considered here, as well as the general applicability of the core algorithm.

evolutionary biology↗