bioRxiv Science⌕ Search

Biology subjects

Andrade, H. S.

Publications and source records attributed to Andrade, H. S..

3 recordsLinked to original sources

Genetic diversity of the LILRB1 and LILRB2 coding regions in an admixed Brazilian population sample

Leukocyte Immunoglobulin (Ig)-like Receptors (LILR) LILRB1 and LILRB2 play a pivotal role in maintaining self-tolerance and modulating the immune response through interaction with classical and non-classical Human Leukocyte Antigen (HLA) molecules. Although both diversity and natural selection patterns over HLA genes have been extensively evaluated, little information is available concerning the genetic diversity and selection signatures on the LIRB1/2 regions. Therefore, we identified the LILRB1/2 genetic diversity using next-generation sequencing in a population sample comprising 528 healthy control individuals from Sao Paulo State, Brazil. We identified 58 LILRB1 Single Nucleotide Variants (SNVs), which gave rise to 13 haplotypes with at least 1% of frequency. For LILRB2, we identified 41 SNVs arranged into 11 haplotypes with frequencies above 1%. We found evidence of either positive or purifying selection on LILRB1/2 coding regions. Some residues in both proteins showed to be under the effect of positive selection, suggesting that amino acid replacements in these proteins resulted in beneficial functional changes. Finally, we have shown that allelic variation (six and five amino acid exchanges in LILRB1 and LILRB2, respectively) affects the structure and/or stability of both molecules. Nonetheless, LILRB2 has shown higher average stability, with no D1/D2 residue affecting protein structure. Taken together, our findings demonstrate that LILRB1 and LILRB2 are highly polymorphic and provide strong evidence supporting the directional selection regime hypothesis.

genetics↗

KIR2DL4 genetic diversity in a Brazilian population sample: implications for transcription regulation and protein diversity in samples with different ancestry backgrounds

KIR2DL4 is an important immune modulator expressed in Natural Killer cells, being HLA-G its main ligand. We characterize KIR2DL4 gene diversity considering the promoter, all exons, and all introns, in a highly admixed Brazilian population sample using massively parallel sequencing. We also introduce a molecular method to amplify and sequence the complete KIR2DL4 gene. To avoid mapping bias and genotype errors commonly observed in gene families, we have developed a bioinformatic pipeline designed to minimize mapping, genotyping, and haplotyping errors. We have applied this method to survey the variability of 220 samples from the State of Sao Paulo, southeastern Brazil. We have also compared the KIR2DL4 genetic diversity in Brazilian samples with the previously reported by the 1000Genomes consortium. KIR2DL4 presents high linkage disequilibrium throughout the gene, with coding sequences associated with specific promoters. There were few, but divergent, promoter haplotypes. We have also detected many new KIR2DL4 sequences, all with nucleotide exchanges in introns and encoding previously described proteins. Exons 3 and 4, which encode the external domains, were the most variable ones. The ancestry background influences KIR2DL4 allele frequencies and must be considered for association studies regarding KIR2DL4.

genetics↗

Whole-genome sequencing of 1,171 elderly admixed individuals from the largest Latin American metropolis (Sao Paulo, Brazil)

As whole-genome sequencing (WGS) becomes the gold standard tool for studying population genomics and medical applications, data on diverse non-European and admixed individuals are still scarce. Here, we present a high-coverage WGS dataset of 1,171 highly admixed elderly Brazilians from a census-based cohort, providing over 76 million variants, of which ~2 million are absent from large public databases. WGS enabled identifying ~2,000 novel mobile element insertions, nearly 5Mb of genomic segments absent from human genome reference, and over 140 novel alleles from HLA genes. We reclassified and curated nearly four hundred variant's pathogenicity assertions in genes associated with dominantly inherited Mendelian disorders and calculated the incidence for selected recessive disorders, demonstrating the clinical usefulness of the present study. Finally, we observed that whole-genome and HLA imputation could be significantly improved compared to available datasets since rare variation represents the largest proportion of input from WGS. These results demonstrate that even smaller sample sizes of underrepresented populations bring relevant data for genomic studies, especially when exploring analyses allowed only by WGS.

genomics↗