bioRxiv ScienceSearch

Biology subjects

Santorico, S. A.

Publications and source records attributed to Santorico, S. A..

2 recordsLinked to original sources

Genome-wide Copy Number Variations in a Large Cohort of Bantu African Children

BackgroundCopy number variations (CNVs) account for a substantial proportion of inter-individual genomic variation. However, a majority of genomic variation studies have focused on single-nucleotide variations (SNVs), with limited genome-wide analysis of CNVs in large cohorts, especially in populations that are under-represented in genetic studies including people of African descent. ResultsIn this study, we carried out a genome-wide analysis in > 3400 healthy Bantu Africans from Tanzania using high density (> 2.5 million probes) genotyping arrays. We identified over 400000 CNVs larger than 1 kilobase (kb), for an average of 120 CNVs (SE = 2.57) per individual. We detected 866 large CNVs ([≥] 300 kb), some of which overlapped genomic regions previously associated with multiple congenital anomaly syndromes, including Prader-Willi/Angelman syndrome (Type1) and 22q11.2 deletion syndrome. Furthermore, several of the common CNVs seen in our cohort ([≥] 5%) overlap genes previously associated with developmental disorders. ConclusionThese findings may help refine the phenotypic outcomes and penetrance of variations affecting genes and genomic regions previously implicated in diseases. Our study provides one of the largest datasets of CNVs from individuals of African ancestry, enabling improved clinical evaluation and disease association of CNVs observed in research and clinical studies in African populations.

genomics

Incorporation of Heterogeneity in a Case-Control Study Through a Mixture Model

Most common human diseases and complex traits are etiologically heterogeneous. Genome-wide Association Studies (GWAS) aim to discover common genetic variants that are associated with complex traits, typically without considering heterogeneity. Heterogeneity, as well as im-precise phenotyping, significantly reduces the power to find genetic variants associated with human diseases and complex traits. Disease subtyping through unsupervised clustering techniques such as latent class analysis can explain some of the heterogeneity; however, subtyping methods do not typically incorporate heterogeneity into the association framework. Here, we use a finite mixture model with logistic regression to incorporate heterogeneity into the association testing framework for a case-control study. In the proposed method, the disease outcome is modeled as a mixture of two binomial distributions. One of the component distributions refers to the subgroup of the population for which the genetic variant is not associated with the disease outcome and another component distribution corresponds to the subgroup for which the genetic variant is associated with the disease outcome. The mixing parameter corresponds to the proportion of the population for which the genetic variant is associated with the disease outcome. A simulation study of a trait with differing levels of prevalence, SNP minor allele frequency, and odds ratio was performed, and effect size estimates compared between the models with and without incorporating heterogeneity. The proposed mixture model yields lower bias of odds ratios while having comparable power compared to classical logistic regression.

genetics