bioRxiv ScienceSearch

Biology subjects

Bustamante, C. D.

Publications and source records attributed to Bustamante, C. D..

16 recordsLinked to original sources

Caring without sharing: Meta-analysis 2.0 for massive genome-wide association studies

Genome-wide association studies have been effective at revealing the genetic architecture of simple traits. Extending this approach to more complex phenotypes has necessitated a massive increase in cohort size. To achieve sufficient power, participants are recruited across multiple collaborating institutions, leaving researchers with two choices: either collect all the raw data at a single institution or rely on meta-analyses to test for association. In this work, we present a third alternative. Here, we implement an entire GWAS workflow (quality control, population structure control, and association) in a fully decentralized setting. Our iterative approach (a) does not rely on consolidating the raw data at a single coordination center, and (b) does not hinge upon large sample size assumptions at each silo. As we show, our approach overcomes challenges faced by meta-studies when it comes to associating rare alleles and when case/control proportions are wildly imbalanced at each silo. We demonstrate the feasibility of our method in cohorts ranging in size from 2K (small) to 500K (large), and recruited across 2 to 10 collaborating institutions.

bioinformatics

Deep learning facilitates rapid cohort identification using human and veterinary clinical narratives

ObjectiveCurrently, dedicated tagging staff spend considerable effort assigning clinical codes to patient summaries for public health purposes, and machine-learning automated tagging is bottlenecked by availability of electronic medical records. Veterinary medical records, a largely untapped data source that could benefit both human and non-human patients, could fill the gap. Materials and MethodsIn this retrospective study, we trained long short-term memory (LSTM) recurrent neural networks (RNNs) on 52,722 human and 89,591 veterinary records. We established relevant baselines by training Decision Trees (DT) and Random Forests (RF) on the same data. We finally investigated the effect of merging data across clinical settings and probed model portability. ResultsWe show that the LSTM-RNNs accurately classify veterinary/human text narratives into top-level categories with an average weighted macro F1, score of 0.735/0.675 respectively. The evaluation metric for the LSTM was 7 and 8% higher than that of the DT and RF models respectively. We generally did not find evidence of model portability albeit moderate performance increases in select categories. DiscussionWe see a strong positive correlation between number of training samples and classification performance, which is promising for future efforts. The use of LSTM-RNN models represents a scalable structure that could prove useful in cohort selection, which could in turn better address emerging public health concerns. ConclusionDigitization of human and veterinary health information will continue to be a reality. Our approach is a step forward for these two domains to learn from, and inform, one another.

epidemiology

Standardized biogeographic grouping system for annotating populations in pharmacogenetic research

The varying frequencies of pharmacogenetic alleles between populations have important implications for the impact of these alleles in different populations. Current population grouping methods to communicate these patterns are insufficient as they are inconsistent and fail to reflect the global distribution of genetic variability. To facilitate and standardize the reporting of variability in pharmacogenetic allele frequencies, we present seven geographically-defined groups: American, Central/South Asian, East Asian, European, Near Eastern, Oceanian, and Sub-Saharan African, and two admixed groups: African American/Afro-Caribbean and Latino. These nine groups are defined by global autosomal genetic structure and based on data from large-scale sequencing initiatives. We recognize that broadly grouping global populations is an oversimplification of human diversity and does not capture complex social and cultural identity. However, these groups meet a key need in pharmacogenetics research by enabling consistent communication of the scale of variability in global allele frequencies and are now used by PharmGKB.

genetics

The Clinical Imperative for Inclusivity: Race, Ethnicity, and Ancestry (REA) in Genomics

The Clinical Genome Resource (ClinGen) Ancestry and Diversity Working Group highlights the need to develop guidance on race, ethnicity, and ancestry (REA) data collection and use in clinical genomics. We present quantitative and qualitative evidence to characterize: 1) acquisition of REA data via clinical laboratory requisition forms, and 2) information disparity across populations in the Genome Aggregation Database (gnomAD) at clinically relevant sites as determined by variants in ClinVar. Our requisition form analysis showed substantial heterogeneity in clinical laboratory ascertainment of REA, as well as marked incongruity among terms used to define REA categories. There was also striking disparity across REA populations in the amount of information available about variants at clinically relevant sites in gnomAD. European ancestral populations constituted the majority of observations (55.8%), allele counts (59.7%), and private alleles (56.1%) in gnomAD at 550 loci with \"pathogenic\" and \"likely pathogenic\" expert-reviewed variants in ClinVar. Our findings highlight the importance of implementing and supporting programs to increase diversity in genome sequencing and clinical genomics, as well as measuring uncertainty around population-level datasets that are used in variant interpretation. Finally, we suggest the need for a standardized REA data collection framework to be developed and adopted across clinical genomics.

genomics

Structural variation detection by proximity ligation from FFPE tumor tissue

The clinical management and therapy of many solid tumor malignancies is dependent on detection of medically actionable or diagnostically relevant genetic variation. However, a principal challenge for genetic assays from tumors is the fragmented and chemically damaged state of DNA in formalin-fixed paraffin-embedded (FFPE) samples. From highly fragmented DNA and RNA there is no current technology for generating long-range DNA sequence data as is required to detect genomic structural variation or long-range genotype phasing. We have developed a high-throughput chromosome conformation capture approach for FFPE samples that we call \"Fix-C\", which is similar in concept to Hi-C. Fix-C enables structural variation detection from fresh and archival FFPE samples. We applied this method to 15 clinical adenocarcinoma and sarcoma specimens spanning a broad range of tumor purities. In this panel, Fix-C analysis achieves a 90% concordance rate with FISH assays - the current clinical gold standard. Additionally, we are able to identify novel structural variation undetected by other methods and recover long-range chromatin configuration information from these FFPE samples harboring highly degraded DNA. This powerful approach will enable detailed resolution of global genome rearrangement events during cancer progression from FFPE material, and inform the development of targeted molecular diagnostic assays for patient care.

genomics

Genetic architecture drives seasonal onset of hibernation in the 13-lined ground squirrel

Hibernation is a highly dynamic phenotype whose timing, for many mammals, is controlled by a circannual clock and accompanied by rhythms in body mass and food intake. When housed in an animal facility, 13-lined ground squirrels exhibit individual variation in the seasonal onset of hibernation, which is not explained by environmental or biological factors, such as body mass and sex. We hypothesized that underlying genetic architecture instead drives variation in this timing. After first increasing the contiguity of the genome assembly, we therefore employed a genotype-by-sequencing approach to characterize genetic variation in 153 13-lined ground squirrels. Combining this with datalogger records, we estimated high heritability (61-100%) for the seasonal onset of hibernation. After applying a genome-wide scan with 46,996 variants, we also identified 21 loci significantly associated with hibernation immergence, which alone accounted for 54% of the variance in the phenotype. The most significant marker (SNP 15, p=3.81x10-6) was located near prolactin-releasing hormone receptor (PRLHR), a gene that regulates food intake and energy homeostasis. Other significant loci were located near genes functionally related to hibernation physiology, including muscarinic acetylcholine receptor M2 (CHRM2), involved in the control of heart rate, exocyst complex component 4 (EXOC4) and prohormone convertase 2 (PCSK2), both of which are involved in insulin signaling and processing. Finally, we applied an expression quantitative loci (eQTL) analysis using existing transcriptome datasets, and we identified significant (q<0.1) associations for 9/21 variants. Our results highlight the power of applying a genetic mapping strategy to hibernation and present new insight into the genetics driving its seasonal onset.

genetics

In-solution Y-chromosome capture-enrichment on ancient DNA libraries

BackgroundAs most ancient biological samples have low levels of endogenous DNA, it is advantageous to enrich for specific genomic regions prior to sequencing. One approach - in-solution capture-enrichment - retrieves sequences of interest and reduces the fraction of microbial DNA. In this work, we implement a capture-enrichment approach targeting informative regions of the Y chromosome in six human archaeological remains excavated in the Caribbean and dated between 200 and 3,000 years BP. We compare the recovery rate of Y-chromosome capture (YCC) alone, whole-genome capture followed by YCC (WGC+Y) versus non-enriched (pre-capture) libraries.\n\nResultsWe recovered 17-4,152 times more targeted unique Y-chromosome sequences after capture, where 0.01-6.2% (WGC+Y) and 0.01-23.5% (YCC) of the sequence reads were on-target, compared to 0.0002-0.004% pre-capture. In samples with endogenous DNA content greater than 0.1%, we found that WGC followed by YCC (WGC+Y) yields lower enrichment due to the loss of complexity in consecutive capture experiments, whereas in samples with lower endogenous content, WGC+Y yielded greater enrichment than YCC alone. Finally, increasing recovery of informative sites enabled us to assign Y-chromosome haplogroups to some of the archeological remains and gain insights about their paternal lineages and origins.\n\nConclusionsWe present to our knowledge the first in-solution capture-enrichment method targeting the human Y-chromosome in aDNA sequencing libraries. YCC and WGC+Y enrichments lead to an increase in the amount of Y-DNA sequences, as compared to libraries not enriched for the Y-chromosome. Our probe design effectively recovers regions of the Y-chromosome bearing phylogenetically informative sites, allowing us to identify paternal lineages with less sequencing than needed for pre-capture libraries. Finally we recommend considering the endogenous content in the experimental design and avoiding consecutive rounds of capture for low-complexity libraries, as clonality increases considerably with each round.

genomics

Genomic insights into the domestication of the chocolate tree, Theobroma cacao L.

Domestication has had a strong impact on the development of modern societies. We sequenced 200 genomes of the chocolate plant Theobroma cacao L. to show for the first time that a single population underwent strong domestication approximately 3,600 years (95% CI: 2481 - 10,903 years ago) ago, the Criollo population. We also show that during the process of domestication, there was strong selection for genes involved in the metabolism of the colored protectants anthocyanins and the stimulant theobromine, as well as disease resistance genes. Our analyses show that domesticated populations of T. cacao (Criollo) maintain a higher proportion of high frequency deleterious mutations. We also show for the first time the negative consequences the increase accumulation of deleterious mutations during domestication on the fitness of individuals (significant negative correlation between Criollo ancestry and Kg of beans per hectare per year, P = 0.000425).

genomics

Germline determinants of the somatic mutation landscape in 2,642 cancer genomes

Cancers develop through somatic mutagenesis, however germline genetic variation can markedly contribute to tumorigenesis via diverse mechanisms. We discovered and phased 88 million germline single nucleotide variants, short insertions/deletions, and large structural variants in whole genomes from 2,642 cancer patients, and employed this genomic resource to study genetic determinants of somatic mutagenesis across 39 cancer types. Our analyses implicate damaging germline variants in a variety of cancer predisposition and DNA damage response genes with specific somatic mutation patterns. Mutations in the MBD4 DNA glycosylase gene showed association with elevated C>T mutagenesis at CpG dinucleotides, a ubiquitous mutational process acting across tissues. Analysis of somatic structural variation exposed complex rearrangement patterns, involving cycles of templated insertions and tandem duplications, in BRCA1-deficient tumours. Genome-wide association analysis implicated common genetic variation at the APOBEC3 gene cluster with reduced basal levels of somatic mutagenesis attributable to APOBEC cytidine deaminases across cancer types. We further inferred over a hundred polymorphic L1/LINE elements with somatic retrotransposition activity in cancer. Our study highlights the major impact of rare and common germline variants on mutational landscapes in cancer.

genomics

An Unexpectedly Complex Architecture for Skin Pigmentation in Africans

Fewer than 15 genes have been directly associated with skin pigmentation variation in humans, leading to its characterization as a relatively simple trait. However, by assembling a global survey of quantitative skin pigmentation phenotypes, we demonstrate that pigmentation is more complex than previously assumed with genetic architecture varying by latitude. We investigate polygenicity in the Khoe and the San, populations indigenous to southern Africa, who have considerably lighter skin than equatorial Africans. We demonstrate that skin pigmentation is highly heritable, but that known pigmentation loci explain only a small fraction of the variance. Rather, baseline skin pigmentation is a complex, polygenic trait in the KhoeSan. Despite this, we identify canonical and non-canonical skin pigmentation loci, including near SLC24A5, TYRP1, SMARCA2/VLDLR, and SNX13 using a genome-wide association approach complemented by targeted resequencing. By considering diverse, under-studied African populations, we show how the architecture of skin pigmentation can vary across humans subject to different local evolutionary pressures.\n\nHighlightsO_LISkin pigmentation in Africans is far more polygenic than light skin pigmentation in Eurasians.\nC_LIO_LIKhoeSan[§] populations, which diverged early in human prehistory from other populations, have lightened skin pigmentation compared to equatorial Africans.\nC_LIO_LISkin color is highly heritable in the KhoeSan, but pigmentation variability is not well explained by previously discovered pigmentation genes.\nC_LIO_LIWe perform the first GWAS for pigmentation in African KhoeSan populations and identify canonical pigmentation loci near TYRP1 and in SLC24A5, as well as novel associations surrounding SMARCA2 and other genes.\nC_LI

genetics

Therapeutic importance of timely immunophenotyping of breast cancer in a resource-constrained setting: a retrospective hospital-based cohort study

BackgroundOrganizations that issue guidance on breast cancer recommend the use of immunohistochemistry (IHC) for providing appropriate and precise care. However, little focus has been directed to the identification of maximum allowable turnaround times for IHC, which is necessary given the diversity of hospital settings in the world. Much less effort has been committed to the development of digital tools that allow hospital administrators to monitor service utilization histories of their patients.\n\nMethodsIn this retrospective cohort study, we reviewed electronic and paper medical records of all suspected breast cancer patients treated at one secondary-care hospital of the Mexican Institute of Social Security (IMSS), located in western Mexico. We then followed three years of medical history of those patients with IHC testing.\n\nResultsIn 2014, there were 402 breast cancer patients, of which 30 were tested for some IHC biomarker (ER, PR, HER2). The subtyping allowed doctors to adjust (56.7 %) or confirm (43.3 %) the initial therapeutic regimen. The average turnaround time was 56 days. Opportune IHC testing was found to be beneficial when it was available before or during the first rounds of chemotherapy.\n\nConclusionsThe use of data mining tools applied to health record data revealed that there is an association between timely immunohistochemistry and improved outcomes in breast cancer patients. Based on this finding, inclusion of turnaround time in clinical guidelines is recommended. As much of the health data in the country becomes digitized, our visualization tools allow a digital dashboard of the hospital service utilization histories.

epidemiology

Neolithization of North Africa involved the migration of people from both the Levant and Europe

The extent to which prehistoric migrations of farmers influenced the genetic pool of western North Africans remains unclear. Archaeological evidence suggests the Neolithization process may have happened through the adoption of innovations by local Epipaleolithic communities, or by demic diffusion from the Eastern Mediterranean shores or Iberia. Here, we present the first analysis of individuals genome sequences from early and late Neolithic sites in Morocco, as well as Early Neolithic individuals from southern Iberia. We show that Early Neolithic Moroccans are distinct from any other reported ancient individuals and possess an endemic element retained in present-day Maghrebi populations, confirming a long-term genetic continuity in the region. Among ancient populations, Early Neolithic Moroccans are distantly related to Levantine Natufian hunter-gatherers ([~]9,000 BCE) and Pre-Pottery Neolithic farmers ([~]6,500 BCE). Although an expansion in Early Neolithic times is also plausible, the high divergence observed in Early Neolithic Moroccans suggests a long-term isolation and an early arrival in North Africa for this population. This scenario is consistent with early Neolithic traditions in North Africa deriving from Epipaleolithic communities who adopted certain innovations from neighbouring populations. Late Neolithic ([~]3,000 BCE) Moroccans, in contrast, share an Iberian component, supporting theories of trans-Gibraltar gene flow. Finally, the southern Iberian Early Neolithic samples share the same genetic composition as the Cardial Mediterranean Neolithic culture that reached Iberia [~]5,500 BCE. The cultural and genetic similarities of the Iberian Neolithic cultures with that of North African Neolithic sites further reinforce the model of an Iberian migration into the Maghreb.\n\nSIGNIFICANCE STATEMENTThe acquisition of agricultural techniques during the so-called Neolithic revolution has been one of the major steps forward in human history. Using next-generation sequencing and ancient DNA techniques, we directly test if Neolithization in North Africa occurred through the transmission of ideas or by demic diffusion. We show that Early Neolithic Moroccans are composed of an endemic Maghrebi element still retained in present-day North African populations and distantly related to Epipaleolithic communities from the Levant. However, late Neolithic individuals from North Africa are admixed, with a North African and a European component. Our results support the idea that the Neolithization of North Africa might have involved both the development of Epipaleolithic communities and the migration of people from Europe.

genetics

Genetic Diversity Turns a New PAGE in Our Understanding of Complex Traits

Summary/AbstractGenome-wide association studies (GWAS) have laid the foundation for investigations into the biology of complex traits, drug development, and clinical guidelines. However, the dominance of European-ancestry populations in GWAS creates a biased view of the role of human variation in disease, and hinders the equitable translation of genetic associations into clinical and public health applications. The Population Architecture using Genomics and Epidemiology (PAGE) study conducted a GWAS of 26 clinical and behavioral phenotypes in 49,839 non-European individuals. Using strategies designed for analysis of multi-ethnic and admixed populations, we confirm 574 GWAS catalog variants across these traits, and find 38 secondary signals in known loci and 27 novel loci. Our data shows strong evidence of effect-size heterogeneity across ancestries for published GWAS associations, substantial benefits for fine-mapping using diverse cohorts, and insights into clinical implications. We strongly advocate for continued, large genome-wide efforts in diverse populations to reduce health disparities.

genetics

Medical relevance of protein-truncating variants across 337,208 individuals in the UK Biobank study

Protein-truncating variants can have profound effects on gene function and are critical for clinical genome interpretation and generating therapeutic hypotheses, but their relevance to medical phenotypes has not been systematically assessed. We characterized the effect of 18,228 protein-truncating variants across 135 phenotypes from the UK Biobank and found 27 associations between medical phenotypes and protein-truncating variants in genes outside the major histocompatibility complex. We performed phenome-wide analyses and directly measured the effect of homozygous carriers, commonly referred to as \"human knockouts,\" across medical phenotypes for genes implicated to be protective against disease or associated with at least one phenotype in our study and found several genes with strong pleiotropic or non-additive effects. Our results illustrate the importance of protein-truncating variants in a variety of diseases.

genetics

DEVELOPING GENE-SPECIFIC META-PREDICTOR OF VARIANT PATHOGENICITY

Rapid, accurate, and inexpensive genome sequencing promises to transform medical care. However, a critical hurdle to enabling personalized genomic medicine is predicting the functional impact of novel genomic variation. Various methods of missense variants pathogenicity prediction have been developed by now. Here we present a new strategy for developing a pathogenicity predictor of improved accuracy by applying and training a supervised machine learning model in a gene-specific manner. Our meta-predictor combines outputs of various existing predictors, supplements them with an extended set of stability and structural features of the protein, as well as its physicochemical properties, and adds information about allele frequency from various datasets. We used such a supervised gene-specific meta-predictor approach to train the model on the CFTR gene, and to predict pathogenicity of about 1,000 variants of unknown significance that we collected from various publicly available and internal resources. Our CFTR-specific meta-predictor based on the Random Forest model performs better than other machine learning algorithms that we tested, and also outperforms other available tools, such as CADD, MutPred, SIFT, and PolyPhen-2. Our predicted pathogenicity probability correlates well with clinical measures of Cystic Fibrosis patients and experimental functional measures of mutated CFTR proteins. Training the model on one gene, in contrast to taking a genome wide approach, allows taking into account structural features specific for a particular protein, thus increasing the overall accuracy of the predictor. Collecting data from several separate resources, on the other hand, allows to accumulate allele frequency information, estimated as the most important feature by our approach, for a larger set of variants. Finally, our predictor will be hosted on the ClinGen Consortium database to make it available to CF researchers and to serve as a feasibility pilot study for other Mendelian diseases.

bioinformatics

Imputation aware tag SNP selection to improve power for multi-ethnic association studies

The emergence of very large cohorts in genomic research has facilitated a focus on genotype-imputation strategies to power rare variant association. Consequently, a new generation of genotyping arrays are being developed designed with tag single nucleotide polymorphisms (SNPs) to improve rare variant imputation. Selection of these tag SNPs poses several challenges as rare variants tend to be continentally-or even population-specific and reflect fine-scale linkage disequilibrium (LD) structure impacted by recent demographic events. To explore the landscape of tag-able variation and guide design considerations for large-cohort and biobank arrays, we developed a novel pipeline to select tag SNPs using the 26 population reference panel from Phase of the 1000 Genomes Project. We evaluate our approach using leave-one-out internal validation via standard imputation methods that allows the direct comparison of tag SNP performance by estimating the correlation of the imputed and real genotypes for each iteration of potential array sites. We show how this approach allows for an assessment of array design and performance that can take advantage of the development of deeper and more diverse sequenced reference panels. We quantify the impact of demography on tag SNP performance across populations and provide population-specific guidelines for tag SNP selection. We also examine array design strategies that target single populations versus multi-ethnic cohorts, and demonstrate a boost in performance for the latter can be obtained by prioritizing tag SNPs that contribute information across multiple populations simultaneously. Finally, we demonstrate the utility of improved array design to provide meaningful improvements in power, particularly in trans-ethnic studies. The unified framework presented will enable investigators to make informed decisions for the design of new arrays, and help empower the next phase of rare variant association for global health.

genomics