bioRxiv Science⌕ Search

Biology subjects

DeFelice, M.

Publications and source records attributed to DeFelice, M..

3 recordsLinked to original sources

A blended genome and exome sequencing method captures genetic variation in an unbiased, high-quality, and cost-effective manner

We deployed the Blended Genome Exome (BGE), a DNA library blending approach that generates low pass whole genome (1-4x mean depth) and deep whole exome (30-40x mean depth) data in a single sequencing run. This technology is cost-effective, empowers most genomic discoveries possible with deep whole genome sequencing, and provides an unbiased method to capture the diversity of common SNP variation across the globe. To evaluate this new technology at scale, we applied BGE to sequence >53,000 samples from the Populations Underrepresented in Mental Illness Associations Studies (PUMAS) Project, which included participants across African, African American, and Latin American populations. We evaluated the accuracy of BGE imputed genotypes against raw genotype calls from the Illumina Global Screening Array. All PUMAS cohorts had R2 concordance [&ge;]95% among SNPs with MAF[&ge;]1%, and never fell below [&ge;]90% R2 for SNPs with MAF<1%. Furthermore, concordance rates among local ancestries within two recently admixed cohorts were consistent among SNPs with MAF[&ge;]1%, with only minor deviations in SNPs with MAF<1%. We also benchmarked the discovery capacity of BGE to access protein-coding copy number variants (CNVs) against deep whole genome data, finding that deletions and duplications spanning at least 3 exons had a positive predicted value of [~]90%. Our results demonstrate BGE scalability and efficacy in capturing SNPs, indels, and CNVs in the human genome at 28% of the cost of deep whole-genome sequencing. BGE is poised to enhance access to genomic testing and empower genomic discoveries, particularly in underrepresented populations.

genomics↗

Sharing Data from the Human Tumor Atlas Network through Standards, Infrastructure, and Community Engagement

The Data Coordinating Center (DCC) of the Human Tumor Atlas Network (HTAN) has played a crucial role in enabling the broad sharing and effective utilization of HTAN data within the scienti[fi]c community. Data from the [fi]rst phase of HTAN are now available publicly. We describe the diverse datasets and modalities shared, multiple access routes to HTAN assay data and metadata, data standards, technical infrastructure and governance approaches, as well as our approach to sustained community engagement. HTAN data can be accessed via the HTAN Portal, explored in visualization tools--including CellxGene, Minerva, and cBioPortal--and analyzed in the cloud through the NCI Cancer Research Data Commons nodes. We have developed a streamlined infrastructure to ingest and disseminate data by leveraging the Synapse platform. Taken together, the HTAN DCCs approach demonstrates a successful model for coordinating, standardizing, and disseminating complex cancer research data via multiple resources in the cancer data ecosystem, offering valuable insights for similar consortia, and researchers looking to leverage HTAN data.

cancer biology↗

Blended Genome Exome (BGE) as a Cost Efficient Alternative to Deep Whole Genomes or Arrays

Genomic scientists have long been promised cheaper DNA sequencing, but deep whole genomes are still costly, especially when considered for large cohorts in population-level studies. More affordable options include microarrays + imputation, whole exome sequencing (WES), or low-pass whole genome sequencing (WGS) + imputation. WES + array + imputation has recently been shown to yield 99% of association signals detected by WGS. However, a method free from ascertainment biases of arrays or the need for merging different data types that still benefits from deeper exome coverage to enhance novel coding variant detection does not exist. We developed a new, combined, "Blended Genome Exome" (BGE) in which a whole genome library is generated, an aliquot of that genome is amplified by PCR, the exome regions are selected and enriched, and the genome and exome libraries are combined back into a single tube for sequencing (33% exome, 67% genome). This creates a single CRAM with a low-coverage whole genome (2-3x) combined with a higher coverage exome (30-40x). This BGE can be used for imputing common variants throughout the genome as well as for calling rare coding variants. We tested this new method and observed >99% r2 concordance between imputed BGE data and existing 30x WGS data for exome and genome variants. BGE can serve as a useful and cost-efficient alternative sequencing product for genomic researchers, requiring ten-fold less sequencing compared to 30x WGS without the need for complicated harmonization of array and sequencing data.

genomics↗