bioRxiv ScienceSearch

Biology subjects

Ghareghani, M.

Publications and source records attributed to Ghareghani, M..

3 recordsLinked to original sources

De novo assembly of 64 haplotype-resolved human genomes of diverse ancestry and integrated analysis of structural variation

Long-read and strand-specific sequencing technologies together facilitate the de novo assembly of high-quality haplotype-resolved human genomes without parent-child trio data. We present 64 assembled haplotypes from 32 diverse human genomes. These highly contiguous haplotype assemblies (average contig N50: 26 Mbp) integrate all forms of genetic variation across even complex loci such as the major histocompatibility complex. We focus on 107,590 structural variants (SVs), of which 68% are inaccessible by short-read sequencing. We identify new SV hotspots (spanning megabases of gene-rich sequence), characterize 130 of the most active mobile element source elements, and find that 63% of all SVs arise by homology-mediated mechanisms--a twofold increase from previous studies. Our resource now enables reliable graph-based genotyping from short reads of up to 50,340 SVs, resulting in the identification of 1,525 expression quantitative trait loci (SV-eQTLs) as well as SV candidates for adaptive selection within the human population.

genetics

A fully phased accurate assembly of an individual human genome

The prevailing genome assembly paradigm is to produce consensus sequences that "collapse" parental haplotypes into a consensus sequence. Here, we leverage the chromosome-wide phasing and scaffolding capabilities of single-cell strand sequencing (Strand-seq)1,2 and combine them with high-fidelity (HiFi) long sequencing reads3, in a novel reference-free workflow for diploid de novo genome assembly. Employing this strategy, we produce completely phased de novo genome assemblies separately for each haplotype of a single individual of Puerto Rican origin (HG00733) in the absence of parental data. The assemblies are accurate (QV > 40), highly contiguous (contig N50 > 25 Mbp) with low switch error rates (0.4%) providing fully phased single-nucleotide variants (SNVs), indels, and structural variants (SVs). A comparison of Oxford Nanopore and PacBio phased assemblies identifies 150 regions that are preferential sites of contig breaks irrespective of sequencing technology or phasing algorithms.

bioinformatics

Single cell tri-channel-processing reveals structural variation landscapes and complex rearrangement processes

Structural variation (SV), where rearrangements delete, duplicate, invert or translocate DNA segments, is a major source of somatic cell variation. It can arise in rapid bursts, mediate genetic heterogenity, and dysregulate cancer-related pathways. The challenge to systematically discover SVs in single cells remains unsolved, with copy-neutral and complex variants typically escaping detection. We developed single cell tri-channel-processing (scTRIP), a computational framework that jointly integrates read depth, template strand and haplotype phase to comprehensively discover SVs in single cells. We surveyed SV landscapes of 565 single cell genomes, including transformed epithelial cells and patient-derived leukemic samples, and discovered abundant SV classes including inversions, translocations and large-scale genomic rearrangements mediating oncogenic dysregulation. We dissected the molecular karyotype of the leukemic samples and examined their clonal structure. Different from prior methods, scTRIP also enabled direct detection and discrimination of SV mutational processes in individual cells, including breakage-fusion-bridge cycles. scTRIP will facilitate studies of clonal evolution, genetic mosaicism and somatic SV formation, and could improve disease classification for precision medicine.

genomics