bioRxiv ScienceSearch

Biology subjects

Jay Shendure

Publications and source records attributed to Jay Shendure.

7 recordsLinked to original sources

Massively multiplex single-cell Hi-C

We present combinatorial single cell Hi-C, a novel method that leverages combinatorial cellular indexing to measure chromosome conformation in large numbers of single cells. In this proof-of-concept, we generate and sequence combinatorial single cell Hi-C libraries for two mouse and four human cell types, comprising a total of 9,316 single cells across 5 experiments. We demonstrate the utility of single-cell Hi-C data in separating different cell types, identify previously uncharacterized cell-to-cell heterogeneity in the conformational properties of mammalian chromosomes, and demonstrate that combinatorial indexing is a generalizable molecular strategy for single-cell genomics.

Genomics

Single-molecule sequencing and conformational capture enable de novo mammalian reference genomes

The decrease in sequencing cost and increased sophistication of assembly algorithms for short-read platforms has resulted in a sharp increase in the number of species with genome assemblies. However, these assemblies are highly fragmented, with many gaps, ambiguities, and errors, impeding downstream applications. We demonstrate current state of the art for de novo assembly using the domestic goat (Capra hircus), based on long reads for contig formation, short reads for consensus validation, and scaffolding by optical and chromatin interaction mapping. These combined technologies produced the most contiguous de novo mammalian assembly to date, with chromosome-length scaffolds and only 663 gaps. Our assembly represents a >250-fold improvement in contiguity compared to the previously published C. hircus assembly, and better resolves repetitive structures longer than 1 kb, supporting the most complete repeat family and immune gene complex representation ever produced for a ruminant species.

Genomics

A systematic comparison reveals substantial differences in chromosomal versus episomal encoding of enhancer activity

Candidate enhancers can be identified on the basis of chromatin modifications, the binding of chromatin modifiers and transcription factors and cofactors, or chromatin accessibility. However, validating such candidates as bona fide enhancers requires functional characterization, typically achieved through reporter assays that test whether a sequence can drive expression of a transcriptional reporter via a minimal promoter. A longstanding concern is that reporter assays are mainly implemented on episomes, which are thought to lack physiological chromatin. However, the magnitude and determinants of differences in cis-regulation for regulatory sequences residing in episomes versus chromosomes remain almost completely unknown. To address this question in a systematic manner, we developed and applied a novel lentivirus-based massively parallel reporter assay (lentiMPRA) to directly compare the functional activities of 2,236 candidate liver enhancers in an episomal versus a chromosomally integrated context. We find that the activities of chromosomally integrated sequences are substantially different from the activities of the identical sequences assayed on episomes, and furthermore are correlated with different subsets of ENCODE annotations. The results of chromosomally-based reporter assays are also more reproducible and more strongly predictable by both ENCODE annotations and sequence-based models. With a linear model that combines chromatin annotations and sequence information, we achieve a Pearsons R2 of 0.347 for predicting the results of chromosomally integrated reporter assays. This level of prediction is better than with either chromatin annotations or sequence information alone and also outperforms predictive models of episomal assays. Our results have broad implications for how cis-regulatory elements are identified, prioritized and functionally validated.

Genomics

Whole organism lineage tracing by combinatorial and cumulative genome editing

Multicellular systems develop from single cells through a lineage, but current lineage tracing approaches scale poorly to whole organisms. Here we use genome editing to progressively introduce and accumulate diverse mutations in a DNA barcode over multiple rounds of cell division. The barcode, an array of CRISPR/Cas9 target sites, records lineage relationships in the patterns of mutations shared between cells. In cell culture and zebrafish, we show that rates and patterns of editing are tunable, and that thousands of lineage-informative barcode alleles can be generated. We find that most cells in adult zebrafish organs derive from relatively few embryonic progenitors. Genome editing of synthetic target arrays for lineage tracing (GESTALT) will help generate large-scale maps of cell lineage in multicellular systems.

Developmental Biology

Autosomal dominant multiple pterygium syndrome is caused by mutations in MYH3

Multiple pterygium syndromes (MPS) are a phenotypically and genetically heterogeneous group of rare Mendelian conditions characterized by multiple pterygia, scoliosis and congenital contractures of the limbs. MPS typically segregates as an autosomal recessive disorder but rare instances of autosomal dominant transmission have been reported. While several mutations causing recessive MPS have been identified, the genetic basis of dominant MPS remains unknown. We identified four families with dominantly transmitted MPS characterized by pterygia, camptodactyly of the hands, vertebral fusions, and scoliosis. Exome sequencing identified predicted protein-altering mutations in embryonic myosin heavy chain (MYH3) in three families. MYH3 mutations underlie distal arthrogryposis types 1, 2A and 2B, but all mutations reported to date occur in the head and neck domains. In contrast, two of the mutations found to cause MPS occurred in the tail domain. The phenotypic overlap among persons with MPS coupled with physical findings distinct from other conditions caused by mutations in MYH3, suggests that the developmental mechanism underlying MPS differs from other conditions and / or that certain functions of embryonic myosin may be perturbed by disruption of specific residues / domains. Moreover, the vertebral fusions in persons with MPS coupled with evidence of MYH3 expression in bone suggests that embryonic myosin plays a previously unknown role in skeletal development.

Genetics

De novo Mutations in NALCN Cause a Syndrome of Congenital Contractures of the Limbs and Face with Hypotonia, and Developmental Delay

Freeman-Sheldon syndrome, or distal arthrogryposis type 2A (DA2A), is an autosomal dominant condition caused by mutations in MYH3 and characterized by multiple congenital contractures of the face and limbs and normal cognitive development. We identified a subset of five simplex cases putatively diagnosed with \"DA2A with severe neurological abnormalities\" in which the proband had Congenital Contractures of the LImbs and FAce, Hypotonia, and global Developmental Delay often resulting in early death, a unique condition that we now refer to as CLIFAHDD syndrome. Exome sequencing identified missense mutations in sodium leak channel, nonselective (NALCN) in four families with CLIFAHDD syndrome. Using molecular inversion probes to screen NALCN in a cohort of 202 DA cases as well as concurrent exome sequencing of six other DA cases revealed NALCN mutations in ten additional families with \"atypical\" forms of DA. All fourteen mutations were missense variants predicted to alter amino acid residues in or near the S5 and S6 pore-forming segments of NALCN, highlighting the functional importance of these segments. In vitro functional studies demonstrated that mutant NALCN nearly abolished the expression of wildtype NALCN, suggesting that mutations that cause CLIFAHDD syndrome have a dominant negative effect. In contrast, homozygosity for mutations in other regions of NALCN has been reported in three families with an autosomal recessive condition characterized mainly by hypotonia and severe intellectual disability. Accordingly, mutations in NALCN can cause either a recessive or dominant condition with varied though overlapping phenotypic features perhaps depending on the type of mutation and affected protein domain(s).

Genetics

MIPSTR: a method for multiplex genotyping of germ-line and somatic STR variation across many individuals

Abstract Short tandem repeats (STRs) are highly mutable genetic elements that often reside in functional genomic regions. The cumulative evidence of genetic studies on individual STRs suggests that STR variation profoundly affects phenotype and contributes to trait heritability. Despite recent advances in sequencing technology, STR variation has remained largely inaccessible across many individuals compared to single nucleotide variation or copy number variation. STR genotyping with short-read sequence data is confounded by (1) the difficulty of uniquely mapping short, low-complexity reads and (2) the high rate of STR amplification stutter. Here, we present MIPSTR, a robust, scalable, and affordable method that addresses these challenges. MIPSTR uses targeted capture of STR loci by single-molecule Molecular Inversion Probes (smMIPs) and a unique mapping strategy. Targeted capture and mapping strategy resolve the first challenge; the use of single molecule information resolves the second challenge. Unlike previous methods, MIPSTR is capable of distinguishing technical error due to amplification stutter from somatic STR mutations. In proof-of-principle experiments, we use MIPSTR to determine germ-line STR genotypes for 102 STR loci with high accuracy across diverse populations of the plant A. thaliana. We show that putatively functional STRs may be identified by deviation from predicted STR variation and by association with quantitative phenotypes. Employing DNA mixing experiments and a mutant deficient in DNA repair, we demonstrate that MIPSTR can detect low-frequency somatic STR variants. MIPSTR is applicable to any organism with a high-quality reference genome and is scalable to genotyping many thousands of STR loci in thousands of individuals.

Genomics