bioRxiv ScienceSearch

Biology subjects

Guryev, V.

Publications and source records attributed to Guryev, V..

7 recordsLinked to original sources

CONSTRUCTION OF WHOLE GENOMES FROM SCAFFOLDS USING SINGLE CELL STRAND-SEQ DATA

Accurate reference genome sequences provide the foundation for modern molecular biology and genomics as the interpretation of sequence data to study evolution, gene expression and epigenetics depends heavily on the quality of the genome assembly used for its alignment. Correctly organising sequenced fragments such as contigs and scaffolds in relation to each other is a critical and often challenging step in the construction of robust genome references. We previously identified misoriented regions in the mouse and human reference assemblies using Strand-seq, a single cell sequencing technique that preserves DNA directionality1, 2. Here we demonstrate the ability of Strand-seq to build and correct full-length chromosomes, by identifying which scaffolds belong to the same chromosome and determining their correct order and orientation, without the need for overlapping sequences. We demonstrate that Strand-seq exquisitely maps assembly fragments into large related groups and chromosome-sized clusters without using new assembly data. Using template strand inheritance as a bi-allelic marker, we employ genetic mapping principles to cluster scaffolds that are derived from the same chromosome and order them within the chromosome based solely on directionality of DNA strand inheritance. We prove the utility of our approach by generating improved genome assemblies for several model organisms including the ferret, pig, Xenopus, zebrafish, Tasmanian devil and the Guinea pig.

genomics

Reduced expression of C/EBPβ-LIP extends health- and lifespan in mice

Ageing is associated with physical decline and the development of age-related diseases such as metabolic disorders and cancer. Few conditions are known that attenuate the adverse effects of ageing, including calorie restriction (CR) and reduced signalling through the mechanistic target of rapamycin complex 1 (mTORC1) pathway. Synthesis of the metabolic transcription factor C/EBP {beta} -LIP is stimulated by mTORC1, which critically depends on a short upstream open reading frame (uORF) in the C/EBP {beta}-mRNA. Here we describe that reduced C/EBP {beta} -LIP expression due to genetic ablation of the uORF delays the development of age-associated phenotypes in mice. Moreover, female C/EBP {beta}{Delta}uORF mice display an extended lifespan. Since LIP levels increase upon aging in wt mice, our data reveal an important role for C/EBP{beta} in the aging process and suggest that restriction of LIP expression sustains health and fitness. Thus, therapeutic strategies targeting C/EBP {beta} -LIP may offer new possibilities to treat age-related diseases and to prolong healthspan.

cell biology

Multi-platform discovery of haplotype-resolved structural variation in human genomes

The incomplete identification of structural variants (SVs) from whole-genome sequencing data limits studies of human genetic diversity and disease association. Here, we apply a suite of long-read, short-read, and strand-specific sequencing technologies, optical mapping, and variant discovery algorithms to comprehensively analyze three human parent-child trios to define the full spectrum of human genetic variation in a haplotype-resolved manner. We identify 818,054 indel variants (<50 bp) and 27,622 SVs ([&ge;]50 bp) per human genome. We also discover 156 inversions per genome--most of which previously escaped detection. Fifty-eight of the inversions we discovered intersect with the critical regions of recurrent microdeletion and microduplication syndromes. Taken together, our SV callsets represent a sevenfold increase in SV detection compared to most standard high-throughput sequencing studies, including those from the 1000 Genomes Project. The method and the dataset serve as a gold standard for the scientific community and we make specific recommendations for maximizing structural variation sensitivity for future large-scale genome sequencing studies.

genomics

BLM helicase suppresses recombination at G-quadruplex motifs in transcribed genes

Bloom syndrome is a cancer predisposition disorder caused by mutations in the BLM helicase gene. Cells from persons with Bloom syndrome exhibit striking genomic instability characterized by excessive sister chromatid exchange events (SCEs). We applied single-cell DNA template strand-sequencing (Strand-seq) to map the genomic locations of SCEs at a resolution that is orders of magnitude better than was previously possible. Our results show that, in the absence of BLM, sister chromatid exchanges in human and murine cells do not occur randomly throughout the genome but are strikingly enriched at coding regions, specifically at sites of putative guanine quadruplex (G4) motifs in transcribed genes. We propose that BLM protects against genome instability by suppressing recombination at sites of G4 structures, particularly in transcribed regions of the genome.

cell biology

Double-strand breaks are not the main cause of spontaneous sister chromatid exchange in wild-type yeast cells

SummaryHomologous recombination involving sister chromatids is the most accurate, and thus most frequently used, form of recombination-mediated DNA repair. Despite its importance, sister chromatid recombination is not easily studied because it does not result in a change in DNA sequence, making recombination between sister chromatids difficult to detect. We have previously developed a novel DNA template strand sequencing technique, called Strand-seq, that can be used to map sister chromatid exchange (SCE) events genome-wide in single cells. An increase in the rate of SCE is an indicator of elevated recombination activity and of genome instability, which is a hallmark of cancer. In this study, we have adapted Strand-seq to detect SCE in the yeast Saccharomyces cerevisiae. Contrary to what is commonly thought, we find that most spontaneous SCE events are not due to the repair of DNA double-strand breaks.

molecular biology

A platform for efficient transgenesis in Macrostomum lignano, a flatworm model organism for stem cell research

Regeneration-capable flatworms are informative research models to study the mechanisms of stem cell regulation, regeneration and tissue patterning. However, the lack of transgenesis methods significantly hampers their wider use. Here we report development of a transgenesis method for Macrostomum lignano, a basal flatworm with excellent regeneration capacity. We demonstrate that microinjection of DNA constructs into fertilized one-cell stage eggs, followed by a low dose of irradiation, frequently results in random integration of the transgene in the genome and its stable transmission through the germline. To facilitate selection of promoter regions for transgenic reporters, we assembled and annotated the M. lignano genome, including genome-wide mapping of transcription start regions, and showed its utility by generating multiple stable transgenic lines expressing fluorescent proteins under several tissue-specific promoters. The reported transgenesis method and annotated genome sequence will permit sophisticated genetic studies on stem cells and regeneration using M. lignano as a model organism.

bioengineering

Dense And Accurate Whole-Chromosome Haplotyping Of Individual Genomes

The diploid nature of the genome is neglected in many analyses done today, where a genome is perceived as a set of unphased variants with respect to a reference genome. Many important biological phenomena such as compound heterozygosity and epistatic effects between enhancers and target genes, however, can only be studied when haplotype-resolved genomes are available. This lack of haplotype-level analyses can be explained by a dearth of methods to produce dense and accurate chromosome-length haplotypes at reasonable costs. Here we introduce an integrative phasing strategy that combines global, but sparse haplotypes obtained from strand-specific single cell sequencing (Strand-seq) with dense, yet local, haplotype information available through long-read or linked-read sequencing. Our experiments provide comprehensive guidance on favorable combinations of Strand-seq libraries and sequencing coverages to obtain complete and genome-wide haplotypes of a single individual genome (NA12878) at manageable costs. We were able to reliably assign > 95% of alleles to their parental haplotypes using as few as 10 Strand-seq libraries in combination with 10-fold coverage PacBio data or, alternatively, 10X Genomics linked-read sequencing data. We conclude that the combination of Strand-seq with different sequencing technologies represents an attractive solution to chart the unique genetic variation of diploid genomes.

bioinformatics