bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Real-time genomic and epidemiological investigation of a multi-institution outbreak of KPC-producing Enterobacteriaceae: a translational study

BackgroundUntil recently, KPC-producing Enterobacteriaceae were rarely identified in Australia. Following an increase in the number of incident cases across the state of Victoria, we undertook a real-time combined genomic and epidemiological investigation. The scope of this study included identifying risk factors and routes of transmission, and investigating the utility of genomics to enhance traditional field epidemiology for informing management of established widespread outbreaks.\n\nMethods and FindingsAll KPC-producing Enterobacteriaceae isolates referred to the state reference laboratory from 2012 onwards were included. Whole-genome sequencing (WGS) was performed in parallel with a detailed descriptive epidemiological investigation of each case, using Illumina sequencing on each isolate. This was complemented with PacBio long-read sequencing on selected isolates to establish high-quality reference sequences and interrogate characteristics of KPC-encoding plasmids. Initial investigations indicated the outbreak was widespread, with 86 KPC-producing Enterobacteriaceae isolates (K. pneumoniae 92%) identified from 35 different locations across metropolitan and rural Victoria between 2012-2015. Initial combined analyses of the epidemiological and genomic data resolved the outbreak into distinct nosocomial transmission networks, and identified healthcare facilities at the epicentre of KPC transmission. New cases were assigned to transmission networks in real-time, allowing focussed infection control efforts. PacBio sequencing confirmed a secondary transmission network arising from inter-species plasmid transmission. Insights from Bayesian transmission inference and analyses of within-host diversity informed the development of state-wide public health and infection control guidelines, including interventions such as an intensive approach to screening contacts following new case detection to minimise unrecognised colonisation.\n\nConclusionsA real-time combined epidemiological and genomic investigation proved critical to identifying and defining multiple transmission networks of KPC Enterobacteriaceae, while data from either investigation alone were inconclusive. The investigation was fundamental to informing infection control measures in real-time and the development of state-wide public health guidelines on carbapenemase producing Enterobacteriaceae management.

genomics

Weighted likelihood inference of genomic autozygosity patterns in dense genotype data

BackgroundGenomic regions of autozygosity (ROA) arise when an individual is homozygous for haplotypes inherited identical-by-descent from ancestors shared by both parents. Over the past decade, they have gained importance for understanding evolutionary history and the genetic basis of complex diseases and traits. However, methods to detect ROA in dense genotype data have not evolved in step with advances in genome technology that now enable us to rapidly create large high-resolution genotype datasets, limiting our ability to investigate their constituent ROA patterns.\n\nResultsWe report a weighted likelihood approach for identifying ROA in dense genotype data that accounts for autocorrelation among genotyped positions and the possibilities of unobserved mutation and recombination events, and variability in the confidence of individual genotype calls in whole genome sequence (WGS) data. Forward-time genetic simulations under two demographic scenarios that reflect situations where inbreeding and its effect on fitness are of interest suggest this approach is better powered than existing state-of-the-art methods to detect ROA at marker densities consistent with WGS and popular microarray genotyping platforms used in human and non-human studies. Moreover, we present evidence that suggests this approach is able to distinguish ROA arising via consanguinity from ROA arising via endogamy. Using subsets of The 1000 Genomes Project Phase 3 data we show that, relative to WGS, intermediate and long ROA are captured robustly with popular microarray platforms, while detection of short ROA is more variable and improves with marker density. Worldwide ROA patterns inferred from WGS data are found to accord well with those previously reported on the basis of microarray genotype data. Finally, we highlight the potential of this approach to detect genomic regions enriched for autozygosity signals in one group relative to another based upon comparisons of per-individual autozygosity likelihoods instead of inferred ROA frequencies.\n\nConclusionsThis weighted likelihood ROA detection approach can assist population- and disease-geneticists working with a wide variety of data types and species to explore ROA patterns and to identify genomic regions with differential ROA signals among groups, thereby advancing our understanding of evolutionary history and the role of recessive variation in phenotypic variation and disease.

genomics

De novo long-read assembly of a complex animal genome

Eukaryotic genome assembly remains a challenge in part because of the prevalence of complex DNA repeats. This is a particularly acute problem for holocentric nematodes because of the large number of satellite DNA sequences found throughout their genomes. These have been recalcitrant to most genome sequencing methods. At the same time, many nematodes are parasites and some represent a serious threat to human health. There is a pressing need for better molecular characterization of animal and plant parasitic nematodes. The advent of long-read DNA sequencing methods offers the promise of resolving complex genomes. Using Nippostrongylus brasiliensis as a test case, applying improved base-calling algorithms and assembly methods, we demonstrate the feasibility of de novo genome assembly matching current community standards using only MinION long reads. In doing so, we uncovered an unexpected diversity of very long and complex DNA repeat sequences, including massive tandem repeats of tRNA genes. The method has the added advantage of preserving haplotypic variants and so has the potential to be used in population analyses.

genomics

KoVariome: Korean National Standard Reference Variome database of whole genomes with comprehensive SNV, indel, CNV, and SV analyses

High-coverage whole-genome sequencing data of a single ethnicity can provide a useful catalogue of population-specific genetic variations. Herein, we report a comprehensive analysis of the Korean population, and present the Korean National Standard Reference Variome (KoVariome). As a part of the Korean Personal Genome Project (KPGP), we constructed the KoVariome database using 5.5 terabases of whole genome sequence data from 50 healthy Korean individuals with an average coverage depth of 31x. In total, KoVariome includes 12.7M single-nucleotide variants (SNVs), 1.7M short insertions and deletions (indels), 4K structural variations (SVs), and 3.6K copy number variations (CNVs). Among them, 2.4M (19%) SNVs and 0.4M (24%) indels were identified as novel. We also discovered selective enrichment of 3.8M SNVs and 0.5M indels in Korean individuals, which were used to filter out 1,271 coding-SNVs not originally removed from the 1,000 Genomes Project data when prioritizing disease-causing variants. CNV analyses revealed gene losses related to bone mineral densities and duplicated genes involved in brain development and fat reduction. Finally, KoVariome health records were used to identify novel disease-causing variants in the Korean population, demonstrating the value of high-quality ethnic variation databases for the accurate interpretation of individual genomes and the precise characterization of genetic variations.

genomics

Inter-Chromosomal Linkage Reveals Errors in the 1000 Genomes Dataset

It is often unavoidable to combine data from different sequencing centers or sequencing platforms when compiling datasets with a large number of individuals. However, the different data are likely to contain specific systematic errors that will appear as SNPs. Here, we devise a method to detect systematic errors in combined datasets. To measure quality differences between individual genomes, we study pairs of variants that reside on different chromosomes and co-occur in individuals. The abundance of these pairs of variants in different genomes is then used to detect systematic errors due to batch effects. Applying our method to the 1000 Genomes dataset, we find that coding regions are enriched for errors, where about 1% of the higher-frequency variants are predicted to be erroneous, whereas errors outside of coding regions are much rarer (<0.001%). As expected, predicted errors are less often found than other variants in a dataset that was generated with a different sequencing technology, indicating that many of the candidates are indeed errors. However, predicted 1000 Genomes errors are also found in other large datasets; our observation is thus not specific to the 1000 Genomes dataset. Our results show that batch effects can be turned into a virtue by using the resulting variation in large scale datasets to detect systematic errors.

genomics

Multi-platform discovery of haplotype-resolved structural variation in human genomes

The incomplete identification of structural variants (SVs) from whole-genome sequencing data limits studies of human genetic diversity and disease association. Here, we apply a suite of long-read, short-read, and strand-specific sequencing technologies, optical mapping, and variant discovery algorithms to comprehensively analyze three human parent-child trios to define the full spectrum of human genetic variation in a haplotype-resolved manner. We identify 818,054 indel variants (<50 bp) and 27,622 SVs ([&ge;]50 bp) per human genome. We also discover 156 inversions per genome--most of which previously escaped detection. Fifty-eight of the inversions we discovered intersect with the critical regions of recurrent microdeletion and microduplication syndromes. Taken together, our SV callsets represent a sevenfold increase in SV detection compared to most standard high-throughput sequencing studies, including those from the 1000 Genomes Project. The method and the dataset serve as a gold standard for the scientific community and we make specific recommendations for maximizing structural variation sensitivity for future large-scale genome sequencing studies.

genomics

A modified sequence capture approach allowing standard and methylation analyses of the same enriched genomic DNA sample

BackgroundBread wheat has a large complex genome that makes whole genome resequencing costly. Therefore, genome complexity reduction techniques such as sequence capture make re-sequencing cost effective. With a high-quality draft wheat genome now available it is possible to design capture probe sets and to use them to accurately genotype and anchor SNPs to the genome. Furthermore, in addition to genetic variation, epigenetic variation provides a source of natural variation contributing to changes in gene expression and phenotype that can be profiled at the base pair level using sequence capture coupled with bisulphite treatment. Here, we present a new 12 Mbp wheat capture probe set, that allows both the profiling of genotype and methylation from the same DNA sample. Furthermore, we present a method, based on Agilent SureSelect Methyl-Seq, that will use a single capture assay as a starting point to allow both DNA sequencing and methyl-seq.\n\nResultsOur method uses a single capture assay that is sequentially split and used for both DNA sequencing and methyl-seq. The resultant genotype and epi-type data is highly comparable in terms of coverage and SNP/methylation site identification to that generated from separate captures for DNA sequencing and methyl-seq. Furthermore, by defining SNP frequencies in a diverse landrace from the Watkins collection we highlight the importance of having genotype data to prevent false positive methylation calls. Finally, we present the design of a new 12 Mbp wheat capture and demonstrate its successful application to re-sequence wheat.\n\nConclusionWe present a cost-effective method for performing both DNA sequencing and methyl-seq from a single capture reaction thus reducing reagent costs, sample preparation time and DNA requirements for these complementary analyses.

genomics

Nanopore Sequencing Reveals High-Resolution Structural Variation in the Cancer Genome

Acquired genomic structural variants (SVs) are major hallmarks of the cancer genome. Their complexity has been challenging to reconstruct from short-read sequencing data. Here, we exploit the long-read sequencing capability of the nanopore platform using our customized pipeline, Picky, to reveal SVs of diverse architecture in a breast cancer model. From modest sequencing coverage, we identified the full spectrum of SVs with superior specificity and sensitivity relative to short-read analyses and uncovered repetitive DNA as the major source of variation. Examination of the genome-wide breakpoints at nucleotide-resolution uncovered micro-insertions as the common structural features associated with SVs. Breakpoint density across the genome is associated with propensity for inter-chromosomal connectivity and transcriptional regulation. Furthermore, an over-representation of reciprocal translocations from chromosomal double-crossovers was observed through phased SVs. The comprehensive characterization of SVs using the robust long-read sequencing approach in cancer cohorts will facilitate strategies to monitor genome stability during tumor evolution and improve therapeutic intervention.

genomics

Summarizing Performance for Genome Scale Measurement of miRNA: Reference Samples and Metrics

BackgroundThe potential utility of microRNA as biomarkers for early detection of cancer and other diseases is being investigated with genome-scale profiling of differentially expressed microRNA. Processes for measurement assurance are critical components of genome-scale measurements. Here, we evaluated the utility of a set of total RNA samples, designed with between-sample differences in the relative abundance of miRNAs, as process controls.\n\nResultsThree pure total human RNA samples (brain, liver, and placenta) and two different mixtures of these components were evaluated as measurement assurance control samples on multiple measurement systems at multiple sites and over multiple rounds. In silico modeling of mixtures provided benchmark values for comparison with physical mixtures. Biomarker development laboratories using next-generation sequencing (NGS) or genome-scale hybridization assays participated in the study and returned data from the samples using their routine workflows. Multiplexed and single assay reverse-transcription PCR (RT-PCR) was used to confirm in silico predicted sample differences. Data visualizations and summary metrics for genome-scale miRNA profiling assessment were developed using this dataset, and a range of performance was observed. These metrics have been incorporated into an online data analysis pipeline and provide a convenient dashboard view of results from experiments following the described design. The website also serves as a repository for the accumulation of performance values providing new participants in the project an opportunity to learn what may be achievable with similar measurement processes.\n\nConclusionsThe set of reference samples used in this study provides benchmark values suitable for assessing genome-scale miRNA profiling processes. Incorporation of these metrics into an online resource allows laboratories to periodically evaluate their performance and assess any changes introduced into their measurement process.

genomics

Sensitivity to sequencing depth in single-cell cancer genomics

BackgroundQuerying cancer genomes at single-cell resolution is expected to provide a powerful framework to understand in detail the dynamics of cancer evolution. However, given the high costs currently associated with single-cell sequencing, together with the inevitable technical noise arising from single-cell genome amplification, cost-effective strategies that maximize the quality of single-cell data are critically needed. Taking advantage of five published single-cell whole-genome and whole-exome cancer datasets, we studied the impact of sequencing depth and sampling effort towards single-cell variant detection, including structural and driver mutations, genotyping accuracy, clonal inference and phylogenetic reconstruction, using recent tools specifically designed for single-cell data.\n\nResultsAltogether, our results suggest that, for relatively large sample sizes (25 or more cells), sequencing single tumor cells at depths >5x does not drastically improve somatic variant discovery, the characterization of clonal genotypes or the estimation of phylogenies from single tumor cells.\n\nConclusionsWe demonstrate that sequencing many individual tumor cells at a modest depth represents an effective alternative to explore the mutational landscape and clonal evolutionary patterns of cancer genomes, without the excessively high costs associated with high-coverage genome sequencing.

genomics

Significant abundance of cis configurations of mutations in diploid human genomes

To fully understand human genetic variation, one must assess the specific distribution of variants between the two chromosomal homologues of genes, and any functional units of interest, as the phase of variants can significantly impact gene function and phenotype. To this end, we have systematically analyzed 18,121 autosomal protein-coding genes in 1,092 statistically phased genomes from the 1000 Genomes Project, and an unprecedented number of 184 experimentally phased genomes from the Personal Genome Project. Here we show that mutations predicted to functionally alter the protein, and coding variants as a whole, are not randomly distributed between the two homologues of a gene, but do occur significantly more frequently in cis-than trans-configurations, with cis/trans ratios of [~]60:40. Significant cis-abundance was observed in virtually all individual genomes in all populations. Nearly all variable genes exhibited either cis, or trans configurations of protein-altering mutations in significant excess, allowing distinction of cis- and trans-abundant genes. These common patterns of phase were largely constituted by a shared, global set of phase-sensitive genes. We show significant enrichment of this global set with gene sets indicating its involvement in adaptation and evolution. Moreover, cis- and trans-abundant genes were found functionally distinguishable, and exhibited strikingly different distributional patterns of protein-altering mutations. This work establishes common patterns of phase as key characteristics of diploid human exomes and provides evidence for their potential functional significance. Thus, it highlights the importance of phase for the interpretation of protein-coding genetic variation, challenging the current conceptual and functional interpretation of autosomal genes.

genomics

A genome-wide association study for host resistance to Ostreid Herpesvirus in Pacific oysters (Crassostrea gigas)

Ostreid herpesvirus (OsHV) can cause mass mortality events in Pacific oyster aquaculture. While various factors impact on the severity of outbreaks, it is clear that genetic resistance of the host is an important determinant of mortality levels. This raises the possibility of selective breeding strategies to improve the genetic resistance of farmed oyster stocks, thereby contributing to disease control. Traditional selective breeding can be augmented by use of genetic markers, either via marker-assisted or genomic selection. The aim of the current study was to investigate the genetic architecture of resistance to OsHV in Pacific oyster, to identify genomic regions containing putative resistance genes, and to inform the use of genomics to enhance efforts to breed for resistance. To achieve this, a population of ~1,000 juvenile oysters were experimentally challenged with a virulent form of OsHV, with samples taken from mortalities and survivors for genotyping and qPCR measurement of viral load. The samples were genotyped using a recently-developed SNP array, and the genotype data were used to reconstruct the pedigree. Using these pedigree and genotype data, the first high density linkage map was constructed for Pacific oyster, containing 20,353 SNPs mapped to the ten pairs of chromosomes. Genetic parameters for resistance to OsHV were estimated, indicating a significant but low heritability for the binary trait of survival and also for viral load measures (h2 0.12 - 0.25). A genome-wide association study highlighted a region of linkage group 6 containing a significant QTL affecting host resistance. These results are an important step towards identification of genes underlying resistance to OsHV in oyster, and a step towards applying genomic data to enhance selective breeding for disease resistance in oyster aquaculture.

genomics

Complete mitochondrial genome of Glomeridesmus spelaeus (Diplopoda), a troglobitic species from Carajas iron-ore caves (Para, Brazil)

We report the complete mitochondrial genome sequence of Glomeridesmus spelaeus, the first sequenced genome of the order Gomeridesmida. The genome is 14,825 pb in length and encodes 37 mitochondrial (13 PCGs, 2 rRNA genes, 22 tRNA) genes and contains a typical AT-rich region. The base composition of the genome was A (40.1%), T (36.4%), C (15.8%), and G (7.6%), with an AT content of 76.5%. Our results indicated that Glomeridesmus spelaeus only distantly related to the other Diplopoda species with available mitochondrial genomes in the public databases. The publication of the mitogenome of G. spelaeus will contribute to the identification of troglobitic invertebrates, a very significant advance for the conservation of the troglofauna.

genomics

Genome-wide association study of suicide death:Results from the first wave of Utah completed suicide data

ObjectiveSuicide death is a highly preventable, yet growing, worldwide health crisis. To date, there has been a lack of adequately powered genomic studies of suicide, with no sizeable suicide death cohorts available for study. To address this limitation, we conducted the first comprehensive genomic analysis of suicide death, using a previously unpublished suicide cohort. MethodsThe analysis sample consisted of 3,413 population-ascertained cases of European ancestry and 14,810 ancestrally matched controls. Analytical methods included principle components analysis for ancestral matching and adjusting for population stratification, linear mixed model genome-wide association testing (conditional on genetic relatedness matrix), gene and gene set enrichment testing, polygenic score analyses, as well as SNP heritability and genetic correlation estimation using LD score regression. ResultsGWAS identified two genome-wide significant loci (6 SNPs, p<5x10-8). Gene-based analyses implicated 19 genes on chromosomes 13, 15, 16, 17, and 19 (q<0.05). Suicide heritability was estimated h2 =0.2463, SE = 0.0356 using summary statistics from a multivariate logistic GWAS adjusting for ancestry. Notably, suicide polygenic scores were robustly predictive of out of sample suicide death, as were polygenic scores for several other psychiatric disorders and psychological traits, particularly behavioral disinhibition and major depressive disorder. ConclusionsIn this report, we identify multiple genome-wide significant loci/genes, and demonstrate robust polygenic score prediction of suicide death case-control status, adjusting for ancestry, in independent training and test sets. Additionally, we report that suicide death cases have increased genetic risk for behavioral disinhibition, major depression, autism spectrum disorder, psychosis, and alcohol use disorder relative to controls. Results demonstrate the ability of polygenic scores to robustly, and multidimensionally, predict suicide death case-control status.

genomics

Sequence variation aware genome references and read mapping with the variation graph toolkit

Reference genomes guide our interpretation of DNA sequence data. However, conventional linear references are fundamentally limited in that they represent only one version of each locus, whereas the population may contain multiple variants. When the reference represents an individuals genome poorly, it can impact read mapping and introduce bias. Variation graphs are bidirected DNA sequence graphs that compactly represent genetic variation, including large scale structural variation such as inversions and duplications.1 Equivalent structures are produced by de novo genome assemblers.2,3 Here we present vg, a toolkit of computational methods for creating, manipulating, and utilizing these structures as references at the scale of the human genome. vg provides an efficient approach to mapping reads onto arbitrary variation graphs using generalized compressed suffix arrays,4 with improved accuracy over alignment to a linear reference, creating data structures to support downstream variant calling and genotyping. These capabilities make using variation graphs as reference structures for DNA sequencing practical at the scale of vertebrate genomes, or at the topological complexity of new species assemblies.

genomics

High quality whole genome sequence of an abundant Holarctic odontocete, the harbour porpoise (Phocoena phocoena)

The harbour porpoise (Phocoena phocoena) is a highly mobile cetacean found in waters across the Northern hemisphere. It occurs in coastal water and inhabits water basins that vary broadly in salinity, temperature, and food availability. These diverse habitats could drive differentiation among populations. Here we report the first harbour porpoise genome, assembled de novo from a Swedish Kattegat individual. The genome is one of the most complete cetacean genomes currently available, with a total size of 2.7 Gb and 50% of the total length found in just 34 scaffolds. Using the largest 122 scaffolds, we were able to validate a high level of homology to the chromosome-level genome assembly of the closest related species for which such resource was available, the domestic cattle (Bos taurus). The draft annotation comprises 22,154 predicted gene models, which we further annotated through matches to the NCBI nucleotide database, GO categorization, and motif prediction. To infer the adaptive abilities of this species, as well as their population history, we performed a Bayesian skyline analysis, and produced results that are concordant with the demographic history of this species, including expansion and fragmentation events. Overall, this genome assembly, together with the draft annotation, represents a crucial addition to the limited genetic markers currently available for the study of porpoises and Phocoenidae conservation, phylogeny, and evolution.

genomics

A high-quality sequence of Rosa chinensis to elucidate genome structure and ornamental traits

Rose is the worlds most important ornamental plant with economic, cultural and symbolic value. Roses are cultivated worldwide and sold as garden roses, cut flowers and potted plants. Rose has a complex genome with high heterozygosity and various ploidy levels. Our objectives were (i) to develop the first high-quality reference genome sequence for the genus Rosa by sequencing a doubled haploid, combining long and short read sequencing, and anchoring to a high-density genetic map and (ii) to study the genome structure and the genetic basis of major ornamental traits.\n\nWe produced a haploid rose line from R. chinensis Old Blush and generated the first rose genome sequence at the pseudo-molecule scale (512 Mbp with N50 of 3.4 Mb and L75 of 97). The sequence was validated using high-density diploid and tetraploid genetic maps. We delineated hallmark chromosomal features including the pericentromeric regions through annotation of TE families and positioned centromeric repeats using FISH. Genetic diversity was analysed by resequencing eight Rosa species. Combining genetic and genomic approaches, we identified potential genetic regulators of key ornamental traits, including prickle density and number of flower petals. A rose APETALA2 homologue is proposed to be the major regulator of petals number in rose. This reference sequence is an important resource for studying polyploidisation, meiosis and developmental processes as we demonstrated for flower and prickle development. This reference sequence will also accelerate breeding through the development of molecular markers linked to traits, the identification of the genes underlying them and the exploitation of synteny across Rosaceae.

genomics

Whole genome sequence of an edible and potential medicinal fungus, Cordyceps guangdongensis

Cordyceps guangdongensis is an edible fungus which has been approved as a Novel Food by the Chinese Ministry of Public Health in 2013. It also has a broad application prospect in pharmaceutical industries with many medicinal activities. In this study, the whole genome of C. guangdongensis GD15, a single spore isolate from a wild strain, was sequenced and assembled with Illumina and PacBio sequencing technology. The generated genome is 29.05 Mb in size, comprising 9 scaffolds with an average GC content of 57.01%. It is predicted to contain a total of 9150 protein-coding genes. Sequence identification and comparative analysis indicated that the assembled scaffolds contained two complete chromosomes and four single-end chromosomes, showing a high level assembly. Gene annotation revealed a diversity of transporters that could contribute to the genome size and evolution. Besides, approximately 15.49% and 13.70% genes involved in metabolic processes were annotated by KEGG and COG respectively. Genes belonging to CAZymes accounted for a proportion of 2.84% of the total genes. In addition, 435 transcription factors (TFs) were identified, which were involved in various biological processes. Among the identified TFs, the fungal transcription regulatory proteins (18.39%) and fungal-specific TFs (19.77%) represented the two largest classes of TFs. These data provided a much needed genomic resource for studying C. guangdongensis, laying a solid foundation for further genetic and biological studies, especially for elucidating the genome evolution and exploring the regulatory mechanism of fruiting body development.

genomics