bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Sensitivity to sequencing depth in single-cell cancer genomics

BackgroundQuerying cancer genomes at single-cell resolution is expected to provide a powerful framework to understand in detail the dynamics of cancer evolution. However, given the high costs currently associated with single-cell sequencing, together with the inevitable technical noise arising from single-cell genome amplification, cost-effective strategies that maximize the quality of single-cell data are critically needed. Taking advantage of five published single-cell whole-genome and whole-exome cancer datasets, we studied the impact of sequencing depth and sampling effort towards single-cell variant detection, including structural and driver mutations, genotyping accuracy, clonal inference and phylogenetic reconstruction, using recent tools specifically designed for single-cell data.\n\nResultsAltogether, our results suggest that, for relatively large sample sizes (25 or more cells), sequencing single tumor cells at depths >5x does not drastically improve somatic variant discovery, the characterization of clonal genotypes or the estimation of phylogenies from single tumor cells.\n\nConclusionsWe demonstrate that sequencing many individual tumor cells at a modest depth represents an effective alternative to explore the mutational landscape and clonal evolutionary patterns of cancer genomes, without the excessively high costs associated with high-coverage genome sequencing.

genomics

Significant abundance of cis configurations of mutations in diploid human genomes

To fully understand human genetic variation, one must assess the specific distribution of variants between the two chromosomal homologues of genes, and any functional units of interest, as the phase of variants can significantly impact gene function and phenotype. To this end, we have systematically analyzed 18,121 autosomal protein-coding genes in 1,092 statistically phased genomes from the 1000 Genomes Project, and an unprecedented number of 184 experimentally phased genomes from the Personal Genome Project. Here we show that mutations predicted to functionally alter the protein, and coding variants as a whole, are not randomly distributed between the two homologues of a gene, but do occur significantly more frequently in cis-than trans-configurations, with cis/trans ratios of [~]60:40. Significant cis-abundance was observed in virtually all individual genomes in all populations. Nearly all variable genes exhibited either cis, or trans configurations of protein-altering mutations in significant excess, allowing distinction of cis- and trans-abundant genes. These common patterns of phase were largely constituted by a shared, global set of phase-sensitive genes. We show significant enrichment of this global set with gene sets indicating its involvement in adaptation and evolution. Moreover, cis- and trans-abundant genes were found functionally distinguishable, and exhibited strikingly different distributional patterns of protein-altering mutations. This work establishes common patterns of phase as key characteristics of diploid human exomes and provides evidence for their potential functional significance. Thus, it highlights the importance of phase for the interpretation of protein-coding genetic variation, challenging the current conceptual and functional interpretation of autosomal genes.

genomics

A genome-wide association study for host resistance to Ostreid Herpesvirus in Pacific oysters (Crassostrea gigas)

Ostreid herpesvirus (OsHV) can cause mass mortality events in Pacific oyster aquaculture. While various factors impact on the severity of outbreaks, it is clear that genetic resistance of the host is an important determinant of mortality levels. This raises the possibility of selective breeding strategies to improve the genetic resistance of farmed oyster stocks, thereby contributing to disease control. Traditional selective breeding can be augmented by use of genetic markers, either via marker-assisted or genomic selection. The aim of the current study was to investigate the genetic architecture of resistance to OsHV in Pacific oyster, to identify genomic regions containing putative resistance genes, and to inform the use of genomics to enhance efforts to breed for resistance. To achieve this, a population of ~1,000 juvenile oysters were experimentally challenged with a virulent form of OsHV, with samples taken from mortalities and survivors for genotyping and qPCR measurement of viral load. The samples were genotyped using a recently-developed SNP array, and the genotype data were used to reconstruct the pedigree. Using these pedigree and genotype data, the first high density linkage map was constructed for Pacific oyster, containing 20,353 SNPs mapped to the ten pairs of chromosomes. Genetic parameters for resistance to OsHV were estimated, indicating a significant but low heritability for the binary trait of survival and also for viral load measures (h2 0.12 - 0.25). A genome-wide association study highlighted a region of linkage group 6 containing a significant QTL affecting host resistance. These results are an important step towards identification of genes underlying resistance to OsHV in oyster, and a step towards applying genomic data to enhance selective breeding for disease resistance in oyster aquaculture.

genomics

Complete mitochondrial genome of Glomeridesmus spelaeus (Diplopoda), a troglobitic species from Carajas iron-ore caves (Para, Brazil)

We report the complete mitochondrial genome sequence of Glomeridesmus spelaeus, the first sequenced genome of the order Gomeridesmida. The genome is 14,825 pb in length and encodes 37 mitochondrial (13 PCGs, 2 rRNA genes, 22 tRNA) genes and contains a typical AT-rich region. The base composition of the genome was A (40.1%), T (36.4%), C (15.8%), and G (7.6%), with an AT content of 76.5%. Our results indicated that Glomeridesmus spelaeus only distantly related to the other Diplopoda species with available mitochondrial genomes in the public databases. The publication of the mitogenome of G. spelaeus will contribute to the identification of troglobitic invertebrates, a very significant advance for the conservation of the troglofauna.

genomics

Genome-wide association study of suicide death:Results from the first wave of Utah completed suicide data

ObjectiveSuicide death is a highly preventable, yet growing, worldwide health crisis. To date, there has been a lack of adequately powered genomic studies of suicide, with no sizeable suicide death cohorts available for study. To address this limitation, we conducted the first comprehensive genomic analysis of suicide death, using a previously unpublished suicide cohort. MethodsThe analysis sample consisted of 3,413 population-ascertained cases of European ancestry and 14,810 ancestrally matched controls. Analytical methods included principle components analysis for ancestral matching and adjusting for population stratification, linear mixed model genome-wide association testing (conditional on genetic relatedness matrix), gene and gene set enrichment testing, polygenic score analyses, as well as SNP heritability and genetic correlation estimation using LD score regression. ResultsGWAS identified two genome-wide significant loci (6 SNPs, p<5x10-8). Gene-based analyses implicated 19 genes on chromosomes 13, 15, 16, 17, and 19 (q<0.05). Suicide heritability was estimated h2 =0.2463, SE = 0.0356 using summary statistics from a multivariate logistic GWAS adjusting for ancestry. Notably, suicide polygenic scores were robustly predictive of out of sample suicide death, as were polygenic scores for several other psychiatric disorders and psychological traits, particularly behavioral disinhibition and major depressive disorder. ConclusionsIn this report, we identify multiple genome-wide significant loci/genes, and demonstrate robust polygenic score prediction of suicide death case-control status, adjusting for ancestry, in independent training and test sets. Additionally, we report that suicide death cases have increased genetic risk for behavioral disinhibition, major depression, autism spectrum disorder, psychosis, and alcohol use disorder relative to controls. Results demonstrate the ability of polygenic scores to robustly, and multidimensionally, predict suicide death case-control status.

genomics

Sequence variation aware genome references and read mapping with the variation graph toolkit

Reference genomes guide our interpretation of DNA sequence data. However, conventional linear references are fundamentally limited in that they represent only one version of each locus, whereas the population may contain multiple variants. When the reference represents an individuals genome poorly, it can impact read mapping and introduce bias. Variation graphs are bidirected DNA sequence graphs that compactly represent genetic variation, including large scale structural variation such as inversions and duplications.1 Equivalent structures are produced by de novo genome assemblers.2,3 Here we present vg, a toolkit of computational methods for creating, manipulating, and utilizing these structures as references at the scale of the human genome. vg provides an efficient approach to mapping reads onto arbitrary variation graphs using generalized compressed suffix arrays,4 with improved accuracy over alignment to a linear reference, creating data structures to support downstream variant calling and genotyping. These capabilities make using variation graphs as reference structures for DNA sequencing practical at the scale of vertebrate genomes, or at the topological complexity of new species assemblies.

genomics

High quality whole genome sequence of an abundant Holarctic odontocete, the harbour porpoise (Phocoena phocoena)

The harbour porpoise (Phocoena phocoena) is a highly mobile cetacean found in waters across the Northern hemisphere. It occurs in coastal water and inhabits water basins that vary broadly in salinity, temperature, and food availability. These diverse habitats could drive differentiation among populations. Here we report the first harbour porpoise genome, assembled de novo from a Swedish Kattegat individual. The genome is one of the most complete cetacean genomes currently available, with a total size of 2.7 Gb and 50% of the total length found in just 34 scaffolds. Using the largest 122 scaffolds, we were able to validate a high level of homology to the chromosome-level genome assembly of the closest related species for which such resource was available, the domestic cattle (Bos taurus). The draft annotation comprises 22,154 predicted gene models, which we further annotated through matches to the NCBI nucleotide database, GO categorization, and motif prediction. To infer the adaptive abilities of this species, as well as their population history, we performed a Bayesian skyline analysis, and produced results that are concordant with the demographic history of this species, including expansion and fragmentation events. Overall, this genome assembly, together with the draft annotation, represents a crucial addition to the limited genetic markers currently available for the study of porpoises and Phocoenidae conservation, phylogeny, and evolution.

genomics

A high-quality sequence of Rosa chinensis to elucidate genome structure and ornamental traits

Rose is the worlds most important ornamental plant with economic, cultural and symbolic value. Roses are cultivated worldwide and sold as garden roses, cut flowers and potted plants. Rose has a complex genome with high heterozygosity and various ploidy levels. Our objectives were (i) to develop the first high-quality reference genome sequence for the genus Rosa by sequencing a doubled haploid, combining long and short read sequencing, and anchoring to a high-density genetic map and (ii) to study the genome structure and the genetic basis of major ornamental traits.\n\nWe produced a haploid rose line from R. chinensis Old Blush and generated the first rose genome sequence at the pseudo-molecule scale (512 Mbp with N50 of 3.4 Mb and L75 of 97). The sequence was validated using high-density diploid and tetraploid genetic maps. We delineated hallmark chromosomal features including the pericentromeric regions through annotation of TE families and positioned centromeric repeats using FISH. Genetic diversity was analysed by resequencing eight Rosa species. Combining genetic and genomic approaches, we identified potential genetic regulators of key ornamental traits, including prickle density and number of flower petals. A rose APETALA2 homologue is proposed to be the major regulator of petals number in rose. This reference sequence is an important resource for studying polyploidisation, meiosis and developmental processes as we demonstrated for flower and prickle development. This reference sequence will also accelerate breeding through the development of molecular markers linked to traits, the identification of the genes underlying them and the exploitation of synteny across Rosaceae.

genomics

Whole genome sequence of an edible and potential medicinal fungus, Cordyceps guangdongensis

Cordyceps guangdongensis is an edible fungus which has been approved as a Novel Food by the Chinese Ministry of Public Health in 2013. It also has a broad application prospect in pharmaceutical industries with many medicinal activities. In this study, the whole genome of C. guangdongensis GD15, a single spore isolate from a wild strain, was sequenced and assembled with Illumina and PacBio sequencing technology. The generated genome is 29.05 Mb in size, comprising 9 scaffolds with an average GC content of 57.01%. It is predicted to contain a total of 9150 protein-coding genes. Sequence identification and comparative analysis indicated that the assembled scaffolds contained two complete chromosomes and four single-end chromosomes, showing a high level assembly. Gene annotation revealed a diversity of transporters that could contribute to the genome size and evolution. Besides, approximately 15.49% and 13.70% genes involved in metabolic processes were annotated by KEGG and COG respectively. Genes belonging to CAZymes accounted for a proportion of 2.84% of the total genes. In addition, 435 transcription factors (TFs) were identified, which were involved in various biological processes. Among the identified TFs, the fungal transcription regulatory proteins (18.39%) and fungal-specific TFs (19.77%) represented the two largest classes of TFs. These data provided a much needed genomic resource for studying C. guangdongensis, laying a solid foundation for further genetic and biological studies, especially for elucidating the genome evolution and exploring the regulatory mechanism of fruiting body development.

genomics

Insights from deconvolution of cell subtype proportions enhance the interpretation of functional genomic data.

Cell subtype proportional differences between samples significantly contribute to variation of functional genomic properties such as gene expression or DNA methylation. Current analytical approaches typically deal with cell subtype proportion influences as a nuisance variable to be eliminated. Here we demonstrate how harvesting information about cell subtype proportions from functional genomics data provides insights into the cellular events in human phenotypes. We note a striking concordance between cell subtype proportions estimated from orthogonal genome-wide assays, and demonstrate the potential for single-cell RNA-seq data to be used in tissues for which reference cell subtype functional genomic datasets are not available. Taken together, our results confirm the importance of estimating cell subtype proportions when testing a model of cellular reprogramming in human phenotypic association studies, and the value of simultaneously testing for systematic cell subtype proportional alterations as a separate phenotypic association, gaining extra insights from functional genomic studies.

genomics

A meta-analysis of the diagnostic sensitivity and clinical utility of genome sequencing, exome sequencing and chromosomal microarray in children with suspected genetic diseases

IMPORTANCEGenetic diseases are a leading cause of childhood mortality. Whole genome sequencing (WGS) and whole exome sequencing (WES) are relatively new methods for diagnosing genetic diseases.\n\nOBJECTIVESCompare the diagnostic sensitivity (rate of causative, pathogenic or likely pathogenic genotypes in known disease genes) and rate of clinical utility (proportion in whom medical or surgical management was changed by diagnosis) of WGS, WES, and chromosomal microarrays (CMA) in children with suspected genetic diseases.\n\nDATA SOURCES AND STUDY SELECTIONSystematic review of the literature (January 2011 - August 2017) for studies of diagnostic sensitivity and/or clinical utility of WGS, WES, and/or CMA in children with suspected genetic diseases. 2% of identified studies met selection criteria.\n\nDATA EXTRACTION AND SYNTHESISTwo investigators extracted data independently following MOOSE/PRISMA guidelines.\n\nMAIN OUTCOMES AND MEASURESPooled rates and 95% Cl were estimated with a random-effects model. Metaanalysis of the rate of diagnosis was based on test type, family structure, and site of testing.\n\nRESULTSIn 36 observational series and one randomized control trial, comprising 20,068 children, the diagnostic sensitivity of WGS (0.41, 95% Cl 0.34-0.48, I2=44%) and WES (0.35, 95% Cl 0.31-0.39, I2=85%) were qualitatively greater than CMA (0.10, 95% Cl 0.08-0.12, I2=81%). Subgroup meta-analyses showed that the diagnostic sensitivity of WGS was significantly greater than CMA in studies published in 2017 (P<.0001, I2=13% and I2=40%, respectively), and the diagnostic sensitivity of WES was significantly greater than CMA in studies featuring within-cohort comparisons (P<001, I2=36%). Evidence for a significant difference in the diagnostic sensitivity of WGS and WES was lacking. In studies featuring within-cohort comparisons of singleton and trio WGS/WES, the likelihood of diagnosis was significantly greater for trios (odds ratio 2.04, 95% Cl 1.62-2.56, I2=12%; P<.0001). The diagnostic sensitivity of WGS/WES with hospital-based interpretation (0.41, 95% Cl 0.38-0.45, I2=50%) was qualitatively higher than that of reference laboratories (0.28, 95% Cl 0.24-0.32, I2=81%); this difference was significant in meta-analysis of studies published in 2017 (P=.004, I2=34% and I2=26%, respectively). The rates of clinical utility of WGS (0.27, 95% Cl 0.17-0.40, I2=54%) and WES (0.18, 95% Cl 0.13-0.24, I2-77%) were higher than CMA (0.06, 95% Cl 0.05-0.07, I2=42%); this difference was significant in meta-analysis of WGS vs CMA (P<.0001).\n\nCONCLUSIONS AND RELEVANCEIn children with suspected genetic diseases, the diagnostic sensitivity and rate of clinical utility of WGS/WES were greater than CMA. Subgroups with higher WGS/WES diagnostic sensitivity were trios and those receiving hospital-based interpretation. WGS/WES should be considered a first-line genomic test for children with suspected genetic diseases.\n\nKey PointsO_ST_ABSQuestionC_ST_ABSWhat is the relative diagnostic sensitivity and clinical utility of different genome tests in children with suspected genetic diseases?\n\nFindingsWhole genome sequencing had greater diagnostic sensitivity and clinical utility than chromosomal microarrays. Testing parent-child trios had greater diagnostic sensitivity than proband singletons. Hospital-based testing had greater diagnostic sensitivity than reference laboratories.\n\nMeaningTrio genomic sequencing is the most sensitive diagnostic test for children with suspected genetic diseases.

genomics

Robustness of Transposable Element regulation but no genomic shock observed in interspecific Arabidopsis hybrids

The merging of two divergent genomes in a hybrid is believed to trigger a \"genomic shock\", disrupting gene regulation and transposable element (TE) silencing. Here, we tested this expectation by comparing the pattern of expression of transposable elements in their native and hybrid genomic context. For this, we sequenced the transcriptome of the Arabidopsis thaliana genotype Col-0, the A. lyrata genotype MN47 and their F1 hybrid. Contrary to expectations, we observe that the level of TE expression in the hybrid is strongly correlated to levels in the parental species. We detect that at most 1.1% of expressed transposable elements belonging to two specific subfamilies change their expression level upon hybridization. Most of these changes, however, are of small magnitude. We observe that the few hybrid-specific modifications in TE expression are more likely to occur when TE insertions are close to genes. In addition, changes in epigenetic histone marks H3K9me2 and H3K27me3 following hybridization do not coincide with TEs with changed expression. Finally, we further examined TE expression in parents and hybrids exposed to severe dehydration stress. Despite the major reorganization of gene and TE expression by stress, we observe that hybridization does not lead to increased disorganization of TE expression in the hybrid. We conclude that TE expression is globally robust to hybridization and that the term \"genomic shock\" is no longerappropriate to describe the anticipated consequences of merging divergent genomes in a hybrid.

genomics

CONSTRUCTION OF WHOLE GENOMES FROM SCAFFOLDS USING SINGLE CELL STRAND-SEQ DATA

Accurate reference genome sequences provide the foundation for modern molecular biology and genomics as the interpretation of sequence data to study evolution, gene expression and epigenetics depends heavily on the quality of the genome assembly used for its alignment. Correctly organising sequenced fragments such as contigs and scaffolds in relation to each other is a critical and often challenging step in the construction of robust genome references. We previously identified misoriented regions in the mouse and human reference assemblies using Strand-seq, a single cell sequencing technique that preserves DNA directionality1, 2. Here we demonstrate the ability of Strand-seq to build and correct full-length chromosomes, by identifying which scaffolds belong to the same chromosome and determining their correct order and orientation, without the need for overlapping sequences. We demonstrate that Strand-seq exquisitely maps assembly fragments into large related groups and chromosome-sized clusters without using new assembly data. Using template strand inheritance as a bi-allelic marker, we employ genetic mapping principles to cluster scaffolds that are derived from the same chromosome and order them within the chromosome based solely on directionality of DNA strand inheritance. We prove the utility of our approach by generating improved genome assemblies for several model organisms including the ferret, pig, Xenopus, zebrafish, Tasmanian devil and the Guinea pig.

genomics

Evaluation of Whole Exome Sequencing as an Alternative of BeadChip and Whole Genome Sequencing in Human Population Genetic Analysis

Understanding the underlying genetic structure of human populations is of fundamental interest to both biological and social sciences. Advances in high-throughput genotyping technology have markedly improved our understanding of global patterns of human genetic variation. The most widely used methods for collecting variant information at the DNA-level include whole genome sequencing, which continues to remain costly, and the more economical solution of array-based techniques, as these are capable of simultaneously genotyping a pre-selected set of variable DNA sites in the human genome. The largest publicly accessible set of human genomic sequence data available today originates from exome sequencing that comprises around 1.2% of the whole genome (approximately 30 million base pairs). In this study, we compared the application of the exome dataset to the array-based dataset and to the gold standard whole genome dataset using the same population genetic analysis methods. Our results draw attention to some of the inherent problems that arise from using pre-selected SNP sets for population genetic analysis. Additionally, we demonstrate that exome sequencing provides a better alternative to the array-based methods for population genetic analysis. In this study, we propose a strategy for unbiased variant collection from exome data and offer a bioinformatics protocol for proper data processing.

genomics

Association of whole-genome and NETRIN1 signaling pathway-derived polygenic risk scores for Major Depressive Disorder and thalamic radiation white matter microstructure in UK Biobank

BackgroundMajor Depressive Disorder (MDD) is a clinically heterogeneous psychiatric disorder with a polygenic architecture. Genome-wide association studies have identified a number of risk-associated variants across the genome, and growing evidence of NETRIN1 pathway involvement. Stratifying disease risk by genetic variation within the NETRIN1 pathway may provide an important route for identification of disease mechanisms by focusing on a specific process excluding heterogeneous risk-associated variation in other pathways. Here, we sought to investigate whether MDD polygenic risk scores derived from the NETRIN1 signaling pathway (NETRIN1-PRS) and the whole genome excluding NETRIN1 pathway genes (genomic-PRS) were associated with white matter integrity.\n\nMethodsWe used two diffusion tensor imaging measures, fractional anisotropy (FA) and mean diffusivity (MD), in the most up-to-date UK Biobank neuroimaging data release (FA: N = 6,401; MD: N = 6,390).\n\nResultsWe found significantly lower FA in the superior longitudinal fasciculus ({beta} = -0.035, pcorrected = 0.029) and significantly higher MD in a global measure of thalamic radiations ({beta} = 0.029, pcorrected = 0.021), as well as higher MD in the superior ({beta} = 0.034, pcorrected = 0.039) and inferior ({beta} = 0.029, pcorrected = 0.043) longitudinal fasciculus and in the anterior ({beta} = 0.025, pcorrected = 0.046) and superior ({beta} = 0.027, pcorrected = 0.043) thalamic radiation associated with NETRIN1-PRS. Genomic-PRS was also associated with lower FA and higher MD in several tracts.\n\nConclusionsOur findings indicate that variation in the NETRIN1 signaling pathway may confer risk for MDD through effects on thalamic radiation white matter microstructure.

genetics

Immuno-genomic PanCancer Landscape Reveals Diverse Immune Escape Mechanisms and Immuno-Editing Histories

Immune reactions in the tumor micro-environment are one of the cancer hallmarks and emerging immune therapies have been proven effective in many types of cancer. To investigate cancer genome-immune interactions and the role of immuno-editing or immune escape mechanisms in cancer development, we analyzed 2,834 whole genomes and RNA-seq datasets across 31 distinct tumor types from the PanCancer Analysis of Whole Genomes (PCAWG) project with respect to key immuno-genomic aspects. We show that selective copy number changes in immune-related genes could contribute to immune escape. Furthermore, we developed an index of the immuno-editing history of each tumor sample based on the information of mutations in exonic regions and pseudogenes. Our immuno-genomic analyses of pan-cancer analyses have the potential to identify a subset of tumors with immunogenicity and diverse background or intrinsic pathways associated with their immune status and immuno-editing history.

genomics

Three invariant Hi-C interaction patterns: applications to genome assembly

Assembly of reference-quality genomes from next-generation sequencing data is a key challenge in genomics. Recently, we and others have shown that Hi-C data can be used to address several outstanding challenges in the field of genome assembly. This principle has since been developed in academia and industry, and has been used in the assembly of several major genomes. In this paper, we explore the central principles underlying Hi-C-based assembly approaches, by quantitatively defining and characterizing three invariant Hi-C interaction patterns on which these approaches can build: Intrachromosomal interaction enrichment, distance-dependent interaction decay and local interaction smoothness. Specifically, we evaluate to what degree each invariant pattern holds on a single locus level in different species, cell types and Hi-C map resolutions. We find that these patterns are generally consistent across species and cell types but are affected by sequencing depth, and that matrix balancing improves consistency of loci with all three invariant patterns. Finally, we overview current Hi-C-based assembly approaches in light of these invariant patterns and demonstrate how local interaction smoothness can be used to easily detect scaffolding errors in extremely sparse Hi-C maps. We suggest that simultaneously considering all three invariant patterns may lead to better Hi-C-based genome assembly methods.

genomics

EquCab3, an Updated Reference Genome for the Domestic Horse

EquCab2, a high-quality reference genome for the domestic horse, was released in 2007. Since then, it has served as the foundation for nearly all genomic work done in equids. Recent advances in genomic sequencing technology and computational assembly methods have allowed scientists to improve reference assemblies of large animal and plant genomes in terms of contiguity and composition. In 2014, the equine genomics research community began a project to improve the reference sequence for the horse, building upon the solid foundation of EquCab2 and incorporating new short-read data, long-read data, and proximity ligation data. The result, EquCab3, is presented here. The count of non-N bases in the incorporated chromosomes is improved from 2.33Gb in EquCab2 to 2.41Gb from EquCab3. Contiguity has also been improved nearly 40-fold with a contig N50 of 4.5Mb and scaffold contiguity enhanced to where all but one of the 32 chromosomes is comprised of a single scaffold.

genomics