bioRxiv ScienceSearch

Biology subjects

Ferguson-Smith, A.

Publications and source records attributed to Ferguson-Smith, A..

2 recordsLinked to original sources

Multiple laboratory mouse reference genomes define strain specific haplotypes and novel functional loci

The most commonly employed mammalian model organism is the laboratory mouse. A wide variety of genetically diverse inbred mouse strains, representing distinct physiological states, disease susceptibilities, and biological mechanisms have been developed over the last century. We report full length draft de novo genome assemblies for 16 of the most widely used inbred strains and reveal for the first time extensive strain-specific haplotype variation. We identify and characterise 2,567 regions on the current Genome Reference Consortium mouse reference genome exhibiting the greatest sequence diversity between strains. These regions are enriched for genes involved in defence and immunity, and exhibit enrichment of transposable elements and signatures of recent retrotransposition events. Combinations of alleles and genes unique to an individual strain are commonly observed at these loci, reflecting distinct strain phenotypes. Several immune related loci, some in previously identified QTLs for disease response have novel haplotypes not present in the reference that may explain the phenotype. We used these genomes to improve the mouse reference genome resulting in the completion of 10 new gene structures, and 62 new coding loci were added to the reference genome annotation. Notably this high quality collection of genomes revealed a previously unannotated gene (Efcab3-like) encoding 5,874 amino acids, one of the largest known in the rodent lineage. Interestingly, Efcab3-like-/- mice exhibit severe size anomalies in four regions of the brain suggesting a mechanism of Efcab3-like regulating brain development.

genomics

Simulation based benchmarking of isoform quantification in single-cell RNA-seq

Single-cell RNA-seq has the potential to facilitate isoform quantification as the confounding factor of a mixed population of cells is eliminated. We carried out a benchmark for five popular isoform quantification tools. Performance was generally good when run on simulated data based on SMARTer and SMART-seq2 data, but was poor for simulated Drop-seq data. Importantly, the reduction in performance for single-cell RNA-seq compared with bulk RNA-seq was small. An important biological insight comes from our analysis of real data which showed that genes that express two isoforms in bulk RNA-seq predominantly express one or neither isoform in individual cells.

bioinformatics