bioRxiv ScienceSearch

Biology subjects

Marchini, J.

Publications and source records attributed to Marchini, J..

5 recordsLinked to original sources

Unified single-cell analysis of testis gene regulation and pathology in 5 mouse strains

By removing the confounding factor of cellular heterogeneity, single cell genomics can revolutionize the study of development and disease, but methods are needed to simplify comparison among individuals. To develop such a framework, we assayed the transcriptome in 62,600 single cells from the testes of wildtype mice, and mice with gonadal defects due to disruption of the genes Mlh3, Hormad1, Cul4a or Cnp. The resulting expression atlas of distinct cell clusters revealed novel markers and new insights into testis gene regulation. By jointly analysing mutant and wildtype cells using a model-based factor analysis method, SDA, we decomposed our data into 46 components that identify novel meiotic gene regulatory programmes, mutant-specific pathological processes, and technical effects. Moreover, we identify, de novo, DNA sequence motifs associated with each component, and show that SDA can be used to impute expression values from single cell data. Analysis of SDA components also led us to identify a rare population of macrophages within the seminiferous tubules of Mlh3-/- and Hormad1-/- testes, an area typically associated with immune privilege. We provide a web application to enable interactive exploration of testis gene expression and components at http://www.stats.ox.ac.uk/~wells/testisAtlas.html

genomics

The spatial correspondence and genetic influence of inter-hemispheric connectivity with white matter microstructure

Microscopic features (i.e., microstructure) of axons affect neural circuit activity through characteristics such as conduction speed. Deeper understanding of structure-function relationships and translating this into human neuroscience has been limited by the paucity of studies relating axonal microstructure in white matter pathways to functional connectivity (synchrony) between macroscopic brain regions. Using magnetic resonance imaging data in 11354 subjects, we constructed multi-variate models that predict the functional connectivity of pairs of brain regions from the microstructural signature of white matter pathways that connect them. Microstructure-derived models provide predictions of functional connectivity that were significant in up to 86% of the brain region pairs considered. These relationships are specific to the relevant white matter pathway and have high reproducibility. The microstructure-function relationships are associated to genetic variants (single-nucleotide polymorphisms), co-located with genes DAAM1 and LPAR1, that have previously been reported to play a role in neural development. Our results demonstrate that variation in white matter microstructure across individuals consistently and specifically predicts functional connectivity, and that this relationship is underpinned by genetic variability.

neuroscience

BGEN: a binary file format for imputed genotype and haplotype data

The impact of modern technology on genetic epidemiology has been significant, with studies comprising millions of individuals assessed at tens of millions of genetic variants now becoming common. Studies on this scale provide logistical and analytic challenges starting with the issue of efficiently storing, transmitting, and accessing underlying data. Here we present a binary file format (the BGEN format) that can store both directly-typed and statistically imputed genotype data, and achieves substantial space savings by data compression and the use of an efficient representation for probabilities. We investigate the properties of this format using imputed data from the UK BiLEVE study, demonstrating both storage efficiency, and fast data loading performance on the order of hundreds of millions of imputed genotypes per second. To make using BGEN as easy as possible, we provide a detailed specification and a freely available reference implementation, and we leverage this by developing additional tools including an indexing tool (bgenix) and an R package (rbgen) that permits loading of BGEN-encoded data into the R statistical programming environment. The UK Biobank is one of a number of projects that have used BGEN for release of imputed data, and we expect the format to continue to be widely implemented and used.

bioinformatics

The genetic basis of human brain structure and function: 1,262 genome-wide associations found from 3,144 GWAS of multimodal brain imaging phenotypes from 9,707 UK Biobank participants

The genetic basis of brain structure and function is largely unknown. We carried out genome-wide association studies of 3,144 distinct functional and structural brain imaging derived phenotypes in UK Biobank (discovery dataset 8,428 subjects). We show that many of these phenotypes are heritable. We identify 148 clusters of SNP-imaging associations with lead SNPs that replicate at p<0.05, when we would expect 21 to replicate by chance. Notable significant and interpretable associations include: iron transport and storage genes, related to changes in T2* in subcortical regions; extracellular matrix and the epidermal growth factor genes, associated with white matter micro-structure and lesion volume; genes regulating mid-line axon guidance development associated with pontine crossing tract organisation; and overall 17 genes involved in development, pathway signalling and plasticity. Our results provide new insight into the genetic architecture of the brain with relevance to complex neurological and psychiatric disorders, as well as brain development and aging. The full set of results is available on the interactive Oxford Brain Imaging Genetics (BIG) web browser.

genetics

Genome-wide genetic data on ~500,000 UK Biobank participants

The UK Biobank project is a large prospective cohort study of ~500,000 individuals from across the United Kingdom, aged between 40-69 at recruitment. A rich variety of phenotypic and health-related information is available on each participant, making the resource unprecedented in its size and scope. Here we describe the genome-wide genotype data (~805,000 markers) collected on all individuals in the cohort and its quality control procedures. Genotype data on this scale offers novel opportunities for assessing quality issues, although the wide range of ancestries of the individuals in the cohort also creates particular challenges. We also conducted a set of analyses that reveal properties of the genetic data - such as population structure and relatedness - that can be important for downstream analyses. In addition, we phased and imputed genotypes into the dataset, using computationally efficient methods combined with the Haplotype Reference Consortium (HRC) and UK10K haplotype resource. This increases the number of testable variants by over 100-fold to ~96 million variants. We also imputed classical allelic variation at 11 human leukocyte antigen (HLA) genes, and as a quality control check of this imputation, we replicate signals of known associations between HLA alleles and many common diseases. We describe tools that allow efficient genome-wide association studies (GWAS) of multiple traits and fast phenome-wide association studies (PheWAS), which work together with a new compressed file format that has been used to distribute the dataset. As a further check of the genotyped and imputed datasets, we performed a test-case genome-wide association scan on a well-studied human trait, standing height.

genetics