bioRxiv ScienceSearch

Biology subjects

Becher, H.

Publications and source records attributed to Becher, H..

3 recordsLinked to original sources

Mitochondrial sequences or Numts - By-catch differs between sequencing methods

Nuclear inserts derived from mitochondrial DNA (Numts) encode valuable information. Being mostly non-functional, and accumulating mutations more slowly than mitochondrial sequence, they act like molecular fossils - they preserve information on the ancestral sequences of the mitochondrial DNA. In addition, changes to the Numt sequence since their insertion into the nuclear genome carry information about the nuclear phylogeny. These attributes cannot be reliably exploited if Numt sequence is confused with the mitochondrial genome (mtDNA). The analysis of mtDNA would be similarly compromised by any confusion, for example producing misleading results in DNA barcoding that used mtDNA sequence. We propose a method to distinguish Numts from mtDNA, without the need for comprehensive assembly of the nuclear genome or the physical separation of organelles and nuclei. It exploits the different biases of long and short-read sequencing. We find that short-read data yield mainly mtDNA sequences, whereas long-read sequencing strongly enriches for Numt sequences. We demonstrate the method using genome-skimming (coverage < 1x) data obtained on Illumina short-read and PacBio long-read technology from DNA extracted from six grasshopper individuals. The mitochondrial genome sequences were assembled from the short-read data despite the presence of Numts. The PacBio data contained a much higher proportion of Numt reads (over 16-fold), making us caution against the use of long-read methods for studies using mitochondrial loci. We obtained two estimates of the genomic proportion of Numts. Finally, we introduce \"tangle plots\", a way of visualising Numt structural rearrangements and comparing them between samples.

evolutionary biology

Patterns of genetic variability in genomic regions with low rates of recombination

Surveys of DNA sequence variation have shown that the level of genetic variability in a genomic region is often strongly positively correlated with its rate of crossing over (CO) [1-3]. This pattern is caused by selection acting on linked sites, which reduces genetic variability and can also cause the frequency distribution of segregating variants to contain more rare variants than expected without selection (skew). These effects of selection may involve the spread of beneficial mutations (selective sweeps, SSWs), the elimination of deleterious mutations (background selection, BGS) or both together, and are expected to be stronger with lower rates of crossing over [1-3]. However, in a recent study of human populations, the skew was reduced in the lowest CO regions compared with regions with somewhat higher CO rates [4]. A similar pattern is seen in the population genomic studies of Drosophila simulans described here. We propose an explanation for this paradoxical observation, and validate it using computer simulations. This explanation is based on the finding that partially recessive, linked deleterious mutations can increase rather than reduce neutral variability when the product of the effective population size (Ne) and the selection coefficient against homozygous carriers of mutations (s) is [&le;] 1, i.e. there is associative overdominance (AOD) rather than BGS [5]. We show that AOD can operate in a genomic region with a low rate of CO, opening up a new perspective on how selection affects patterns of variability at linked sites.

evolutionary biology

Common breast cancer risk loci predispose to distinct tumor subtypes

BackgroundGenome-wide association studies (GWAS) have identified multiple common breast cancer susceptibility variants. Many of these variants have differential associations by estrogen receptor (ER), but how these variants relate with other tumor features and intrinsic molecular subtypes is unclear. MethodsAmong 106,571 invasive breast cancer cases and 95,762 controls of European ancestry with data on 173 breast cancer variants identified in previous GWAS, we used novel two-stage polytomous logistic regression models to evaluate variants in relation to multiple tumor features (ER, progesterone receptor (PR), human epidermal growth factor receptor 2 (HER2) and grade) adjusting for each other, and to intrinsic-like subtypes. ResultsEighty-five of 173 variants were associated with at least one tumor feature (false discovery rate <5%), most commonly ER and grade, followed by PR and HER2. Models for intrinsic-like subtypes found nearly all of these variants (83 of 85) associated at P<0.05 with risk for at least one luminal-like subtype, and approximately half (41 of 85) of the variants were associated with risk of at least one non-luminal subtype, including 32 variants associated with triple-negative (TN) disease. Ten variants were associated with risk of all subtypes in different magnitude. Five variants were associated with risk of luminal A-like and TN subtypes in opposite directions. ConclusionThis report demonstrates a high level of complexity in the etiology heterogeneity of breast cancer susceptibility variants and can inform investigations of subtype-specific risk prediction.

genetics