bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Estimating the quality of eukaryotic genomes recovered from metagenomic analysis

Eukaryotes make up a large fraction of microbial biodiversity. However, the field of metagenomics has been heavily biased towards the study of just the prokaryotic fraction. This focus has driven the necessary methodological developments to enable the recovery of prokaryotic genomes from metagenomes, which has reliably yielded genomes from thousands of novel species. More recently, microbial eukaryotes have gained more attention, but there is yet to be a parallel explosion in the number of eukaryotic genomes recovered from metagenomic samples. One of the current deficiencies is the lack of a universally applicable and reliable tool for the estimation of eukaryote genome quality. To address this need, we have developed EukCC, a tool for estimating the quality of eukaryotic genomes based on the dynamic selection of single copy marker gene sets, with the aim of applying it to metagenomics datasets. We demonstrate that our method outperforms current genome quality estimators and have applied EukCC to datasets from two different biomes to enable the identification of novel genomes, including a eukaryote found on the human skin and a Bathycoccus species obtained from a marine sample.

bioinformatics↗

Photosynthetic protein classification using genome neighborhood-based machine learning feature

Identification of novel photosynthetic proteins is important for understanding and improving photosynthetic efficiency. Synergistically, genomic context such as genome neighborhood can provide additional useful information to identify the photosynthetic proteins. We, therefore, expected that applying the computational approach, particularly machine learning (ML) with the genome neighborhood-based feature should facilitate the photosynthetic function assignment. Our results revealed a functional relationship between photosynthetic genes and their genomic neighbors, indicating the possibility to assign functions from their genome neighborhood profile. Therefore, we created a new method for extracting the patterns based on genome neighborhood network (GNN) and applied for the photosynthetic protein classification using ML algorithms. Random forest (RF) classifier using genome neighborhood-based features achieved the highest accuracy up to 94% in the classification of photosynthetic proteins and also showed better performance (Mathews correlation coefficient = 0.852) than other available tools including the sequence similarity search (0.497) and ML-based method (0.512). Furthermore, we demonstrated the ability of our model to identify novel photosynthetic proteins comparing to the other methods. Our classifier is available at http://bicep.kmutt.ac.th/photomod_standalone, https://bit.ly/2S0I2Ox and DockerHub: https://hub.docker.com/r/asangphukieo/photomod

bioinformatics↗

Millipede genomes reveal unique adaptation of genes and microRNAs during myriapod evolution

The Myriapoda including millipedes and centipedes is of major importance in terrestrial ecology and nutrient recycling. Here, we sequenced and assembled two chromosomal-scale genomes of millipedes Helicorthomorpha holstii (182 Mb, N50 18.11 Mb mainly on 8 pseudomolecules) and Trigoniulus corallinus (449 Mb, N50 26.78 Mb mainly on 15 pseudomolecules). Unique defense systems, genomic features, and patterns of gene regulation in millipedes, not observed in other arthropods, are revealed. Millipedes possesses a unique ozadene defensive gland unlike the venomous forcipules in centipedes. Sets of genes associated with anti-microbial activity are identified with proteomics, suggesting that the ozadene gland is not primarily an antipredator adaptation (at least in T. corallinus). Macro-synteny analyses revealed highly conserved genomic blocks between centipede and the two millipedes. Tight Hox and the first loose ecdysozoan ParaHox homeobox clusters are identified, and a myriapod-specific genomic rearrangement including Hox3 is also observed. The Argonaute proteins for loading small RNAs are duplicated in both millipedes, but unlike insects, an argonaute duplicate has become a pseudogene. Evidence of post-transcriptional modification in small RNAs, including species-specific microRNA arm switching that provide differential gene regulation is also obtained. Millipede genomes reveal a series of unique genomic adaptations and microRNA regulation mechanisms have occurred in this major lineage of arthropod diversity. Collectively, the two millipede genomes shed new light on this fascinating but poorly understood branch of life, with a highly unusual body plan and novel adaptations to their environment.

evolutionary biology↗

Genome-wide search for parent-of-origin allele specific expression in Bombus terrestris.

Genomic imprinting is the differential expression of alleles in diploid individuals, with the expression being dependent upon the sex of the parent from which it was inherited. Haigs kinship theory hypothesizes that genomic imprinting is due to an evolutionary conflict of interest between alleles from the mother and father. In social insects, it has been suggested that genomic imprinting should be widespread. One recent study identified parent-of-origin expression in honeybees and found evidence supporting the kinship theory. However, little is known about genomic imprinting in insects and multiple theoretical predictions must be tested to avoid single-study confirmation bias. We, therefore, tested for parent-of-origin expression in a primitively eusocial bee. We found equal numbers of maternally and paternally biased expressed alleles. The most highly biased alleles were maternally expressed, offering support for the kinship theory. We also found low conservation of potentially imprinted genes with the honeybee, suggesting rapid evolution of genomic imprinting in Hymenoptera. Impact summaryGenomic imprinting is the differential expression of alleles in diploid individuals, with the expression being dependent upon the sex of the parent from which it was inherited. Genomic imprinting is an evolutionary paradox. Natural selection is expected to favour expression of both alleles in order to protect against recessive mutations that render a gene ineffective. What then is the benefit of silencing one copy of a gene, making the organism functionally haploid at that locus? Several explanations for the evolution of genomic imprinting have been proposed. Haigs kinship theory is the most developed and best supported. Haigs theory is based on the fact that maternally (matrigene) and paternally (patrigene) inherited genes in the same organism can have different interests. For example, in a species with multiple paternity, a patrigene has a lower probability of being present in siblings that are progeny of the same mother than does a matrigene. As a result, a patrigene will be selected to value the survival of the organism it is in more highly, compared to the survival of siblings. This is not the case for a matrigene. Kinship theory is central to our evolutionary understanding of imprinting effects in human health and plant breeding. Despite this, it still lacks a robust, independent test. Colonies of social bees consist of diploid females (queens and workers) and haploid males created from unfertilised eggs. This along with their social structures allows for novel predictions of Haigs theory. In this paper, we find parent of origin allele specific expression in the important pollinator, the buff-tailed bumblebee. We also find, as predicted by Haigs theory, a balanced number of genes showing matrigenic or patrigenic bias with the most extreme bias been found in matrigenically biased genes.

evolutionary biology↗

StoHi-C: Using t-Distributed Stochastic Neighbor Embedding (t-SNE) to predict 3D genome structure from Hi-C Data

In order to comprehensively understand the structure-function relationship of the genome, 3D genome structures must first be predicted from biological data (like Hi-C) using computational tools. Many of these existing tools rely partially or completely on multi-dimensional scaling (MDS) to embed predicted structures in 3D space. MDS is known to have inherent problems when applied to high-dimensional datasets like Hi-C. Alternatively, t-Distributed Stochastic Neighbor Embedding (t-SNE) is able to overcome these problems but has not been applied to predict 3D genome structures. In this manuscript, we present a new workflow called StoHi-C (pronounced "stoic") that uses t-SNE to predict 3D genome structure from Hi-C data. StoHi-C was used to predict 3D genome structures for multiple, independent existing fission yeast Hi-C datasets. Overall, StoHi-C was able to generate 3D genome structures that more clearly exhibit the established principles of fission yeast 3D genomic organization.

bioinformatics↗

Variance-stabilized units for sequencing-based genomic signals

MotivationA sequencing-based genomic assay such as ChIP-seq outputs a real-valued signal for each position in the genome that measures the strength of activity at that position. Most genomic signals lack the property of variance stabilization. That is, a difference between 100 and 200 reads usually has a very different statistical importance from a difference between 1,100 and 1,200 reads. A statistical model such as a negative binomial distribution can account for this pattern, but learning these models is computationally challenging. Therefore, many applications--including imputation and segmentation and genome annotation (SAGA)--instead use Gaussian models and use a transformation such as log or inverse hyperbolic sine (asinh) to stabilize variance. ResultsWe show here that existing transformations do not fully stabilize variance in genomic data sets. To solve this issue, we propose VSS, a method that produces variance-stabilized signals for sequencingbased genomic signals. VSS learns the empirical relationship between the mean and variance of a given signal data set and produces transformed signals that normalize for this dependence. We show that VSS successfully stabilizes variance and that doing so improves downstream applications such as SAGA. VSS will eliminate the need for downstream methods to implement complex mean-variance relationship models, and will enable genomic signals to be easily understood by eye. Contactmaxwl@sfu.ca. Availabilityhttps://github.com/faezeh-bayat/Variance-stabilized-units-for-sequencing-based-genomic-signals.

bioinformatics↗

Unified methods for feature selection in large-scale genomics data with censored survival outcomes

One of the major goals in large-scale genomic studies is to identify genes with a prognostic impact on time-to-event outcomes which provide insight into the diseases process. With rapid developments in high-throughput genomic technologies in the past two decades, the scientific community is able to monitor the expression levels of tens of thousands of genes and proteins resulting in enormous data sets where the number of genomic features is far greater than the number of subjects. Methods based on univariate Cox regression are often used to select genomic features related to survival outcome; however, the Cox model assumes proportional hazards (PH), which is unlikely to hold for each feature. When applied to genomic features exhibiting some form of non-proportional hazards (NPH), these methods could lead to an under- or over-estimation of the effects. We propose a broad array of marginal screening techniques that aid in feature ranking and selection by accommodating various forms of NPH. First, we develop an approach based on Kullback-Leibler information divergence and the Yang-Prentice model that includes methods for the PH and proportional odds (PO) models as special cases. Next, we propose R2 indices for the PH and PO models that can be interpreted in terms of explained randomness. Lastly, we propose a generalized pseudo-R2 measure that includes PH, PO, crossing hazards and crossing odds models as special cases and can be interpreted as the percentage of separability between subjects experiencing the event and not experiencing the event according to feature expression. We evaluate the performance of our measures using extensive simulation studies and publicly available data sets in cancer genomics. We demonstrate that the proposed methods successfully address the issue of NPH in genomic feature selection and outperform existing methods. The proposed information divergence, R2 and pseudo-R2 measures were implemented in R (www.R-project.org) and code is available upon request.

bioinformatics↗

Agewise mapping of genomic oxidative DNA modification demonstrates oxidative-driven reprogramming of pro-longevity genes

The accumulation of unrepaired oxidatively damaged DNA can influence both the rate of ageing and life expectancy of an organism. Mapping oxidative DNA damage sites at whole-genome scale will help us to recognize the damage-prone sequence and genomic feature information, which is fundamental for ageing research. Here, we developed an algorithm to map the whole-genome oxidative DNA damage at single-base resolution using Single-Molecule Real-Time (SMRT) sequencing technology. We sequenced the genomic oxidative DNA damage landscape of C. elegans at different age periods to decipher the potential impact of genomic DNA oxidation on physiological ageing. We observed an age-specific pattern of oxidative modification in terms of motifs, chromosomal distribution, and genomic features. Integrating with RNA-Seq data, we demonstrated that oxidative modification in promoter regions was negatively associated with the expression of pro-longevity genes, denoting that oxidative modification in pro-longevity genes may exert epigenetic potential and thus affect lifespan determination. Together, our study opens up a new field for exploration of "oxigenetics," that focuses on the mechanisms of redox-mediated ageing. SummaryO_LIWe developed an algorithm to map the oxidative DNA damage at single-base resolution. C_LIO_LIOxidative DNA damage landscape in C. elegans illustrated an age-specific pattern in terms of motifs, chromosomal distribution, and genomic features. C_LIO_LIOxidative modification in older worms occurred higher frequency at the sex chromosome, with the preference for promoter and exon regions. C_LIO_LIOxidative modification in promoter regions of pro-longevity genes was negatively associated with their expression, suggesting the oxidative-driven transcript reprogramming of pro-longevity genes in physiological ageing. C_LI

genetics↗

Multiplexed conditional genome editing with Cas12a in Drosophila

CRISPR-Cas genome engineering has revolutionised biomedical research by enabling targeted genome modification with unprecedented ease. In the popular model organism Drosophila melanogaster gene editing has so far relied exclusively on the prototypical CRISPR nuclease Cas9. The availability of additional CRISPR systems could expand the genomic target space, offer additional modes of regulation and enable the independent manipulation of genes in different cell populations of the same animal. Here we describe a platform for efficient Cas12a gene editing in Drosophila. We show that Cas12a from Lachnospiraceae bacterium, but not Acidaminococcus spec., can mediate robust gene editing in vivo. In combination with most crRNAs, LbCas12a activity is strongly suppressed at lower temperatures, enabling control of gene editing by simply modulating temperature. LbCas12a can directly utilize compact crRNAs arrays that are substantially easier to construct than Cas9 sgRNA arrays, facilitating multiplex genome engineering of several target sites in parallel. Targeting genes with arrays of three crRNAs results in the induction of loss-of function phenotypes with comparable efficiencies than a state-of-the-art Cas9 system. Lastly, we show that cell type-specific expression of LbCas12a is sufficient to mediate tightly controlled gene editing in a variety of tissues, allowing detailed analysis of gene function in this multicellular organism. Cas12a gene editing substantially expands the genome engineering toolbox in this organism and will be a powerful method for the functional annotation of the Drosophila genome. This work also lays out principles for the development of multiplexed transgenic Cas12a genome engineering systems in other genetically tractable organisms.

genetics↗

CRISPR-Csy4-mediated editing of rotavirus double-stranded RNA genome

CRISPR-nucleases have been widely applied for editing cellular and viral genomes, but nuclease-mediated genome editing of double-stranded RNA (dsRNA) viruses has not yet been reported. Here, by engineering CRISPR-Csy4 nuclease to localise to rotavirus viral factories, we achieved the first nuclease-mediated genome editing of rotavirus, an important human and livestock pathogen with a multi-segmented dsRNA genome. Rotavirus replication intermediates cleaved by Csy4 were repaired through the formation of defined deletions in the targeted genome segments in a single replication cycle. Using CRISPR-Csy4-mediated editing of rotavirus genome, we labelled for the first time the products of rotavirus secondary transcription made by newly assembled viral particles during rotavirus replication, demonstrating that this step largely contributes to the overall production of viral proteins. We anticipate that the nuclease-mediated cleavage of dsRNA virus genomes will promote a new level of understanding of viral replication and host-pathogen interactions, offering the opportunity to develop new therapeutics.

microbiology↗

Analysis procedures for assessing recovery of high quality, complete, closed genomes from Nanopore long read metagenome sequencing

New long read sequencing technologies offer huge potential for effective recovery of complete, closed genomes from complex microbial communities. Using long read (MinION) obtained from an ensemble of activated sludge enrichment bioreactors, we 1) describe new methods for validating long read assembled genomes using their counterpart short read metagenome assembled genomes; 2) assess the influence of different correction procedures on genome quality and predicted gene quality and 3) contribute 21 new closed or complete genomes of community members, including several species known to play key functional roles in wastewater bioprocesses: specifically microbes known to exhibit the polyphosphate- and glycogen-accumulating organism phenotypes (namely Accumulibacter and Dechloromonas, and Micropruina and Defluviicoccus, respectively), and filamentous bacteria (Thiothrix) associated with the formation and stability of activated sludge flocs. Our findings further establish the feasibility of long read metagenome-assembled genome recovery, and demonstrate the utility of parallel sampling of moderately complex enrichments communities for recovery of genomes of key functional species relevant for the study of complex wastewater treatment bioprocesses.

bioinformatics↗

Heterotrophic Thaumarchaeota with ultrasmall genomes are widespread in the ocean

The Thaumarchaeota comprise a diverse archaeal phylum including numerous lineages that play key roles in global biogeochemical cycling, particularly in the ocean. To date, all genomically-characterized marine Thaumarchaeota are reported to be chemolithoautotrophic ammonia-oxidizers. In this study, we report a group of heterotrophic marine Thaumarchaeota (HMT) with ultrasmall genome sizes that is globally abundant in deep ocean waters, apparently lacking the ability to oxidize ammonia. We assemble five HMT genomes from metagenomic data derived from both the Atlantic and Pacific Oceans, including two that are >95% complete, and show that they form a deeply-branching lineage sister to the ammonia-oxidizing archaea (AOA). Metagenomic read mapping demonstrates the presence of this group in mesopelagic samples from all major ocean basins, with abundances reaching up to 6% that of AOA. Surprisingly, the predicted sizes of complete HMT genomes are only 837-908 Kbp, and our ancestral state reconstruction indicates this lineage has undergone substantial genome reduction compared to other related archaea. The genomic repertoire of HMT indicates a highly reduced metabolism for aerobic heterotrophy that, although lacking the carbon fixation pathway typical of AOA, includes a divergent form III-a RuBisCO that potentially functions in a nucleotide scavenging pathway. Despite the small genome size of this group, we identify 13 encoded pyrroloquinoline quinone (PQQ)-dependent dehydrogenases that are predicted to shuttle reducing equivalents to the electron transport chain, suggesting these enzymes play an important role in the physiology of this group. Our results suggest that heterotrophic Thaumarchaeota are widespread in the ocean and potentially play key roles in global chemical transformations. ImportanceIt has been known for many years that marine Thaumarchaeota are abundant constituents of dark ocean microbial communities, where their ability to couple ammonia oxidation and carbon fixation plays a critical role in nutrient dynamics. In this study we describe an abundant group of heterotrophic marine Thaumarchaeota (HMT) in the ocean with physiology distinct from their ammonia-oxidizing relatives. HMT lack the ability to oxidize ammonia and fix carbon via the 3-hydroxypropionate/4-hydroxybutyrate pathway, but instead encode a form III-a RuBisCO and diverse PQQ-dependent dehydrogenases that are likely used to generate energy in the dark ocean. Our work expands the scope of known diversity of Thaumarchaeota in the ocean and provides important insight into a widespread marine lineage.

microbiology↗

HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata

We present HO_SCPLOWARVESTMANC_SCPLOW, a method that takes advantage of hierarchical relationships among the possible biological interpretations and representations of genomic variants to perform automatic feature learning, feature selection, and model building. We demonstrate that HO_SCPLOWARVESTMANC_SCPLOW scales to thousands of genomes comprising more than 84 million variants by processing phase 3 data from the 1000 Genomes Project, the largest publicly available collection of whole genome sequences. Next, using breast cancer data from The Cancer Genome Atlas, we show that HO_SCPLOWARVESTMANC_SCPLOW selects a rich combination of representations that are adapted to the learning task, and performs better than a binary representation of SNPs alone. Finally, we compare HO_SCPLOWARVESTMANC_SCPLOW to existing feature selection methods and demonstrate that our method selects smaller and less redundant feature subsets, while maintaining accuracy of the resulting classifier. The data used is available through either the 1000 Genomes Project or The Cancer Genome Atlas. Access to TCGA data requires the completion of a Data Access Request through the Database of Genotypes and Phenotypes (dbGaP). Binary releases of HO_SCPLOWARVESTMANC_SCPLOW compatible with Linux, Windows, and Mac are available for download at https://github.com/cmlh-gp/Harvestman-public/releases

bioinformatics↗

Discovery of a role for Rab3b in habituation and cocaine induced locomotor activation in mice using heterogeneous functional genomic analysis

Substance use disorders are prevalent and present a tremendous societal cost but the mechanisms underlying addiction behavior are poorly understood and few biological treatments exist. One strategy to identify novel molecular mechanisms of addiction is through functional genomic experimentation. However, results from individual experiments are often noisy. To address this problem, the convergent analysis of multiple genomic experiments can prioritize signal from these studies. In the present study, we examine genetic loci identified in the recombinant inbred (BXD RI) genetic reference population that modulate the locomotor response to cocaine. We then applied the GeneWeaver software system for heterogeneous functional genomic analysis to integrate and aggregate multiple studies of addiction genomics, resulting in the identification of Rab3b, as a functional correlate of the locomotor response to cocaine in rodents. This gene encodes a member of the RAB family of Ras-like GTPases known to be involved in trafficking of secretory and endocytic vesicles in eukaryotic cells. The convergent evidence for a role of Rab3b was included co-occurrence in previously published genetic mapping studies of cocaine related behaviors; methamphetamine response and Cartpt (Cocaine- and amphetamine-regulated transcript prepropeptide) abundance; evidence related to other addictive substances; density of polymorphisms; and its expression pattern in reward pathways. To evaluate this finding, we examined the effect of RAB3 complex perturbation in cocaine response. B6;129-Rab3btm1Sud Rab3ctm1sud Rab3dtm1sud triple null mice (Rab3bcd-/-) exhibited significant deficits in habituation, and increased acute and repeated cocaine responses. This previously unidentified mechanism of the behavioral predisposition and response to cocaine is an example of many that can be identified and validated using aggregate genomic studies. Many genetic and genomic studies have been performed over the past few decades, representing a wealth of data on the underlying neurobiological and genetic basis of multiple complex behaviors. However, these studies, particularly legacy studies using older technologies and resources lack precision. By aggregating multiple studies, convergent evidence for shared molecular mechanisms of multiple behaviors can be found, for example the widely reported relations among psychostimulant use and novelty response behavior. Here a legacy genetic mapping result for a cocaine related trait mapped in mice was refined using data from 113 different experimental gene sets related to addiction in the GeneWeaver system for heterogeneous functional genomic analysis. Convergent evidence revealed a role for Rab3b in this and other traits including multiple psychostimulant responses and CART expression. Experimental perturbation of the RAB complex revealed effects on habituation to a novel environment, cocaine induced activation and Carpt expression. The analysis of aggregate data thus revealed a molecular mechanism that influences the relationship between response to novel situations and cocaine-related phenotypes.

neuroscience↗

Host association induces genome changes in Candida albicans which alters its virulence

Candida albicans is an opportunistic fungal pathogen of humans that is typically diploid yet, has a highly labile genome that is tolerant of large-scale perturbations including chromosomal aneuploidy and loss-of-heterozygosity events. The ability to rapidly generate genetic variation is crucial for C. albicans to adapt to changing or stress environments, like those encountered in the host. Genetic variation occurs via stress-induced mutagenesis or can be generated through its parasexual cycle, which includes mating between diploids or stress-induced mitotic defects to produce tetraploids and non-meiotic ploidy reduction. However, it remains largely unknown how genetic background contributes to C. albicans genome instability in vitro or in vivo. Here, we tested how genetic background, ploidy and host environment impact C. albicans genome stability. We found that host association induced both loss-of-heterozygosity events and genome size changes, regardless of genetic background or ploidy. However, the magnitude and types of genome changes varied across C. albicans strains. We also assessed whether host-induced genomic changes resulted in any consequences on growth rate and virulence phenotypes and found that many host derived isolates had significant changes compared to their parental strains. Interestingly, host derivatives from diploid C. albicans predominantly displayed increased virulence, whereas host derivatives from tetraploid C. albicans had mostly reduced virulence. Together, these results are important for understanding how host-induced genomic changes in C. albicans alter the relationship between the host and C. albicans.

microbiology↗

Transposable Elements activity and role in Meloidogyne incognita genome dynamic and adaptability

AO_SCPLOWBSTRACTC_SCPLOWDespite reproducing without sexual recombination, the root-knot nematode Meloidogyne incognita is adaptive and versatile. Indeed, this species displays a global distribution, is able to parasitize a large range of plants and can overcome plant resistance in a few generations. The mechanisms underlying this adaptability without sex remain poorly known and only low variation at the single nucleotide polymorphism level have been observed so far across different geographical isolates with distinct ranges of compatible hosts. Hence, other mechanisms than the accumulation of point mutations are probably involved in the genomic dynamics and plasticity necessary for adaptability. Transposable elements (TEs), by their repetitive nature and mobility, can passively and actively impact the genome dynamics. This is particularly expected in polyploid hybrid genomes such as the one of M. incognita. Here, we have annotated the TE content of M. incognita, analyzed the statistical properties of this TE content, and used population genomics approach to estimate the mobility of these TEs across 12 geographical isolates, presenting phenotypic variations. The TE content is more abundant in DNA transposons and the distribution of TE copies identity to their consensuses sequence suggests they have been at least recently active. We have identified loci in the genome where the frequencies of presence of a TE showed variations across the different isolates. Compared to the M. incognita reference genome, we detected the insertion of some TEs either within genic regions or in the upstream regulatory regions. These predicted TEs insertions might thus have a functional impact. We validated by PCR the insertion of some of these TEs, confirming TE movements probably play a role in the genome plasticity with possible functional impacts.

evolutionary biology↗

Haplocheck: Phylogeny-based Contamination Detection in Mitochondrial and Whole-Genome Sequencing Studies

Within-species contamination is a major issue in sequencing studies, especially for mitochondrial studies. Contamination can be detected by analysing the nuclear genome or by inspecting the heteroplasmic sites in the mitochondrial genome. Existing methods using the nuclear genome are computationally expensive, and no suitable tool for detecting contamination in large-scale mitochondrial datasets is available. Here we present haplocheck, a tool that requires only the mitochondrial genome to detect contamination in both mitochondrial and whole-genome sequencing studies. Haplocheck is able to distinguish between contaminated and real heteroplasmic sites using the mitochondrial phylogeny. By applying haplocheck to the 1000 Genomes Project data, we show (1) high concordance in contamination estimates between mitochondrial and nuclear DNA and (2) quantify the impact of mitochondrial copy numbers on the mitochondrial based contamination results. Haplocheck complements leading nuclear DNA based contamination tools, and can therefore be used as a proxy tool in nuclear genome studies. Haplocheck is available both as a command-line tool at https://github.com/genepi/haplocheck and as a cloud web-service producing interactive reports that facilitates the navigation through the phylogeny of contaminated samples.

bioinformatics↗

Discordant evolution of organellar genomes in peas (Pisum L.)

Plastids and mitochondria have their own small genomes which do not undergo meiotic recombination and may have evolutionary fate different from each other and nuclear genome, thus highlighting interesting phenomena in plant evolution. We for the first time sequenced mitochondrial genomes of pea (Pisum L.), in 38 accessions mostly representing diverse wild germplasm from all over pea geographical range. Six structural types of pea mitochondrial genome were revealed. From the same accessions, plastid genomes were sequenced. Bayesian phylogenetic trees based on the plastid and mitochondrial genomes were compared. The topologies of these trees were highly discordant implying not less than six events of hybridisation of diverged wild peas in the past, with plastids and mitochondria differently inherited by the descendants. Such discordant inheritance of organelles is supposed to have been driven by plastid-nuclear incompatibility, known to be widespread in pea wide crosses and apparently shaping the organellar phylogenies. The topology of a phylogenetic tree based on the nucleotide sequence of a nuclear gene His5 coding for a histone H1 subtype corresponds to the current taxonomy and resembles that based on the plastid genome. Wild peas (Pisum sativum subsp. elatius s.l.) inhabiting Southern Europe were shown to be of hybrid origin resulting from crosses of peas similar to those presently inhabiting south-east and north-east Mediterranean in broad sense.

evolutionary biology↗