bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

CRISPR/Cas9 Targeted Capture Of Mammalian Genomic Regions For Characterization By NGS

The robust detection of structural variants in mammalian genomes remains a challenge. It is particularly difficult in the case of genetically unstable Chinese hamster ovary (CHO) cell lines with only draft genome assemblies available. We explore the potential of the CRISPR/Cas9 system for the targeted capture of genomic loci containing integrated vectors in CHO-K1-based cell lines, and compare it to popular target-enrichment methods and to whole genome sequencing (WGS). The CRISPR/Cas9-based techniques allow for amplification-free capture of genomic regions, which reduces the possibility of sequencing artifacts. Other advantages of these methods are the ease of bioinformatics analysis, potential for multiplexing, and the production of longer sequencing templates for real-time sequencing. The utility of these protocols has been proven by identification of transgene integration sites and flanking sequences in a number of CHO cell lines. However, data produced by these and other targeted capture methods are not always sufficient to analyze complex genomic rearrangements (CGRs) or unexpected sequences introduced into genome by vector integration events. In contrast, WGS provides complete information about vector integration sites, vector copy number, CGRs, and foreign DNA-but despite these benefits, WGS is not easily implemented due to the cost and complexity of the analysis.

genomics

An open source platform for analyzing and sharing worm behavior data

Animal behavior is increasingly being recorded in systematic imaging studies that generate large data sets. To maximize the usefulness of these data there is a need for improved resources for analyzing and sharing behavior data that will encourage re-analysis and method development by computational scientists1. However, unlike genomic or protein structural data, there are no widely used standards for behavior data. It is therefore desirable to make the data available in a relatively raw form so that different investigators can use their own representations and derive their own features. For computational ethology to approach the level of maturity of other areas of bioinformatics, we need to address at least three challenges: storing and accessing video files, defining flexible data formats to facilitate data sharing, and making software to read, write, browse, and analyze the data. We have developed an open resource to begin addressing these challenges using worm tracking as a model.

animal behavior and cognition

Initial Characterization of the Two ClpP Paralogs of Chlamydia trachomatis Suggests Unique Functionality for Each

Chlamydia is an obligate intracellular bacterium that differentiates between two distinct functional and morphological forms during its developmental cycle: elementary bodies (EBs) and reticulate bodies (RBs). EBs are non-dividing, small electron dense forms that infect host cells. RBs are larger, non-infectious replicative forms that develop within a membrane-bound vesicle, termed an inclusion. Given the unique properties of each developmental form of this bacterium, we hypothesized that the Clp protease system plays an integral role in proteomic turnover by degrading specific proteins from one developmental form or the other. Chlamydia has five uncharacterized clp genes: clpX, clpC, two clpP paralogs, and clpB. In other bacteria, ClpC and ClpX are ATPases that unfold and feed proteins into the ClpP protease to be degraded, and ClpB is a deaggregase. Here, we focused on characterizing the ClpP paralogs. Transcriptional analyses and immunoblotting determined these genes are expressed mid-cycle. Bioinformatic analyses of these proteins identified key residues important for activity. Over-expression of inactive clpP mutants in Chlamydia suggested independent function of each ClpP paralog. To further probe these differences, we determined interactions between the ClpP proteins using bacterial two-hybrid assays and native gel analysis of recombinant proteins. Homotypic interactions of the ClpP proteins, but not heterotypic interactions between the ClpP paralogs, were detected. Interestingly, ClpP2, but not ClpP1, protease activity was detected in vitro. This activity was stimulated by antibiotics known to activate ClpP, which also blocked chlamydial growth. Our data suggest the chlamydial ClpP paralogs likely serve distinct and critical roles in this important pathogen.\n\nImportanceChlamydia trachomatis is the leading cause of preventable infectious blindness and of bacterial sexually transmitted infections worldwide. Chlamydiae are developmentally regulated, obligate intracellular pathogens that alternate between two functional and morphologic forms with distinct repertoires of proteins. We hypothesize that protein degradation is a critical aspect to the developmental cycle. A key system involved in protein turnover in bacteria is the Clp protease system. Here, we characterized the two chlamydial ClpP paralogs by examining their expression in Chlamydia, their ability to oligomerize, and their proteolytic activity. This work will help understand the evolutionarily diverse Clp proteases in the context of intracellular organisms, which may aid in the study of other clinically relevant intracellular bacteria.

microbiology

Continuous State HMMs for Modeling Time Series Single Cell RNA-Seq Data

MotivationMethods for reconstructing developmental trajectories from time series single cell RNA-Seq (scRNA-Seq) data can be largely divided into two categories. The first, often referred to as pseudotime ordering methods, are deterministic and rely on dimensionality reduction followed by an ordering step. The second learns a probabilistic branching model to represent the developmental process. While both types have been successful, each suffers from shortcomings that can impact their accuracy.\n\nResultsWe developed a new method based on continuous state HMMs (CSHMMs) for representing and modeling time series scRNA-Seq data. We define the CSHMM model and provide efficient learning and inference algorithms which allow the method to determine both the structure of the branching process and the assignment of cells to these branches. Analyzing several developmental single cell datasets we show that the CSHMM method accurately infers branching topology and correctly and continuously assign cells to paths, improving upon prior methods proposed for this task. Analysis of genes based on the continuous cell assignment identifies known and novel markers for different cell types.\n\nAvailabilitySoftware and Supporting website: www.andrew.cmu.edu/user/chiehll/CSHMM/\n\nContactzivbj@cs.cmu.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Quick and efficient approach to develop genomic resources in orphan species: application in Lavandula angustifolia

Next-Generation Sequencing (NGS) technologies, by reducing the cost and increasing the throughput of sequencing, have opened doors of research efforts to generate genomic data to a range of previously poorly studied species. In this study, we proposed a method for the rapid development of a large scale molecular resources for orphan species. We studied as an example Lavandula angustifolia, a perennial sub-shrub plant native from the Mediterranean region and whose essential oil have numerous applications in cosmetics, pharmaceuticals, and alternative medicines.\n\nWe first built a Maillette reference Unigene, compound of coding sequences, thanks to de novo RNA-seq assembly. Then, we reconstructed the complete genes sequences (with exons and introns) using a transcriptome-guided DNA-seq assembly approach in order to maximize the possibilities of finding polymorphism between genetically close individuals. Finally, we used these resources for SNP mining within a collection of 16 lavender clones and tested the SNP within the scope of a phylogeny analysis. We obtained a cleaned reference of 8, 030 functionally annotated genes (in silico annotation). We found up to 400K polymorphic sites, depending on the genotype analyzed, and observed a high SNP frequency (mean of 1 SNP per 90 bp) and a high level of heterozygosity (more than 60% of heterozygous SNP per genotype). We found similar genetic distances between pairs of clones, related to the out-crossing nature of the species, the restricted area of cultivation and the clonal propagation of the varieties.\n\nThe method propose is transferable to other orphan species, requires little bioinformatics resources and can be realized within a year. This is the first reported large-scale SNP development on Lavandula angustifolia. All this data provides a rich pool of molecular resource to explore and exploit biodiversity in breeding programs.

genomics

Genome-wide identification and functional analysis of circRNAs in Zea mays

Circular RNAs (circRNAs) are a class of endogenous noncoding RNAs, which increasingly drawn researchers attention in recent years as their importance in regulating gene expression at the transcriptional and post-transcriptional levels. With the development of high-throughput sequencing and bioinformatics, circRNAs have been widely analysed in animals, but the understanding of characteristics and function of circRNAs is limited in plants, especially in maize. Here, 3715 unique circRNAs were predicted in Zea mays systematically, and 8 of 12 circRNAs were validated by experiments. By analysing circRNA sequence, the events of alternative circularization phenomenon were found prevailed in maize. By comparing circRNAs in different species, it showed that part circRNAs are conserved across species, for example, there are 273 circRNAs conserved between maize and rice. Although most of the circRNAs have low expression levels, we found 213 differential expressed circRNAs responding to heat, cold, or drought, and 1782 tissue-specific expressed circRNAs. The results showed that those circRNAs may have potential biological functions in specific situations. Finally, two different methods were used to search circRNA functions, which were based on circRNAs originated from protein-coding genes and circRNAs as miRNA decoys. 346 circRNAs could act as miRNA decoys, which might modulate the effects of multiple molecular functions, including binding, catalytic activity, oxidoreductase activity, and transmembrane transporter activity. Maize circRNAs were identified, classified and characterized systematically. We also explored circRNA functions, suggesting that circRNAs are involved in multiple molecular processes and play important roles in regulating of gene expression. Our results provide a rich resource for further study of maize circRNAs.

genetics

Etiology of fever in Ugandan children: identification of microbial pathogens using metagenomic next-generation sequencing and IDseq, a platform for unbiased metagenomic analysis

BackgroundFebrile illness is a major burden in African children, and non-malarial causes of fever are uncertain. We built and employed IDseq, a cloud-based, open access, bioinformatics platform and service to identify microbes from metagenomic next-generation sequencing of tissue samples. In this pilot study, we evaluated blood, nasopharyngeal, and stool specimens from 94 children (aged 2-54 months) with febrile illness admitted to Tororo District Hospital, Uganda.\n\nResultsThe most common pathogens identified were Plasmodium falciparum (51.1% of samples) and parvovirus B19 (4.4%) from blood; human rhinoviruses A and C (40%), respiratory syncytial virus (10%), and human herpesvirus 5 (10%) from nasopharyngeal swabs; and rotavirus A (50% of those with diarrhea) from stool. Among other potential pathogens, we identified one novel orthobunyavirus, tentatively named Nyangole virus, from the blood of a child diagnosed with malaria and pneumonia, and Bwamba orthobunyavirus in the nasopharynx of a child with rash and sepsis. We also identified two novel human rhinovirus C species.\n\nConclusionsThis exploratory pilot study demonstrates the utility of mNGS and the IDseq platform for defining the molecular landscape of febrile infectious diseases in resource limited areas. These methods, supported by a robust data analysis and sharing platform, offer a new tool for the surveillance, diagnosis, and ultimately treatment and prevention of infectious diseases.

genomics

Common ancestry of heterodimerizing TALE homeobox transcription factors across Metazoa and Archaeplastida

Homeobox transcription factors (TFs) in the TALE superclass are deeply embedded in the gene regulatory networks that orchestrate embryogenesis. Knotted-like homeobox (KNOX) TFs, homologous to animal MEIS, have been found to drive the haploid-to-diploid transition in both unicellular green algae and land plants via heterodimerization with other TALE superclass TFs, representing remarkable functional conservation of a developmental TF across lineages that diverged one billion years ago. To delineate the ancestry of TALE-TALE heterodimerization, we analyzed TALE endowment in the algal radiations of Archaeplastida, ancestral to land plants. Homeodomain phylogeny and bioinformatics analysis partitioned TALEs into two broad groups, KNOX and non-KNOX. Each group shares previously defined heterodimerization domains, plant KNOX-homology in the KNOX group and animal PBC-homology in the non-KNOX group, indicating their deep ancestry. Protein-protein interaction experiments showed that the TALEs in the two groups all participated in heterodimerization. These results indicate that the TF dyads consisting of KNOX/MEIS and PBC-containing TALEs must have evolved early in eukaryotic evolution, a likely function being to accurately execute the haploid-to-diploid transitions during sexual development.\n\nAuthor summaryComplex multicellularity requires elaborate developmental mechanisms, often based on the versatility of heterodimeric transcription factor (TF) interactions. Highly conserved TALE-superclass homeobox TF networks in major eukaryotic lineages suggest deep ancestry of developmental mechanisms. Our results support the hypothesis that in early eukaryotes, the TALE heterodimeric configuration provided transcription-on switches via dimerization-dependent subcellular localization, ensuring execution of the haploid-to-diploid transition only when the gamete fusion is correctly executed between appropriate partner gametes, a system that then diversified in the several lineages that engage in complex multicellular organization.

evolutionary biology

Flux balance analysis predicts NADP phosphatase and NADH kinase are critical to balancing redox during xylose fermentation in Scheffersomyces stipitis

Xylose is the second most abundant sugar in lignocellulose and can be used as a feedstock for next-generation biofuels by industry. Saccharomyces cerevisiae, one of the main workhorses in biotechnology, is unable to metabolize xylose natively but has been engineered to ferment xylose to ethanol with the xylose reductase (XR) and xylitol dehydrogenase (XDH) genes from Scheffersoymces stipitis. In the scientific literature, the yield and volumetric productivity of xylose fermentation to ethanol in engineered S. cerevisiae still lags S. stipitis, despite expressing of the same XR-XDH genes. These contrasting phenotypes can be due to differences in S. cerevisiaes redox metabolism that hinders xylose fermentation, differences in S. stipitis redox metabolism that promotes xylose fermentation, or both. To help elucidate how S. stipitis ferments xylose, we used flux balance analysis to test various redox balancing mechanisms, reviewed published omics datasets, and studied the phylogeny of key genes in xylose fermentation. In vivo and in silico xylose fermentation cannot be reconciled without NADP phosphatase (NADPase) and NADH kinase. We identified eight candidate genes for NADPase. PHO3.2 was the sole candidate showing evidence of expression during xylose fermentation. Pho3.2p and Pho3p, a recent paralog, were purified and characterized for their substrate preferences. Only Pho3.2p was found to have NADPase activity. Both NADPase and NAD(P)H-dependent XR emerged from recent duplications in a common ancestor of Scheffersoymces and Spathaspora to enable efficient xylose fermentation to ethanol. This study demonstrates the advantages of using metabolic simulations, omics data, bioinformatics, and enzymology to reverse engineer metabolism.

microbiology

Integrative analysis of transcriptomic and clinical data uncovers the tumor suppressive activity of MITF in prostate cancer.

The dysregulation of gene expression is an enabling hallmark of cancer. Computational analysis of transcriptomics data from human cancer specimens, complemented with exhaustive clinical annotation, provides an opportunity to identify core regulators of the tumorigenic process. Here we exploit well-annotated clinical datasets of prostate cancer for the discovery of transcriptional regulators relevant to prostate cancer. Following this rationale, we identify Microphthalmia-associated transcription factor (MITF) as a prostate tumor suppressor among a subset of transcription factors. Importantly, we further interrogate transcriptomics and clinical data to refine MITF perturbation-based empirical assays and unveil Crystallin Alpha B (CRYAB) as an unprecedented direct target of the transcription factor that is, at least in part, responsible for its tumor suppressive activity in prostate cancer. This evidence was supported by the enhanced prognostic potential of a signature based on the concomitant alteration of MITF and CRYAB in prostate cancer patients. In sum, our study provides proof-of-concept evidence of the potential of the bioinformatics screen of publicly available cancer patient databases as discovery platforms, and demonstrates that the MITF-CRYAB axis controls prostate cancer biology.

cancer biology

Therapeutic effects of Hypoxia-Inducible Factor-1α (HIF-1α) on bone formation around implants in diabetic mice

Patients with uncontrolled diabetes are susceptible to implant failure due to impaired bone metabolism. Hypoxia-Inducible Factor 1 (HIF-1), a transcription factor that is up-regulated in response to reduced oxygen condition during the bone repair process after fracture or osteotomy, is known to mediate angiogenesis and osteogenesis. However, its function is inhibited under hyperglycemic conditions in diabetic patients. The aim of this study is to evaluate the effects of exogenous HIF-1 on bone formation around implants by applying HIF-1 to diabetic mice via a novel PTD-mediated DNA delivery system. Smooth surface implants (1mm in diameter; 2mm in length) were placed in the both femurs of diabetic and normal mice. HIF-1 and placebo gels were injected to implant sites of the right and left femurs, respectively: Normal mouse with HIF-1 gel (NH), Normal mouse with placebo gel (NP), Diabetic mouse with HIF-1 gel (DH), and Diabetic mouse with placebo gel (DP). RNA sequencing was performed 4 days after surgery. Based on RNA sequencing, Differentially Expressed Genes (DEGs) were identified and HIF-1 target genes were selected. Histologic and histomorphometric results were evaluated 2 weeks after the surgery. The results showed that bone-to-implant contact (BIC) and bone volume (BV) were significantly greater in the DH group than the DP group (p < 0.05). A total of 216 genes were differentially expressed in DH group compared to DP group. On the other hand, there were 95 DEGs in the case of normal mice. Twenty-one target genes of HIF-1 were identified in diabetic mice through bioinformatic analysis of DEGs. Among the target genes, NOS2, GPNMB, CCL2, CCL5, CXCL16 and TRIM63 were manually found to be associated with wound healing-related genes. In conclusion, local administration of HIF-1 via PTD may help bone formation around the implant and induce gene expression more favorable to bone formation in diabetic mice.

molecular biology

The core genome m5C methyltransferase JHP1050 (M.Hpy99III) plays an important role in orchestrating gene expression in Helicobacter pylori

Helicobacter pylori encodes a large number of Restriction-Modification (R-M) systems despite its small genome.R-M systems have been described as \"primitive immune systems\" in bacteria, but the role of methylation in bacterial gene regulation and other processes is increasingly accepted. Every H.pylori strain harbours a unique set of R-M systems resulting in a highly diverse methylome. We identified a highly conserved GCGC-specific m5C MTase (JHP1050) that was predicted to be active in all of 459 H.pylori genome sequences analyzed. Transcriptome analysis of two H.pylori strains and their respective MTase mutants showed that inactivation of the MTase led to changes in the expression of 225 genes in strain J99, and 29 genes in strain BCM-300.10 genes were differentially expressed in both mutated strains. Combining bioinformatic analysis and site-directed mutagenesis, we demonstrated that motifs overlapping the promoter influence the expression of genes directly, while methylation of other motifs might cause secondary effects.Thus, m5C methylation modifies the transcription of multiple genes, affecting important phenotypic traits that include adherence to host cells, natural competence for DNA uptake, bacterial cell shape, and susceptibility to copper.

microbiology

Association of Prevotella enterotype with polysomnographic data in obstructive sleep apnea/hypopnea syndrome patients

Intermittent hypoxia and sleep fragmentation are critical pathophysiological processes involved in obstructive sleep apnea/hypopnea syndrome (OSAHS). These manifestation independently affect similar brain regions and contribute to OSAHS-related comorbidities that are known to be related to the host gut alteration microbiota. We hypothesized that microbiota disruption influences the pathophysiological processes of OSAHS through a microbiota-gut-brain axis. Thus, we aim to survey enterotypes and polysomnographic data of OSAHS patients. Subjects were diagnosed by polysomnography, from whom fecal samples were obtained and analyzed for the microbiome composition by variable regions 3-4 of 16S rRNA pyrosequencing and bioinformatic analyses. We examined blood cytokines level of all subjects. Three enterotypes Bacteroides (n=73), Ruminococcus (n=14), and Prevotella (n=26) were identified. Central apnea indices, mixed apnea indices, N1 sleep stage, mean apnea-hypopnea duration, and arousal indices were increased in apnea-hypopnea indices (AHI) [&ge;]15 patients with the Prevotella enterotype. However, for AHI<15 subjects, obstructive apnea indices and systolic blood pressure were significantly observed in Ruminococcus and Prevotella enterotypes, respectively. The present study indicates the possibility of pathophysiological interplay between enterotypes and sleep structure disruption in sleep apnea through a microbiota-gut-brain axis and offers some new insight toward the pathogenesis of OSAHS.\n\nImportanceIntermittent hypoxia (IH) and sleep fragmentation (SF) are hallmarks of are the predominant mechanism underlying obstructive sleep apnea/hypopnea syndrome (OSAHS). Moreover, IH and SF of pathophysiological roles in the gut microbiota dysbiosis in OSAHS have been demonstrated. We hypothesized that gut microbiota disruption may cross-talk the brain function via microbiota-gut-brain axis. Indeed, we observed central apnea indices and other parameters of disturbances during sleep were significantly elevated in AHI[&ge;]15 patients with the Prevotella enterotype. This enterotype prone to endotoxin production, driving systemic inflammation, ultimately contributes to OSAHS-linked comorbidities. Vice versa, increasing the arousal index leads to systemic inflammatory changes and accompanies metabolic dysfunction. We highlight that the possibility that the microbiota-gut-brain axis operates a bidirectional effect on the development of OSAHS pathology.

microbiology

Circular RNAs regulate cancer stem cells by FMRP against CCAR1 complex in hepatocellular carcinoma

Circular RNA (circRNA) possesses great pre-clinical diagnostic and therapeutic potentials in multiple cancers. However, the underlying correlation between circRNAs and cancer stem cells (CSCs) has not been reported. The absence of circZKSCAN1 endowed several malignant properties including cancer stemness and tightly correlated with worse overall and recurrence-free survival rate in HCC cells in vitro and in vivo. Bioinformatics analysis and RNA immunoprecipitation-sequencing (RIP-seq) results revealed that circZKSCAN1 exerted its inhibitive role by competitively binding FMRP, therefore, block the binding between FMRP and {beta}-catenin-binding protein-cell cycle and apoptosis regulator 1 (CCAR1) mRNA, and subsequently restraining the transcriptional activity of Wnt signaling. In addition, RNA-splicing protein Quaking 5 was found downregulated in HCC tissues and responsible for the reduction of circZKSCAN1. Collectively, this study revealed the mechanisms underlying the regulatory role of circZKSCAN1 in HCC CSCs and identified the newly discovered Qki5- circZKSCAN1-FMRP-CCAR1-Wnt signaling axis as a potentially important therapeutic target for HCC treatment.\n\nStatement of significanceO_LICircZKSCAN1, a novel negative regulator for cancer stem cells, was firstly identified with reverse correlation with HCC prognosis.\nC_LIO_LICircZKSCAN1 directly targets FMRP, and competitive binding with {beta}-catenin-binding protein cell cycle and apoptosis regulator 1 (CCAR1) for its activity.\nC_LIO_LIThe decreased expression of Quaking 5 is responsible for the absence of circZKSCAN1 in HCC.\nC_LI

cancer biology

Topokaryotyping demonstrates single cell variability and stress dependent variations in nuclear envelope associated domains

Analysis of large-scale interphase genome positioning with reference to a nuclear landmark has recently been studied using sequencing-based single cell approaches. However, these approaches are dependent upon technically challenging, time consuming and costly high throughput sequencing technologies, requiring specialized bioinformatics tools and expertise. Here, we propose a novel, affordable and robust microscopy-based single cell approach, termed Topokaryotyping, to analyze and reconstruct the interphase positioning of genomic loci relative to a given nuclear landmark, detectable as banding pattern on mitotic chromosomes. This is accomplished by proximity-dependent histone labeling, where biotin ligase BirA fused to nuclear envelope marker Emerin was coexpressed together with Biotin Acceptor Peptide (BAP)-histone fusion followed by (i) biotin labeling, (ii) generation of mitotic spreads, (iii) detection of the biotin label on mitotic chromosomes and (iv) their identification by karyotyping. Using Topokaryotyping, we identified both cooperativity and stochasticity in the positioning of emerin-associated chromatin domains in individual cells. Furthermore, the chromosome-banding pattern showed dynamic changes in emerin-associated domains upon physical and radiological stress. In summary, Topokaryotyping is a sensitive and reliable technique to quantitatively analyze spatial positioning of genomic regions interacting with a given nuclear landmark at the single cell level in various experimental conditions.

cell biology

Pervasive contaminations in sequencing experiments are a major source of false genetic variability: a Mycobacterium tuberculosis meta-analysis

Contaminant DNA is a well-known confounding factor in molecular biology and in genomic repositories. Strikingly, analysis workflows for whole-genome sequencing (WGS) data usually neglect the errors introduced by potential contaminations. We performed a comprehensive evaluation of the extent and impact of contaminant DNA in WGS by analyzing more than 4,000 bacterial samples from 20 different studies. We found that contaminations are pervasive and can introduce large biases in variant analysis. We showed that these biases can translate in hundreds of false positive and negative SNPs, even for samples with slight contaminations. Studies investigating complex biological traits from sequencing data can be completely biased if contaminations are neglected during the bioinformatic analysis. We used both real and simulated data to evaluate and implement reliable, contamination-aware analysis pipelines. Our results urge for the implementation of such pipelines as sequencing technologies consolidate as a precision tool in the research and clinical context.

genomics

Chiral DNA sequences as commutable reference standards for clinical genomics

Chirality is a geometric property describing any object that is inequivalent to a mirror image of itself. Due to its 5-3 directionality, a DNA sequence is distinct from a mirrored sequence arranged in reverse nucleotide order, and is therefore chiral. A given sequence and its opposing chiral partner sequence share many properties, such as nucleotide composition and sequence entropy. Here we demonstrate that chiral DNA sequence pairs also perform equivalently during molecular and bioinformatic techniques that underpin modern genetic analysis, including PCR amplification, hybridization, whole-genome, target-enriched and nanopore sequencing, sequence alignment and variant detection. Given these shared properties, synthetic DNA sequences that directly mirror clinically relevant and/or analytically challenging regions of the human genome are ideal reference standards for clinical genomics. We show how the addition of chiral DNA standards to patient tumor samples can prevent false-positive and false-negative mutation detection and, thereby, improve diagnosis. Accordingly, we propose that chiral DNA standards can fulfill the unmet need for commutable internal reference standards in precision medicine.

genomics

Eps8 is a convergence point integrating EGFR and integrin trafficking and crosstalk

Crosstalk between adhesion and growth factor receptors plays a critical role in tissue morphogenesis and repair, and aberrations contribute substantially to neoplastic disease. However, the mechanisms by which adhesion and growth factor receptor signalling are integrated, spatially and temporally, are unclear.\n\nWe used adhesion complex enrichment coupled with quantitative proteomic analysis to identify rapid changes to adhesion complex composition and signalling following growth factor stimulation. Bioinformatic network and ontological analyses revealed a substantial decrease in the abundance of adhesion regulatory proteins and co-ordinators of endocytosis within 5 minutes of EGF stimulation. Together these data suggested a mechanism of EGF-induced receptor endocytosis and adhesion complex turnover.\n\nCombinatorial interrogation of the networks allowed a global and dynamic view of adhesion and growth factor receptor crosstalk to be assembled. By interrogating network topology we identified Eps8 as a putative node integrating 5{beta}1 integrin and EGFR functions. Importantly, EGF stimulation promoted internalisation of both 5{beta}1 and EGFR. However, perturbation of Eps8 increased constitutive internalisation of 5{beta}1 and EGFR; suggesting that Eps8 constrains 5{beta}1 and EGFR endocytosis in the absence of EGF stimulation. Consistent with this, Eps8 regulated Rab5 activity and was required for maintenance of adhesion complex organisation and for EGF-dependent adhesion complex disassembly. Thus, by co-ordinating 5{beta}1 and EGFR trafficking mechanisms, Eps8 is able to control adhesion receptor and growth factor receptor bioavailability and cellular contractility.\n\nWe propose that during tissue morphogenesis and repair, Eps8 functions to spatially and temporally constrain endocytosis, and engagement, of 5{beta}1 and EGFR in order to precisely co-ordinate adhesion disassembly, cytoskeletal dynamics and cell migration.

cell biology