bioRxiv ScienceSearch

Biology subjects

Vidal, M.

Publications and source records attributed to Vidal, M..

4 recordsLinked to original sources

Network-based prediction of protein interactions

As biological function emerges through interactions between a cells molecular constituents, understanding cellular mechanisms requires us to catalogue all physical interactions between proteins [1-4]. Despite spectacular advances in high-throughput mapping, the number of missing human protein-protein interactions (PPIs) continues to exceed the experimentally documented interactions [5, 6]. Computational tools that exploit structural, sequence or network topology information are increasingly used to fill in the gap, using the patterns of the already known interactome to predict undetected, yet biologically relevant interactions [7-9]. Such network-based link prediction tools rely on the Triadic Closure Principle (TCP) [10-12], stating that two proteins likely interact if they share multiple interaction partners. TCP is rooted in social network analysis, namely the observation that the more common friends two individuals have, the more likely that they know each other [13, 14]. Here, we offer direct empirical evidence across multiple datasets and organisms that, despite its dominant use in biological link prediction, TCP is not valid for most protein pairs. We show that this failure is fundamental - TCP violates both structural constraints and evolutionary processes. This understanding allows us to propose a link prediction principle, consistent with both structural and evo-lutionary arguments, that predicts yet uncovered protein interactions based on paths of length three (L3). A systematic computational cross-validation shows that the L3 principle significantly outperforms existing link prediction methods. To experimentally test the L3 predictions, we perform both large-scale high-throughput and pairwise tests, finding that the predicted links test positively at the same rate as previously known interactions, suggesting that most (if not all) predicted interactions are real. Combining L3 predictions with experimen-tal tests provided new interaction partners of FAM161A, a protein linked to retinitis pigmentosa, offering novel insights into the molecular mechanisms that lead to the disease. Because L3 is rooted in a fundamental biological principle, we expect it to have a broad applicability, enabling us to better understand the emergence of biological function under both healthy and pathological conditions.\n\nSummaryWe unveil a fundamental organizing principle of biological networks and demonstrate its predictive power for uncovering novel protein interactions.

systems biology

Whole genome sequence of Mapuche-Huilliche Native Americans

BackgroundWhole human genome sequencing initiatives provide a compendium of genetic variants that help us understand population history and the basis of genetic diseases. Current data mostly focuses on Old World populations and information on the genomic structure of Native Americans, especially those from the Southern Cone is scant.\n\nResultsHere we present a high-quality complete genome sequence of 11 Mapuche-Huilliche individuals (HUI) from Southern Chile (85% genomic and 98% exonic coverage at > 30X), with 96-97% high confidence calls. We found approximately 3.1x106 single nucleotide variants (SNVs) per individual and identified 403,383 (6.9%) of novel SNVs that are not included in current sequencing databases. Analyses of large-scale genomic events detected 680 copy number variants (CNVs) and 4,514 structural variants (SVs), including 398 and 1,910 novel events, respectively. Global ancestry composition of HUI genomes revealed that the cohort represents a marginally admixed population from the Southern Cone, whose genetic component is derived from early Native American ancestors. In addition, we found that HUI genomes display highly divergent and novel variants with potential functional impact that converge in ontological categories essential in cell metabolic processes.\n\nConclusionsMapuche-Huilliche genomes contain a unique set of small- and large-scale genomic variants in functionally linked genes, which may contribute to susceptibility for the development of common complex diseases or traits in admixed Latinos and Native American populations. Our data represents an ancestral reference panel for population-based studies in Native and admixed Latin American populations.

genomics

Controllability in an islet specific regulatory network identifies the transcriptional factor NFATC4, which regulates Type 2 Diabetes associated genes

Probing the dynamic control features of biological networks represents a new frontier in capturing the dysregulated pathways in complex diseases. Here, using patient samples obtained from a pancreatic islet transplantation program, we constructed a tissue-specific gene regulatory network and used the control centrality (Cc) concept to identify the high control centrality (HiCc) pathways, which might serve as key pathobiological pathways for Type 2 Diabetes (T2D). We found that HiCc pathway genes were significantly enriched with modest GWAS p-values in the DIAbetes Genetics Replication And Meta-analysis (DIAGRAM) study. We identified variants regulating gene expression (expression quantitative loci, eQTL) of HiCc pathway genes in islet samples. These eQTL genes showed higher levels of differential expression compared to non-eQTL genes in low, medium and high glucose concentrations in rat islets. Among genes with highly significant eQTL evidence, NFATC4 belonged to four HiCc pathways. We asked if the expressions of T2D-associated candidate genes from GWAS and literature are regulated by Nfatc4 in rat islets. Extensive in vitro silencing of Nfatc4 in rat islet cells displayed reduced expression of 16, and increased expression of 4 putative downstream T2D genes. Overall, our approach uncovers the mechanistic connection of NFATC4 with downstream targets including a previously unknown one, TCF7L2, and establishes the HiCc pathways relationship to T2D.

systems biology

Expanding the Atlas of Functional Missense Variation for Human Genes

Although we now routinely sequence human genomes, we can confidently identify only a fraction of the sequence variants that have a functional impact. Here we developed a deep mutational scanning framework that produces exhaustive maps for human missense variants by combining random codon-mutagenesis and multiplexed functional variation assays with computational imputation and refinement. We applied this framework to four proteins corresponding to six human genes: UBE2I (encoding SUMO E2 conjugase), SUMO1 (small ubiquitin-like modifier), TPK1 (thiamin pyrophosphokinase), and CALM1/2/3 (three genes encoding the protein calmodulin). The resulting maps recapitulate known protein features, and confidently identify pathogenic variation. Assays potentially amenable to deep mutational scanning are already available for 57% of human disease genes, suggesting that DMS could ultimately map functional variation for all human disease genes.

molecular biology