bioRxiv Science⌕ Search

Biology subjects

Riba, M.

Publications and source records attributed to Riba, M..

6 recordsLinked to original sources

Anonymized Somatic Tumor Twins (STTs) enable open genome data sharing and use in research and clinical oncology

The study of somatic variants from tumor genomes is fundamental to cancer research and clinical decision-making. However, existing data protection frameworks impose restrictions on the use and sharing of these variants in conjunction with sensitive germline information. To overcome these challenges, we developed GenomeAnonymizer, the first method to anonymize short-read DNA sequences from tumor-normal pairs. This generates Somatic Tumor Twins (STTs), an anonymized version of the original data that preserves the donors privacy while retaining somatic tumor information and sequencing noise. This method successfully removed all detectable germline variants from the 47 PCAWG-Pilot samples. We further demonstrate that Whole-Genome Sequencing (WGS) STTs preserve more than 98% of the original somatic variants, enabling reliable downstream analysis that replicates somatic-related findings from the original samples, including cancer driver genes, mutational signatures, and intratumor heterogeneity. Importantly, we also show that STTs can reproduce the identification of actionable genes and downstream clinical interpretations and decision-making. We generated a cancer cohort of STTs matched with synthetic clinical data that could be openly shared and used across projects and centers worldwide. This paradigm-shifting approach will accelerate discovery and clinical translation in oncology and enable the robust benchmarking of genome analysis and large-scale data infrastructures.

bioinformatics↗

A methodological framework for accommodating Cancer Genomics Information in OMOP-CDM using Variation Representation Specification (VRS).

The OMOP Common Data Model (OMOP CDM) in which observational health data are organized and stored is a broadly accepted data standard which helps clinical research facilitating federation study protocols. In case of cancer studies, there is a growing need to incorporate cancer genomics data in a standardized way. Starting from a brief overview of the basic features of the OMOP CDM, we imagine a path of increasing complexity for including known biomarker genomic data coming from pathology or reports or clinical laboratory findings, towards storing thousands of known and unknown variants coming from genome sequencing data. Data should be stored using standardized identifiers, including those defined by the Global Alliance for Genomics and Health (GA4GH). We propose a scalable strategy for storing genomics variants in increasingly complex scenarios and present KOIOS-VRS, a pipeline that automates the conversion of VCF files into OMOP compatible format.

bioinformatics↗

Biomarkers in cerebrospinal fluid sediments

Cerebrospinal fluid (CSF) biomarkers for neurodegenerative diseases have been extensively studied over the years. However, CSF samples are routinely centrifuged, and the resulting sediment or pellet is typically discarded to remove cellular debris and high-density particles. This standard practice raises a critical question: could these discarded sediments harbour potential biomarkers relevant to the diagnosis and prognosis of certain brain diseases? In this study, we analysed CSF pellets from various cases and identified, entrapped among undetermined remnants, brain-derived structures such as wasteosomes and psammoma bodies. Furthermore, we observed that disease-relevant proteins can become deposited in the sediment, as is the case for both tau and A{beta}42 in Alzheimers disease or tau in progressive supranuclear palsy disease. These findings suggest that some potential biomarkers might accumulate or be hidden in the sediment and, taken as a whole, the results underscore the need to broaden the scope of biomarker research. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/657984v1_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@1d5b55corg.highwire.dtl.DTLVardef@175ae02org.highwire.dtl.DTLVardef@f30e42org.highwire.dtl.DTLVardef@12d326b_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphical abstractC_FLOATNO C_FIG

neuroscience↗

Genomic signatures of climate-driven (mal)adaptation in an iconic conifer, the English yew (Taxus baccata L.)

The risk of climate maladaptation is increasing for numerous species, including trees. Developing robust methods to assess population maladaptation remains a critical challenge. Genomic offset approaches aim to predict climate maladaptation by characterising the genomic changes required for populations to maintain their fitness under changing climates. In this study, we assessed the risk of climate maladaptation in European populations of English yew (Taxus baccata), a long-lived tree with a patchy distribution across Europe, the Atlas Mountains, and the Near East, where many populations are small or threatened. We found evidence suggesting local climate adaptation by analysing 8,616 SNPs in 475 trees from 29 European T. baccata populations, with climate explaining 18.1% of genetic variance and 100 unlinked climate-associated loci identified via genotype- environment association (GEA). Then, we evaluated the deviation of populations from the overall gene-climate association to assess variability in local adaptation or different adaptation trajectories across populations and found the highest deviations in low latitude populations. Moreover, we predicted genomic offsets and successfully validated these predictions using fitness proxies assessed in plants from 26 populations grown in a comparative experiment. Finally, we integrated information from current local adaptation, genomic offset, historical genetic differentiation and effective migration rates to show that Mediterranean and high-elevation T. baccata populations face higher vulnerability to climate change than low-elevation Atlantic and continental populations. Our study demonstrates the practical use of the genomic offset framework in conservation genetics, offers insights for its further development, and highlights the need for a population-centred approach that incorporates additional statistics and data sources to credibly assess climate vulnerability in wild plant populations.

genomics↗

The Minimal Dataset for Cancer of the 1+Million Genomes Initiative

For a real impact on healthcare, precision cancer medicine requires accessibility and interoperability of clinical and genomic data across centres and countries. Due to the heterogeneous digitization in Europe and worldwide, the definition of models for standardised data collection and usability becomes mandatory if countries want to work together on this mission. The European Union 1+Million Genomes (1+MG) initiative, supported by the Horizon 2020 Beyond 1 Million Genomes project, aims at outlining data models, guidance, best practices, and technical infrastructures for transnational access to sequenced genomes, including cancer genomes. Within the framework of the cancer-focused Working Group 9, we developed the 1+MG-Minimal Dataset for Cancer (1+MG-MDC)-a data model encompassing 140 items and organized in eight conceptual domains for the collection of cancer-related clinical information and genomics metadata. The 1+MG-MDC, which results from a multidisciplinary effort, leverages pre-existing models and emphasizes the annotation and traceability of multiple aspects relevant to the complex longitudinal path of the cancer disease and its treatment. We strived to make the 1+MG-MDC easy to adopt, yet comprehensive, addressing the needs of both clinicians and researchers. We will periodically revise and update it to ensure it remains fit for purpose. We propose the 1+MG-MDC as a model to create homogeneous databases, which would, in turn, guide discussions on clinical and genomic features with prognostic or therapeutic value and foster real-world data research.

cancer biology↗

A fluorescent reporter model for the visualization and characterization of TDC

TDC are hematopoietic cells that combine dendritic cell (DC) and conventional T cell markers and functional properties. They were identified in secondary lymphoid organs (SLOs) of naive mice as cells expressing CD11c, major histocompatibility molecule (MHC)-II, and the T cell receptor (TCR) {beta} chain. Despite thorough characterization as to their potential functional properties, a physiological role for TDC remains to be determined. Unfortunately, using CD11c as a marker for TDC has the caveat of its upregulation on different cells, including T cells, upon activation. Therefore, a more specific marker is needed to further investigate TDC functions in peripheral organs in different pathological settings. Here we took advantage of Zbtb46-GFP reporter mice to explore the frequency and localization of TDC in peripheral tissues at steady state and upon viral infection. RNA sequencing analysis confirmed that TDC identified with this reporter model have a gene signature that is distinct from conventional T cells and DC. In addition, frequency and total numbers of TDC in the SLOs recapitulated those found using CD11c as a marker. This reporter model allowed for identification of TDC in situ not only in SLOs but also in the liver and lung of naive mice. Interestingly, we found that TDC numbers in the SLOs increased upon viral infection, suggesting that TDC might play a role during viral infections. In conclusion, we propose a visualization strategy that might shed light on the physiological role of TDC in several pathological contexts, including infection and cancer.

immunology↗