bioRxiv Science⌕ Search

Biology subjects

Biddie, S. C.

Publications and source records attributed to Biddie, S. C..

5 recordsLinked to original sources

Benchmarking tissue- and cell type-of-origin deconvolution in cell-free transcriptomics

Plasma cell-free RNA (cfRNA) reflects tissue- and cell-type-specific activity across pathological states and is a promising biomarker for organ injury and disease. Computational deconvolution methods are widely used to infer organ and cell-type contributions to cfRNA profiles. However, most were originally developed for single-tissue bulk transcriptomes and their performance in body-wide cfRNA settings, where any tissue or cell type can contribute, remains poorly characterised. Here, we present a systematic benchmarking of tissue- and cell type-of-origin deconvolution for plasma cfRNA that considers both methodological and reference-related sources of variability under realistic cfRNA simulation settings. We evaluated seven commonly used deconvolution methods across distinct algorithmic classes and multi-organ reference configurations derived from bulk and single-cell atlases. We assessed performance using simulation frameworks that model multi-organ mixtures, technical noise, and transcript degradation. We further examined deconvolution methods across multiple previously published clinical cfRNA cohorts spanning diverse disease contexts. Across both tissue- and cell-type-level analyses, deconvolution performance was strongly influenced by both method choice and reference parameters. Tissue-of-origin inference was comparatively robust across simulated and clinical datasets, recovering disease-associated organ signals and concordance with biochemical markers. In contrast, cell type-of-origin inference showed greater variability and reduced consistency across analytical settings, leading to divergent interpretations in both simulations and published clinical cfRNA cohorts. Together, these findings demonstrate that methodological and reference-related variability are major sources of uncertainty in cfRNA deconvolution, with tissue-level inference being more robust than cell-type-level inference. Our benchmarking framework provides guidance for reference selection and comparative interpretation in cfRNA deconvolution.

bioinformatics↗

Identifying severe COVID-19 risk variants modulating enhancer reporter activity in lung cells

Common genetic variants contribute to risk for complex human diseases. However, despite thousands of associations, variants modulating disease risk and their functional impact remain largely unknown. This includes SARS-CoV-2 infection, where outcomes range from asymptomatic to fatal. Most host risk variants associated with COVID-19 disease, identified through genome wide association studies, are located in the non-coding genome and may function by altering gene expression in disease-relevant cells and tissues. To address this at scale, we tested >4800 severe COVID-19-associated variants to determine the impact of individual variants and variant combinations on regulatory activity using Self-Transcribing Active Regulatory Region sequencing, a massively-parallel reporter assay, in a lung epithelial cell line (A549). We identify 166 variants within active sequences, of which 29 modulate activity allele-specifically. Evaluating variant combinations, we observe both additive and non-additive effects on regulatory activity. We employ state-of-the-art deep learning models to interpret allele-specific variant effects on regulatory activity and endogenous genomic features. Our work provides a set of prioritised severe COVID-19-associated variants that modulate regulatory activity in lung epithelial cells, candidate transcription factors, and candidate target genes with potential to be disease modifying.

molecular biology↗

Disease-associated genetic variants can cause mutations in tissue-specific protein isoforms

Genetic variants can cause protein-coding mutations that result in disease. Variants are typically interpreted using the reference transcript for a gene. However, most human multi-exon genes encode alternative isoforms. Here, we show that coding exons in alternative isoforms harbour more population variants than exons of reference isoforms, consistent with their reduced evolutionary constraint, and that these variants are more likely to cause nonsynonymous coding mutations. Common and rare disease-associated variants mapping to alternative transcripts can lead to amino acid substitutions predicted to be structurally damaging in the corresponding protein isoform. The alternative transcripts to which disease-associated variants map demonstrate high tissue-specific expression, with many unannotated in reference human genomes, revealed only by long-read RNA-sequencing. As an example, we report an unannotated alternative transcript of the inflammasome regulator DPP9 that is lung epithelium-specific and which harbours a common genetic variant associated with severe COVID-19 and lung fibrosis. The variant causes a p.Leu8Pro missense mutation in an alternative first exon, predicted to disrupt the encoded alpha helix. These findings highlight the importance of considering alternative isoforms, their tissue-specific expression, and full-length transcripts in variant interpretation, with implications for uncovering underappreciated mechanisms of both common and rare disease.

genomics↗

A structure-guided approach to non-coding variant evaluation for transcription factor binding using AlphaFold 3

Non-coding single-nucleotide variants (SNVs) that alter transcription factor (TF) binding can affect gene expression and contribute to disease. Sequence-based methods can excel at predicting TF binding, but rely on training data and can exhibit TF-specific biases. Here we propose a structure-guided approach for non-coding SNVs, using AlphaFold 3 (AF3) to model TF-DNA complexes and FoldX for downstream physics-based assessment. Benchmarked against SNP-SELEX data for six TFs (SPIB, ELK3, ETV4, SF-1, PAX5 and MEIS2), the FoldX-based strategy showed good agreement with experimental allele preferences. Interestingly, differences in AF3s interface predicted template modelling (ipTM) score aligned even more closely with SNP-SELEX results, generally surpassing energy-based metrics. Application to known disease-associated variants recapitulated most reported effects for TFs including NKX2-5, GATA3 and USF2A-USF1. In these examples, considering both {Delta}ipTM and FoldX energies proved more reliable than either metric alone. While less accurate than state-of-the-art sequence-based methods, this work demonstrates that structural modelling can yield interpretable insights into how non-coding variants influence TF binding. By highlighting both the promise and limitations of AF3 in this context, our study provides a framework for complementary structural evaluation of regulatory variants.

bioinformatics↗

DNA-binding factor footprints and enhancer RNAs identify functional non-coding genetic variants

Genome-wide association studies (GWAS) have revealed a multitude of candidate genetic variants affecting the risk of developing complex traits and diseases. However, these highlighted regions are typically in the non-coding genome, and uncovering the functional causative single nucleotide variants (SNVs) is challenging. Prioritisation of variants is commonly based on functional genomic annotation with markers of active regulatory elements, but current approaches still poorly predict functional variants. To address this, we systematically analyse six markers of active regulatory elements for their ability to identify functional variants. We benchmark against molecular quantitative trait loci (molQTL) from assays of regulatory element activity that identify allelic effects on DNA-binding factor occupancy, reporter assay expression, and chromatin accessibility. We identify the combination of DNase footprints and divergent enhancer RNA as markers for functional variants. This signature provides high precision, trading-off low recall, thus substantially reducing candidate variant sets to prioritise variants for functional validation. We present this as a framework called FINDER - Functional SNV IdeNtification using DNase footprints and Enhancer RNA, and demonstrate its utility to prioritise variants using leukocyte count trait and analyse variants in linkage disequilibrium with a lead variant to predict a functional variant in asthma. Our findings have implications for prioritising variants from GWAS, in development of predictive scoring algorithms, and for functionally informed fine mapping approaches.

genomics↗