bioRxiv Science⌕ Search

Biology subjects

Foschini, L.

Publications and source records attributed to Foschini, L..

2 recordsLinked to original sources

The SEA-AD DREAM Challenge: Community benchmarking human and AI agent solutions for Alzheimer's disease neuropathology prediction from single-nucleus transcriptomics

Single-nucleus transcriptomic atlases offer an unprecedented opportunity to connect cellular molecular states with Alzheimer's disease (AD) neuropathology, but whether these profiles encode reproducible, predictive information about pathological burden remains unclear. We present the SEA-AD DREAM Challenge, an open, international, model-to-data competition built on the Seattle Alzheimer's Disease Brain Cell Atlas to predict Alzheimer's disease neuropathological severity from single-nucleus RNA-sequencing data. Participants developed containerized models to predict categorical neuropathological staging, including overall Alzheimer's disease neuropathologic change, Braak stage, Thal phase, and CERAD score, as well as quantitative amyloid-{beta} and phospho-tau burden measured by 6E10 and AT8 immunohistochemistry. Across 17 eligible teams from 15 countries, the crowdsourcing framework enabled systematic comparison of diverse computational approaches and surfaced a broad landscape of modeling strategies and candidate predictive features. Top-performing methods achieved near-perfect prediction of categorical staging, with the best submission reaching a quadratic weighted kappa of 1.0 for the Overall AD Neuropathological Change score (ADNC), and competitive prediction of quantitative pathological burden in held-out data, with a best concordance correlation coefficient of 0.48. Post hoc perturbation analyses revealed that top categorical-stage predictions relied heavily on donor-level metadata-driven signals rather than transcriptomic features, whereas quantitative pathology prediction was more robust and supported by transcriptomic and cell-type-associated features with potential biological relevance to AD progression. The challenge also introduced the first AI Agent Track in a DREAM Challenge, providing an early benchmark for autonomous and human-guided agentic model development in single-cell neuroscience. This work demonstrates that single-nucleus transcriptomes encode substantial information about Alzheimer's disease pathology, establishes a reproducible benchmark for molecular neuropathology prediction, and highlights critical principles for designing privacy-preserving, leakage-aware community challenges using deeply phenotyped human brain data.

neuroscience↗

Towards Useful and Private Synthetic Omics: Community Benchmarking of Generative Models for Transcriptomics Data

BackgroundThe synthesis of anonymized data derived from real-world cohorts offers a promising strategy for regulatory-compliant and privacy-preserving biological data sharing, potentially facilitating model development that can improve predictive performance. However, the extent to which generative models can preserve biological signals while remaining resilient to adversarial privacy attacks in high-dimensional omics contexts remains underexplored. To address this gap, the CAMDA 2025 Health Privacy Challenge launched a community-driven effort to systematically benchmark synthetic and privacy-preserving data generation for bulk RNA-seq cohorts. ResultsBuilding on this initiative, we systematically benchmarked 11 generative methods across two cancer cohorts ([~]1,000 and [~]5,000 patients) over 978 landmark genes. Methods were evaluated across complementary axes of distributional fidelity, downstream utility, biological plausibility and empirical privacy risk, with emphasis on trade-offs between vulnerability to membership inference attacks (MIA) and other evaluation dimensions. Expressive deep generative models achieved strong predictive utility and differential expression recovery, but were often more vulnerable to membership inference risk. Differentially private methods improved resistance to attacks at the cost of reduced utility, while simpler statistical approaches offered competitive utility with moderate privacy risk and fast training. ConclusionsSynthetic bulk RNA-seq quality is inherently multi-dimensional and shaped by trade-offs between utility, biological preservation and privacy. Our results indicate that differences in model architecture drive distinct trade-offs across these axes, suggesting that model choice should align with dataset characteristics, intended downstream use and privacy requirements. Privacy risk should also be assessed using multiple complementary attack methods and, where possible, formal differential privacy protection.

bioinformatics↗