bioRxiv Science⌕ Search

Biology subjects

Jürgens, M.

Publications and source records attributed to Jürgens, M..

2 recordsLinked to original sources

Reliable Molecular Retrieval from Mass Spectra using Conformal Prediction

A key task in the computational analysis of liquid chromatography-tandem mass spectrometry (LC- MS/MS) data is identifying the molecular structure underlying a measured spectrum. A common approach ranks candidate molecules retrieved from chemical databases using predicted fingerprint similarities, yet standard metrics such as top-k accuracy summarize performance only at the dataset level and provide no spectrum-specific reliability statement. In this work, we apply conformal prediction to candidate-based molecular retrieval to construct spectrum-specific prediction sets that contain the true molecule with a user-specified probability. We evaluate marginal and conditional conformal prediction across three experimental scenarios representing in-distribution, partially shifted, and fully out-of-distribution settings on the MassSpecGym benchmark. When calibration and test data are aligned, conformal prediction attains the target coverage with small candidate sets for most spectra. Under distribution shift, prediction sets become larger as rankings grow more ambiguous, although candidates can still be reduced when calibration remains representative. Conditional conformal prediction improves subgroup reliability across spectra of different difficulty, with the best gains obtained using confidence-based grouping. Overall, conformal prediction turns candidate rankings into reliable, spectrum-specific candidate sets with an explicit reliability-efficiency trade-off.

bioinformatics↗

POTTR: Identifying Recurrent Trajectories in Evolutionary and Developmental Processes using Posets

Multiple biological processes, including cancer evolution and organismal development, are described as a sequence of events with a temporal ordering. While cancer evolves independently in each patient, DNA sequencing studies have shown that in some cancers different patients share specific orders of mutations and these correlate with distinct morphology, drug response, and treatment outcomes. Several methods have been developed to identify such recurrent trajectories of genetic events from phylogenetic trees, but this is complicated by high intra- and inter-tumor heterogeneity as well as uncertainty in the inferred tumor phylogenies including the ambiguous orders between some mutations. We formalize the problem of finding recurrent mutation trajectories using a novel framework of incomplete partially ordered sets (posets), which generalize representations used in previous works and explicitly account for the uncertainty in tumor phylogenies. We define the problem of identifying the largest recurrent trajectories shared in at least k input phylogenies as the maximum k-common induced incomplete subposet (MkCIIS) problem, which we show is NP-hard. We present a combinatorial algorithm, POsets for Temporal Trajectory Resolution (POTTR), to solve the MkCIIS problem using a conflict graph that models recurrent trajectories as independent sets. Thereby we identify maximum recurrent trajectories while resolving multiple sources of uncertainty, like mutation clusters, in the phylogenetic data. We apply POTTR to TRACERx non-small cell lung cancer bulk sequencing and acute myeloid leukemia single-cell sequencing data and through resolution of mutation clusters discover previously unreported trajectories of high statistical significance. On lineage tracing data of an in vitro embryoid model, POTTR identifies conserved differentiation routes across biological replicates and how these routes change in response to chemical perturbations.

bioinformatics↗