bioRxiv · 10.1101/2025.07.10.664057
Protein Abundance Inference via Expectation Maximization in Fluorosequencing
Abstract
1Fluorosequencing produces millions of single-peptide reads, yet a principled strategy for converting these data into quantitative protein abundances has been lacking. We introduce a probabilistic framework that adapts expectation maximization to the fluorosequencing measurement process, estimating relative protein abundances with peptide inference results delivered by previously developed peptide-classification tools. The algorithm iteratively updates protein abundances, maximising the likelihood of the observed reads by obtaining more accurate protein abundance estimations. We first assess performance on simulated five-protein mixtures that reflect realistic labelling and system errors. A simple Python implementation processes one million reads in under ten seconds on a standard work-station and lowers the mean absolute error in relative abundance by more than an order of magnitude compared with a uniform-abundance guess, demonstrating robustness in protein inference for small-scale settings. Scalability is then evaluated with simulations of the complete human proteome (20 642 proteins). Ten million reads are processed in less than four hours on a NVIDIA DGX system using one Tesla V100 GPU, confirming that the method remains tractable at proteome scale. Using error rates characteristic of current fluorosequencing, the algorithm produces marginal improvements in relative abundance accuracy. However, when error rates were artificially lowered, estimation error decreased significantly. This result suggests that improvements in fluorosequencing chemistry could directly translate into substantially more accurate quantitative proteomics with this computational framework. Together, these results establish EM-based inference as a scalable model-driven bridge between peptide-level classification and protein-level quantification in fluorosequencing, laying computational groundwork for high-throughput single-molecule proteomics. Furthermore, the proposed protein inference framework can also be used as a refinement step within other inference methods, enhancing their protein abundance estimates. 2 Author summaryProteins carry out many of the functions inside our cells, but measuring how many copies of each protein are present remains a difficult task. The protein sequencing technology called fluorosequencing observes individual protein fragments one by one, producing millions of short "snapshots" in a single experiment. While this flood of data holds great promise, it is not yet clear how to turn the snapshots into reliable counts of the original proteins. In this study we developed a mathematical procedure, based on a classic "guess-and-improve" strategy called expectation maximization, that bridges this gap. Our program takes the computers best guesses for each fragment and repeatedly refines protein abundances until they best explain all of the measurements. Using realistic computer-generated data, we show that the method is both fast and accurate: it finishes in seconds for small mixtures and in only a few hours for the entire human set of proteins, both with accessible hardware. Because our approach can work with any future improvements in fragment interpretation, it lays the essential groundwork for bringing single-molecule protein sequencing into routine biological and medical use.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kipen, J., Smith, M. B., Blom, T., Zhou, S. B., Marcotte, E. M., Jalden, J.. 2025-07-14. Protein Abundance Inference via Expectation Maximization in Fluorosequencing. https://doi.org/10.1101/2025.07.10.664057
Cite the original work for its findings. Save a collection to share your selection of sources.