bioRxiv · 10.64898/2026.09.06.749687
MutCleaner: Cleaning and Standardizing Biological Mutation Datasets for Variant Effect Prediction
Abstract
Summary: Protein mutation datasets are widely used in variant effect prediction and protein engineering, but datasets from different sources often lack consistent conventions for mutation representation, sequence representation, data organization, and experimental labels, making these resources difficult to integrate directly and limiting their use in downstream analyses and modeling. MutCleaner is an extensible Python framework that cleans, validates, and standardizes protein- and codon-level mutation datasets through composable cleaning pipelines, unified sequence and mutation data structures, and dataset-specific cleaners. It provides standardized resources covering 16 protein and codon mutation datasets with more than 12.14 million mutation records. Availability and implementation: MutCleaner is an open-source Python package released under the Apache License 2.0. The source code is available on https://github.com/xulab-research/MutCleaner, the package is distributed through https://pypi.org/project/mutcleaner/, and the documentation is available https://xulab-research.github.io/MutCleaner/.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shi, Z., Tang, Y., Yang, M., Yu, S., Shi, Y., Xu, Y.. 2026-09-08. MutCleaner: Cleaning and Standardizing Biological Mutation Datasets for Variant Effect Prediction. https://doi.org/10.64898/2026.09.06.749687
Cite the original work for its findings. Save a collection to share your selection of sources.