bioRxiv · 10.1101/2020.10.13.323618
DIMA: Data-driven selection of a suitable imputation algorithm
Abstract
MotivationImputation is a prominent strategy when dealing with missing values (MVs) in proteomics data analysis pipelines. However, the performance of different imputation methods is difficult to assess and varies strongly depending on data characteristics. To overcome this issue, we present the concept of a data-driven selection of a suitable imputation algorithm (DIMA). ResultsThe performance and broad applicability of DIMA is demonstrated on 121 quantitative proteomics data sets from the PRIDE database and on simulated data consisting of 5 - 50% MVs with different proportions of missing not at random and missing completely at random values. DIMA reliably suggests a high-performing imputation algorithm which is always among the three best algorithms and results in a root mean square error difference ({Delta}RMSE) [≤] 10% in 84% of the cases. Availability and ImplementationSource code is freely available for download at github.com/clemenskreutz/OmicsData.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Egert, J., Warscheid, B., Kreutz, C.. 2020-10-14. DIMA: Data-driven selection of a suitable imputation algorithm. https://doi.org/10.1101/2020.10.13.323618
Cite the original work for its findings. Save a collection to share your selection of sources.