J Proteome Res. 2026 Jul 22.
Missing values (MVs) remain a significant barrier to reliable proteomics analysis, particularly in single-cell proteomics, where small amounts of starting material and limits in detection drive missing-not-at-random (MNAR) sparsity. Existing imputation methods typically target either missing-at-random (MAR) or MNAR mechanisms, resulting in a trade-off between replicate consistency and preservation of biological variation, and are largely designed for bulk data. Here, we introduce SoftHybrid, a data-driven imputation framework that jointly models missingness and protein abundance to estimate the probability of MNAR, enabling continuous weighting between MAR- and MNAR-oriented strategies. SoftHybrid requires no external priors (cell type labels, group annotations, predefined missingness assumptions, etc.), enabling fully unsupervised applications. Across ground truth benchmarks and real single-cell proteomics data sets, SoftHybrid outperforms existing methods at low input and matches or exceeds their performance at the minibulk level. By preserving the proteomic structure and abundance accuracy, it enhances the recovery of biologically meaningful signals. SoftHybrid is implemented as an R package and is freely available at GitHub.
Keywords: benchmarking; imputation; label-free proteomics; missing values; single-cell proteomics