
EB
Eug\`ene Berta, Sacha Braun, Francis Bach, Michael I. Jordan, David Holzm\"uller
· 1 min read
ResearcharXiv cs.LG
V-ECE: Estimating General Expected Calibration Errors
arXiv:2602.24230v2 Announce Type: replace-cross
Abstract: In probabilistic classification, calibration error (CE) measures the average divergence of predicted probabilities $f(X)$ from $\mathbb{P}(Y|f(X))$, the true class distribution for that predicted probability. While being a useful diagnostic tool, it is hard to estimate: popular binning-based estimators are often inconsistent and scale poorly beyond two classes. Recent work rewrites the CE as the excess risk of a model compared to the best recalibration of its own predictions, measured with a proper loss. However, this only works for Bregman-divergence-based calibration errors like the squared error, excluding the more popular $L_1$-distance-based CE. We show that using prediction-dependent proper scores can alleviate this restriction, allowing us to estimate CEs with general convex divergences, including $L_p$ distances with closed-form losses in the binary and multiclass settings. To estimate the excess risk, we introduce a more accurate recalibrator that fits a residual to temperature scaling with gradient boosting. The resulting variational estimator, V-ECE, needs no bins or clusters and lower-bounds the true calibration error in expectation. On a benchmark of semi-synthetic tasks built from real classifiers, with known true CE, V-ECE is among the most accurate binary estimators for every calibration error and significantly outperforms all multiclass estimators. Our results are accompanied by additional theory on $L_p$ CE, estimator bias, and over- or under-confidence estimation.
Original source
This story was published by arXiv cs.LG and written by Eug\`ene Berta, Sacha Braun, Francis Bach, Michael I. Jordan, David Holzm\"uller. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


