Invertible Logits Transformation (InvLT), which applies a learned scalar MLP element-wise to the pre-softmax logits, consistently outperforms a broad set of post-hoc baselines on standard calibration metrics.
Abstract
Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining. An ideal calibrator should correct nonlinear miscalibration, scale gracefully to large label spaces, and preserve the original predictions; existing methods typically violate at least one of these properties---temperature scaling lacks expressivity, more flexible parametric alternatives introduce parameters that grow with the number of classes $C$, and other expressive methods do not preserve the rank ordering of class scores and may alter the predicted class. We propose \textbf{Invertible Logits Transformation (InvLT)}, which applies a learned scalar MLP $f:\mathbb{R}\to\mathbb{R}$ element-wise to the pre-softmax logits. Sharing $f$ across all logit dimensions makes the parameter count independent of $C$. Monotonicity of $f$---and hence preservation of the argmax prediction---is softly encouraged via a paired inverse network rather than enforced through the numerical integration required by prior monotone calibrators; this avoids their computational overhead while empirically preserving the original classification accuracy in every setting we evaluate. Across standard image classification benchmarks and a range of architectures, InvLT consistently outperforms a broad set of post-hoc baselines on standard calibration metrics.
Outcomes are regressed on a calibrated probability vector for unobserved class membership. Under a structural conditional mean excluding the score and conditional calibration, the observed-data model reduces to a partially linear regression. The probability vector is a Berkson-type surrogate for membership, so the effect vector $\tau$ is identified without attenuation. In practice the vector is often coarsened to a hard label - an argmax, a confidence threshold - and that label need not retain the Berkson property. For any coarsening the plug-in estimator converges to $\mathcal{A}\tau$, where $\mathcal{A}-I$ is determined by the regression of the discarded part of the score on the retained part. Coarsening therefore leaves $\tau$ undistorted exactly when that regression vanishes, and otherwise distorts some contrasts far more than others. The same operator determines the bias that drives coverage loss. Where that bias is of the order of the standard error, the Wald interval has limiting coverage $\Phi(z-\nu)-\Phi(-z-\nu)$, with $\nu$ their ratio. A fixed bias sends coverage to zero. Operator, standard error, and - through the uncoarsened estimator - the bias are estimable from observed data, so the implied coverage can be approximated before the interval is reported. Simulations show severe coverage loss after argmax coarsening. Three real-data audits exhibit the direction-specific distortion.
A novel Ante-hoc Explainable AI framework designed to bridge the interpretability-accuracy trade-off in high-stakes financial prognosis, specifically within credit scoring systems, which provides a verifiable and robust solution for modern, regulatory-compliant financial environments.
Deniz NoorMohammadzadehMaleki, Mahdi Baghaei Oskouei, Alireza Taheri et al.· AI and Ethics· 0 citations
Classical likelihood-ratio tests and $\Delta$AIC exacerbate the statistical significance crisis by scaling with sample size, often flagging negligible improvements as highly significant. While causal estimands like the average treatment effect (ATE) quantify practical magnitude, their reliance on the expectation operator ties them to the data's original coordinate scale. Furthermore, existing pseudo-$R^2$ metrics are inadequate: variance-based measures ignore higher-order distributional changes, and current formulations lack invariance to monotone transformations. We resolve these limitations by introducing Entropic Variance (EV) as a rigorous, scale-independent generalization of error variance in ordinary least squares. We define the population EV-based parameter, $\rho^2_V$, which projects unbounded cross-entropy onto a standardized $[0,1]$ scale, and establish that the EV-based $F_\text{V}$ statistic asymptotically follows an $F$-distribution. Building on these distributional properties, we propose two estimators: the empirical population $R^2_{\text{SV}}$ and the out-of-sample predictive $R^2_{\text{SVP}}$. Both are derived by exponentiating per-observation cross-entropy and incorporate a degrees-of-freedom correction for training optimism. Leveraging the $F_\text{V}$-distribution, we derive refined $p$-values and confidence intervals for $\rho^2_V$ without requiring intractable Fisher information matrices. Simulation studies and a Parkinson's disease microbiome application demonstrate the superiority of variable selection via these EV-$R^2$ metrics. Notably, evaluating the $R^2_{\text{SVP}}$ of a LASSO path via data-splitting reduced false discovery rates from 80% to 6% in simulations while fully preserving signal recall.
This work derives a finite-dimensional dual formulation of PrO inference that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error and uses an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a unique predictive distribution even though the parameter distributions have no weak limit on the original parameter space.
Aurya Javeed, D. Kouri, Teresa Portone et al.· 0 citations
FBO, which uses a closed-form adjoint the authors derive for the squared-loss case to obtain an exact hypergradient, and ITD, which differentiates through unrolled inner steps and extends beyond squared loss, consistently match or outperform strong HSIC, adversarial, linear-dependence, and generalized-DP baselines.
Ieva Petrulionyte, Julien Mairal, Michael Arbel· 0 citations
The contribution is accordingly not a better estimator but a characterization of \emph{when prior-based correction is justified}, plus designs that supply the missing information when it is not, plus designs that supply the missing information when it is not.
Jian Xu, Delu Zeng, John W. Paisley et al.· 0 citations