A recourse method is built on a gradient-boosted ensemble that retains 58% of its validity where the strongest baseline retains 41%, a distinction the standard evaluation cannot see because it never asks whether a recommendation can be carried out.
Abstract
A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than as intermediate results, and every instance becomes a point in R^M on which the model acts linearly: the score is the sum of the coordinates. This small change of view makes contrastive explanation exact. The difference between two instances is a vector that is identically zero wherever they share a leaf, so the gap between a rejected applicant and an accepted one is carried by a handful of coordinates, each traceable to a real split in a real tree. Nothing is fitted, sampled, or assumed additive in features -- the additivity is already there, in the right space. We build a recourse method on this representation and evaluate it on five tabular datasets under repeated cross-validation. Its recommendation reconstructs the model's own decision to 6.2 x 10^-15, so an auditor can re-check the arithmetic without the model. On the credit datasets it is Pareto-non-dominated on effort against realism. And when recommendations are restricted to changes the subject could actually make -- not their age, not a settled delinquency -- it retains 58% of its validity where the strongest baseline retains 41%, a distinction the standard evaluation cannot see because it never asks whether a recommendation can be carried out.
Modern AI models such as tabular foundation models and gradient-boosted ensembles can outpredict classical methods, but provide little basis for reasoning about their predictions. High-stakes decisions call for models that are both accurate and interpretable as built. Local linear modeling offers a path forward: a smooth regression function is locally well approximated by a linear one, allowing a linear fit near each query point to achieve high accuracy without sacrificing transparency. The challenges lie in learning what is"local"and developing statistical tools for interpretation. Here, we propose local distillation, in which a black-box"teacher"guides a regularized linear"student"model at each query point. The teacher (1) defines locality by upweighting training observations with similar predicted outcomes, and (2) anchors the fit with its prediction at the query point, included as a pseudo-observation whose weight is estimated from the data. For interpretation, we add a small amount of Gaussian randomization to the local objective and use refits to assess stability: selection frequencies identify reliable features at a query point, and clustering the randomized fits identifies stable subgroups across the data. Under the lasso penalty, we prove that this randomization yields feature-selection probabilities that are stable under small perturbations of the training responses. Across 17 benchmark datasets, local distillation nearly matches its AI teacher's accuracy while producing a sparse linear model at each test point. In a high-dimensional cancer gene expression example, the framework identifies patient subgroups whose local models use different genes; this heterogeneity is invisible to a global linear model, and difficult to surface in a black-box model.
The Forward-Forward algorithm trains each layer locally, so that a scalar goodness - the sum of squared activations - is high on real inputs and low on contrastive ones. Under an explicit generative model this goodness is the sufficient statistic of a likelihood-ratio test, and the pairwise form of the objective admits a gauge: a layer can lower its loss by inflating the scale of its weights rather than by separating the two populations. The analysis prescribes the repair - a whitened, scale-invariant goodness trained online within each layer - which we evaluate as a training procedure. Across three corpora, three depths and a fourfold range of layer width (13 seeds per cell), it raises linear-probe accuracy over the standard pairwise objective in every measured cell - by 4 to 7 points on eight of nine corpus-depth combinations - and closes 16-61% of the gap to end-to-end backpropagation. A control isolates the mechanism: Hinton's fixed-threshold loss also bounds the runaway, to a factor of 1.4 against 133, yet tracks the unmodified baseline - invariance to the gauge, not a bound on it, is what pays. Against the strongest published alternative - a sparse, top-k goodness - the derived objective is statistically indistinguishable on two corpora of three, yet only it removes the runaway: sparsity and gauge-invariance are independent axes, and the published variant recovers accuracy while leaving the pathology in place. We state the boundaries we measured, and every prediction was recorded before its experiment with the refutations reported.
It is shown that the contradiction in experience with neural surrogates in derivative-free optimisation dissolves once three factors are stated, and that these, rather than the fit accuracy a training curve reports, are what delimit when a learned local model pays.
In multiple instance regression (MIR) data are organized into bags (collections of instances in feature space) and the goal is to learn a mapping that assigns labels to bags. A typical assumption is that there is a so-called concept point in feature space, the proximity to which dictates the bag label. Motivated by modern MIR architectures which are based on attention, we study a softmax model that decouples the concept point and the labeling scheme. The two are respectively determined by a query direction and a value direction value in feature space. This problem isolates a basic challenge of learning both the query and value vectors from bag-level supervision. From this model we derive a parametric family of iterations in the noiseless limit, which generalizes a method known as the EM-DD algorithm. We then derive concentration results for the MLE estimators of the query and value vectors obtained from a random selection of instances. Our result for the value vector shows that a single random initialization of the value vector already points in the correct direction on average, so that a polynomial (in the number of instances per bag and the feature dimension) number of bags is enough for the EM algorithm to converge in $O(1)$ steps with high probability. A key aspect of this analysis is the interplay between concentration of empirical covariance matrices and extremal statistics arising from the selection rule.
Gradient-boosted trees dominate tabular machine learning, yet canonical correlation analysis has always relied on linear or neural encoders. We propose \textbf{TreeCCA}, the first method to train gradient-boosted tree ensembles end-to-end as CCA encoders, inheriting their plug-and-play reliability: no architecture design, familiar hyperparameters, and strong performance with defaults. The technical enabler is the Eckart-Young (EY) loss, which supplies closed-form per-sample gradients that slot directly into any standard GBT library (XGBoost, LightGBM) as a custom objective. TreeCCA is the first CCA method to combine nonlinear accuracy with native interpretability: every tree split selects one feature, so gain importances reveal which inputs drive cross-view correlation at no extra cost. We demonstrate these properties on synthetic benchmarks, where TreeCCA matches or exceeds Deep CCA (2.61 vs.\ 2.43 on Signed Power; 2.93 vs.\ 2.89 on Hermite), and on a sparse benchmark with zero linear cross-view covariance, where TreeCCA recovers the true support with $\text{Precision@}S = 1.00$ at $p=50$ while PMD finds no signal. On the UCI HAR sensor-fusion benchmark, TreeCCA achieves comparable accuracy to Deep CCA at $5\times$ lower cost, while XGBoost gain importances directly validate a physics-motivated hypothesis about the data --- an interpretation not readily available with neural encoders. Across five popular tabular multi-view datasets, TreeMCCA consistently matches or exceeds linear CCA in both nonlinear correlation extraction and downstream classification accuracy.