Auditing a Frozen Neural Goodness-of-Fit Test for Logistic Regression: Validity, Certified Consistency and Local Power of DeepGOF-1
Abstract
Reproducibility archive (version 2.1) for the paper Auditing a Frozen Neural Goodness-of-Fit Test for Logistic Regression: Validity, Certified Consistency and Local Power of DeepGOF-1. Version 1.0.0 of the DeepGOF-1 archive is 10.5281/zenodo.22113220. DeepGOF-1 is a goodness-of-fit test for logistic regression whose statistic is a convolutional network trained once, offline, on simulated departures and shipped frozen (18,273 parameters). The network reads the fitted model's residual map, a 6×6 grid of standardized residual sums over the ranks of the two strongest covariates (or, in the all-pairs reading, of every pair of covariates, taking the largest score), and the p-value is the rank of the observed score inside the analyst's own parametric bootstrap (B = 199). New in version 2.1.0. A bug fix in the deployed test: up to version 2.0.0 the bootstrap refitted the model formula on the model frame, which fails for any term that transforms a covariate (log, spline, polynomial) and returned p = 1; the refits now use the design matrix. The all-pairs reading as an option. New studies: the shipped network on the level-study nulls and on the benchmark with both readings, a U-shaped covariate among linear ones, the projection test on the corrupted-record datasets, and the SUPPORT in-hospital mortality application. Contents. R/, inst/extdata/: the deployable test in base R with the frozen weights; ties among covariate values are broken at random, once per call, and each bootstrap sample is refitted on the fitted model's design matrix, as in ebrahim.gof 2.9.0. training/: the training corpus (8,400 labelled maps), its generator, the network and training loop, the export to R and the cross-language check. benchmark/: self-contained harnesses for the two simulation studies (the 60-cell design grid and the four settings of Liu et al. 2024), with a smoke test that regenerates datasets and reproduces the deposited p-values exactly. results/: every per-replicate p-value behind the benchmark tables and figures, and the FDIC application's results. theory/: scripts, raw simulation output and notes behind the theoretical results of the paper and of its the Supporting Information (validity and the finite-sample size bound, the data-computable calibration-gap bound and its rate, conditional consistency and the blind cone, local asymptotic power and covariates that carry no signal, finite-sample power certificates, the Neyman–Pearson optimality gap, one corrupted record at small n, the added comparison tests, the simulation with two active covariates among ten, the real-design check, the tie-breaking study). bagoft/: the paired comparison with BAGofT: every per-replicate p-value, timings, the verification of the fast implementation and the one-line repair the released BAGofT code needs for a single-covariate model. application/fdic/: the bank-failure application, rebuilt from two keyless FDIC BankFind API calls. tables/: one script that recomputes every generated number of the paper's tables from the deposited files. figures/: the scripts of the paper's figures. README.md maps every table and figure of the paper and every result of the Supporting Information to its files. The deployable test alone is also distributed as deepgof1() in the R package ebrahim.gof (version 2.9.0 or later). Code is MIT-licensed; data, results, corpora and weights are CC BY 4.0 (see LICENSE).