Skip to content
#small language model Open access

Auditing a Frozen Neural Goodness-of-Fit Test for Logistic Regression: Validity, Certified Consistency and Local Power of DeepGOF-1

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Reproducibility archive (version 2.3) for the paper Auditing a Frozen Neural Goodness-of-Fit Test for Logistic Regression: Validity, Certified Consistency and Local Power of DeepGOF-1. Version 1.0.0 of the DeepGOF-1 archive is 10.5281/zenodo.22113220. DeepGOF-1 is a goodness-of-fit test for logistic regression whose statistic is a convolutional network trained once, offline, on simulated departures and shipped frozen (18,273 parameters). The network reads the fitted model's residual map, a 6×6 grid of standardized residual sums over the ranks of the two strongest covariates (or, in the all-pairs reading, of every pair of covariates, taking the largest score), and the p-value is the rank of the observed score inside the analyst's own parametric bootstrap (B = 199). New in version 2.3.0. The network against a chi-square statistic and the largest cell of its own map, and the combined reading; locating the misfit against cumulative residuals and a spline screen; external validation of a published risk model by an exact Monte Carlo test of frozen predictions (deepgof1.external() in ebrahim.gof 2.9.0); on the SUPPORT data, the ladder under the final readings, a plasmode up to all 8,873 patients, the computing times of DeepGOF-1 and the projection test, and a split-sample external validation; and corrupted records at n = 500 and 1,000 for DeepGOF-1, the projection test and BAGofT. New in version 2.2.0. A study of very small samples and wrong links: the size of five tests at n = 20 to 100 and their power against misfit in the covariates and against four wrong links. New in version 2.1.0. A bug fix in the deployed test: up to version 2.0.0 the bootstrap refitted the model formula on the model frame, which fails for any term that transforms a covariate (log, spline, polynomial) and returned p = 1; the refits now use the design matrix. The all-pairs reading as an option. New studies: the shipped network on the level-study nulls and on the benchmark with both readings, a U-shaped covariate among linear ones, the projection test on the corrupted-record datasets, and the SUPPORT in-hospital mortality application. Contents. R/, inst/extdata/: the deployable test in base R with the frozen weights; ties among covariate values are broken at random, once per call, and each bootstrap sample is refitted on the fitted model's design matrix, as in ebrahim.gof 2.9.0. training/: the training corpus (8,400 labelled maps), its generator, the network and training loop, the export to R and the cross-language check. benchmark/: self-contained harnesses for the two simulation studies (the 60-cell design grid and the four settings of Liu et al. 2024), with a smoke test that regenerates datasets and reproduces the deposited p-values exactly. results/: every per-replicate p-value behind the benchmark tables and figures, and the FDIC application's results. theory/: scripts, raw simulation output and notes behind the theoretical results of the paper and of its Supporting Information (validity and the finite-sample size bound, the data-computable calibration-gap bound and its rate, conditional consistency and the blind cone, local asymptotic power and covariates that carry no signal, finite-sample power certificates, the Neyman–Pearson optimality gap, one corrupted record at small n, the added comparison tests, the simulation with two active covariates among ten, the real-design check, the tie-breaking study). bagoft/: the paired comparison with BAGofT: every per-replicate p-value, timings, the verification of the fast implementation and the one-line repair the released BAGofT code needs for a single-covariate model. application/fdic/: the bank-failure application, rebuilt from two keyless FDIC BankFind API calls. tables/: one script that recomputes every generated number of the paper's tables from the deposited files. figures/: the scripts of the paper's figures. README.md maps every table and figure of the paper and every result of the Supporting Information to its files. The deployable test alone is also distributed as deepgof1() in the R package ebrahim.gof (version 2.9.0 or later). Code is MIT-licensed; data, results, corpora and weights are CC BY 4.0 (see LICENSE).

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.