Evaluating LLM-Based Credit Rating Systems: A Pre-Registered Specification-Curve Analysis — reproducibility artifact
Reproducibility artifact for the manuscript Evaluating LLM-Based Credit Rating Systems: A Pre-Registered Specification-Curve Analysis. A pre-registered specification-curve (multiverse) study of large-language-model credit-rating instability. From a frozen corpus of 8,640 elicitations (90-item firm-profile battery with an objective Altman Z'' benchmark, crossed with a 32-specification factorial grid of elicitation choices spanning provider, model version, temperature, prompt paraphrase, output format, few-shot exemplars and answer presentation, three seeds, two providers) the analysis regenerates the OLS (Type-II ANOVA) variance decomposition, the honored-determinism permutation test, within/cross-vendor Fleiss-kappa agreement, the granularity-kappa ladder, and the economic translation (portfolio turnover, Cornaggia-anchored spread, a post-hoc Basel CRE20 capital recast, and the RCAP calibration). A decontaminated real-firm arm (31 anonymized, per-issuer-perturbed issuers benchmarked to disclosed agency ratings, under a pre-registered fingerprinting gate) shows the WATCH-stratum flip-share replicates (consistent with the constructed battery; the difference is not distinguishable from zero, though a formal TOST equivalence test is inconclusive at this sample size). A post-hoc flagship model-tier arm (gpt-5.4 and gemini-3.1-pro-preview on the frozen 12-specification grid, 1,620 elicitations) finds no evidence that the instability attenuates on larger models. The artifact also derives and validates a deployable specification-instability score (ROC-AUC 0.94 on held-out specifications; 0.70 on the real-issuer arm), with a self-consistency curve and a per-issuer cost model. The reproducible run is offline, deterministic and free; live model capture (vendor API keys) is intentionally excluded.