Skip to content
Open access

A Reproducible, Provenance-Preserving Architecture for Sequential Intervention-Threshold Evaluation in Large Language Models

Sep 2026 · Informatica · 0 citations · 17 references

Abstract

Sequential evaluation of large language models (LLMs) needs an auditable way to analyse when intervention first occurs. The architecture presented here uses cumulative scenarios and reconstructs first-ACT/censoring and at-risk datasets from a master record for discrete-time event-history analysis. Its methodological contribution is the integration of established first-event reconstruction with a frozen instrument, independent current-state calls and serving provenance, while retaining later responses for audit. In a thirteen-model demonstration, 26,000 stage-level records yielded 25,999 decision-valid responses, 3,120 scenario-runs and 10,088 at-risk rows. The derived endpoint files and all 2,000 prompt coordinates were regenerated exactly from deposited records. Complete conditional hazards were stage-dependent but non-monotone; within-group linear trend estimates were OR 12.591 (95% CI 8.481–18.694) in Group A, 10.346 (9.292–11.518) in Group B and 4.377 (3.815– 5.022) in Group C, and a quadratic sensitivity confirmed nonlinearity. Later WAIT responses occurred in 21.05% of ACT-containing runs. The reconstruction checks establish analytical reproducibility of one realised first-intervention path under recorded serving conditions; they do not establish a stable latent threshold, normative safety validity or exact recreation of historical provider responses.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.