What the Karpowicz Theorem Does Not Prove: A Three-Resource Theory of the LLM Einstein Test
Two AI systems can reach the same scientific conclusion for different reasons: one may generate the candidate sooner, another may check it more cheaply, and a third may obtain decisive evidence earlier. An end-to-end score records success while hiding which interface made success possible, which resource an intervention changed, and what the result warrants. This paper develops an interface-sensitive three-resource theory for the Einstein Test for large language models. A running laboratory example follows a team choosing between faster candidate generation and earlier access to a distinguishing experiment. Generation effort, computational verification and empirical time form separate coordinates. A finite-budget theorem gives sufficient conditions that connect them: positive target support, consistent witness-producing experiments, bounded screening and complete verification. A complementary theorem composes resource floors for a specified serial procedure. The empirical analysis distinguishes strict refutation from sequential statistical acceptance and allows instruments and experimental opportunities to change the completion frontier. Computational recognition depends on representation: broad recursively axiomatised classes admit undecidability reductions, while suitable real-closed-field representations permit decision procedures. The worked example gives a quantitative success guarantee and shows that generator and instrument improvements alter different costs. Historical cases explain how to choose a data cutoff and acceptance rule without treating an observed discovery interval as a universal lower bound. Publicly deposited on 13 May 2026, the account predates several later 2026 studies that independently foreground these interfaces. The framework states the interfaces required for success, the resource changed by an intervention, and the evidential conclusion supported by the result.