When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that woul...