Preprint
Jul 2026
Moral Hazard in Multi-Agent Language Models
CREDIT (Counterfactual Replay for Evidence-Driven Information Transfer), a mechanism-aligned multi-agent prompt-optimization algorithm that uses matched hidden-state twins and total-action replay to reward robust causal contribution rather than query frequency is introduced.
Dane Malenfant
· 0 citations