In long agent executions, an early error can persist through later actions and checks, while evidence needed to trace its origin is dispersed across the history. Short histories offer limited tests of recovering error origins across substantial subsequent execution. We introduce LongRCA Bench: 1,140 complete failed tra...
Yun-Fei Zhang, Bo-Yu Feng, Changhua Pei et al.· 0 citations
DUOTRACE follows a detect-before-attribute paradigm: it first detects anomalous executions and then supplies focused trajectory evidence to downstream LLM-based attribution methods, which improves agent-level and step-level attribution accuracy.
Jia-Yi Zhang, Ze-Xin Wang, DecisionMakingRon Sun et al.· 0 citations
This work introduces LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors, and presents Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions.
Yun-Fei Zhang, Bo-Yu Feng, Changhua Pei et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.