As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model weights fixed can substantially alter task performance. Current agent evaluations typically report downstream performance under a chosen harness, leaving a model's ability to develop the harness itself comparatively underexplored. We introduce HarnessDev, a benchmark that shifts the unit of evaluation from task outputs to runnable infrastructure. HarnessDev covers two stages. In Creation, the agent starts from a minimal seed and a small number of cases, then builds a complete execution system. In Evolution, it starts from its own created harness and iteratively revises it using downstream execution feedback, with the goal of improving benchmark performance. We then evaluate each constructed harness on capability (task success on held-out benchmarks) and efficiency (execution-token cost). The reported Creation results cover six creator LLMs, four domains, and five downstream benchmarks totaling 2,207 unique downstream instances, with hidden evaluation tasks withheld from development. We find that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost. Evolution produces some performance gains, but they are unstable and transfer only partially to held-out tasks. Experiments with a fixed runtime model further show that the gains depend strongly on the model executing the harness, indicating limited transfer across models.
Yu-Hao Wu, Jing-Yuan Zhang, Jia-Jun Shi et al.· 0 citations
ObjectiveThis study examined whether cognitive load produces selective effects on different trust updating pathways in AI-assisted decision making.BackgroundAlthough cognitive load affects trust in automation, its influence on the mechanisms of trial-by-trial trust updating remains unclear.MethodsA dual-task paradigm embedded in a mining exploration task manipulated cognitive load while capturing dynamic trust calibration. Guided by a dual-pathway framework, we operationalized process-based (analytical evaluation of AI recommendation correctness) and outcome-based (heuristic reliance on task outcomes) trust updating pathways. Trust dynamics and behavioral reliance were examined using linear mixed-effects models.ResultsCognitive load shifted the relative influence of the two trust updating pathways. Process-based updating was attenuated under high cognitive load, indicating reduced sensitivity to AI recommendation correctness during trust updating. Outcome-based information gained greater influence under high load, amplifying outcome-driven bias regardless of recommendation correctness. Asymmetric trust updating was evident overall, although the influence of cognitive load on this asymmetry depended on task outcomes. Overall, high cognitive load elevated both subjective trust and behavioral reliance on AI.ConclusionCognitive load shapes trust calibration through mechanism-level reconfiguration rather than global impairment. By revealing how cognitive constraints rebalance dual trust pathways-weakening analytic evaluation while amplifying heuristic outcome reliance-this study advances theoretical understanding of dynamic trust in human-AI collaboration.ApplicationThe results provide practical guidance for the design of AI systems in high-stakes settings, highlighting the need to support analytic trust updating and mitigate over-reliance under cognitive strain.
Xiaojiao Chen, Yonghan Liu, Yiran Ma et al.· Human Factors· 0 citations