Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.
Tianyi Guan, Yiding Wang, Haotong Yang et al.· 1 citation
The construction safety of workers in hydraulic construction sites that are crowded and difficult to manage is very serious. When personnel movement is frequent, the status of workers wearing safety helmets is difficult to monitor in real time. The focus of this study is the design of HDS-DETR model which is aimed to improve the safety recognition in hydraulic construction projects. Improvements were achieved by integrating the C2f-HDRAB Module to the RT-DETR model to strengthen the model's ability to detect features, the D-Attention mechanism to improve the model's ability to recognize important features, and SlimNeck architecture was implemented to improve the model's ability to efficiently fuse features. The results of the experiments reflect that the accuracy achieved was 94.1% and 89.6% of the improved model offered by the dedicated dataset in recall, and 94.9% of the mean Average Precision at IoU threshold 0.5, which is a 3.7% increase in the original model. The ablation tests demonstrate the effectiveness of the correction of modules and the proposed design is aimed at the complex nature of hydraulic construction, and provides real-time hard hat wearing monitoring. Safety management of the hydraulic engineering construction project provides support and improves the safety condition recognition in smart water conservancy construction projects.
Shousong Liu, Qiulei Zhang, J. Mi et al.· International Conference on...· 0 citations