A load-aware GPU-based dynamic graph pattern matching scheme is proposed to make full use of GPU computing resources and a task overhead prediction model is proposed to guide task allocation to alleviate the load imbalance between multiple GPU devices.
Yu Zhang, Yu-Luo Guo, Fu-Bing Mao et al.· IEEE Transactions on Knowled...· 0 citations
Memory disaggregation provides key-value stores larger memory capacity at low cost. Emerging compute express link (CXL) enables efficient memory disaggregation. It, however, dramatically slows down the system performance as disaggregated memory accesses are considerably slower than local memory accesses. This paper pre...
Chencheng Ye, Yuanchao Xu, Xipeng Shen et al.· ACM Transactions on Architec...· 0 citations
Experiments on tool-augmented agent workloads show that CoAct improves per-step execution efficiency and resource utilization while achieving competitive or superior task accuracy, demonstrating that contrastive online dispatch can expose substantial parallelism in LLM-agent workflows without retraining the underlying...
Yuyang Peng, Yanling Xu, Shu-Yi Wang et al.· Proceedings of the 32nd ACM...· 1 citation
LLM agents solve complex tasks by executing multi-step workflows that interleave LLM inference with external tool calls, yet execution efficiency is often the dominant bottleneck in real deployments because LLM-generated workflows are typically chain-structured and inherently sequential, limiting parallelism and underu...
Yuyang Peng, Yanling Xu, Shuyi Wang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.