HA-NPU is presented, the first system to enable efficient hybrid attention LLM inference on edge NPUs without modifying the underlying algorithms, and enhances execution efficiency by reorganizing the dataflow of the LA components across three levels.
Yin-Yuan Zhang, Da-Liang Xu, Xiao-Long Huang et al.· 0 citations
PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services, is built.
Cheng-Hua Wang, Daliang Xu, Dongqi Cai et al.· 1 citation
A systems vision for AI infrastructure in space is developed as the systems layer that manages AI capabilities across spacecraft, orbital networks, ground stations, and cloud backends, while treating orbital and physical state as part of the resource model.
Qing Li, Qi-Yang Zhang, Da-Liang Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.