HYMELL is introduced, a hybrid three-level framework for estimating LLM inference latency and energy by combining analytical modeling with machine learning (ML), which enables fast, hardware-free design space exploration and energy-efficient optimization.
Abstract
The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware design. In this work, we introduce Hybrid Modeling for Energy and Latency of LLMs (HYMELL), a hybrid three-level framework for estimating LLM inference latency and energy by combining analytical modeling with machine learning (ML). HYMELL models LLM execution through a three-level hierarchy: analytical estimation of primitive operations, ML prediction of higher-level components, and an end-to-end model that captures system-level overheads across both prefill and decode phases. The framework supports diverse architectures, including dense and mixture-of-experts (MoE) feed-forward networks (FFNs), as well as multi-head attention (MHA) and grouped-query attention (GQA) mechanisms. Evaluated on an NVIDIA H100 graphics processing unit (GPU), HYMELL achieves high predictive accuracy; notably, for LLaMA 3 8B, it attains less than 5% error for both prefill and decode phases. By predicting execution costs directly from architectural parameters, it enables fast, hardware-free design space exploration and energy-efficient optimization.
This study introduces an edge-native framework for optimizing latency and energy efficiency in LLM-enabled autonomous mobile agents and shows decreased communication overhead, increased operational continuity, and faster response times without significantly lowering language comprehension or decision-making precision.
A. Rajalakshmi, D. Saveetha, S. V. Manikanthan et al.· International Journal of Int...· 0 citations
Results show that attention mechanism is the primary factor governing how decode energy scales with context length, and model size primarily determines absolute energy consumption, while batching reduces both energy per generated token and request latency by up to 87%.
Molka Chkir, Syed Muhammad Danish, Jos Höll et al.· 0 citations
HA-NPU is presented, the first system to enable efficient hybrid attention LLM inference on edge NPUs without modifying the underlying algorithms, and enhances execution efficiency by reorganizing the dataflow of the LA components across three levels.
Yin-Yuan Zhang, Da-Liang Xu, Xiao-Long Huang et al.· 0 citations
Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajectories extend. While in practice, many queries do not require the capabilities of the largest available model, and routinely directing such queries to a high-capability model...
Muhammad Abdur Rab Siddiqui, Daniel Rojas, Chen Yang et al.· 0 citations
The three-way intersection of on-device AI inference optimization, retrieval-augmented generation, and Green AI / sustainability has not previously been drawn together into a unified system-design perspective, and this survey undertakes that cross-domain synthesis.
Zhiyuan Cheng, Longying Lai, Yue Liu et al.· 6 citations
The design of transformer-based Large Language Models (LLMs) is being radically changed through new architectures that are able to overcome scalability limitations of previous designs, including Mixture-of-Experts (MoE), Multi-Head Latent Attention (MLA), and Multi-Token Prediction (MTP). As an open-weighted model rele...
Yassine Zouhdi, B. Hdioud· EPJ Web of Conferences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.