Skip to content
Preprint

Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

HYMELL is introduced, a hybrid three-level framework for estimating LLM inference latency and energy by combining analytical modeling with machine learning (ML), which enables fast, hardware-free design space exploration and energy-efficient optimization.

Abstract

The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware design. In this work, we introduce Hybrid Modeling for Energy and Latency of LLMs (HYMELL), a hybrid three-level framework for estimating LLM inference latency and energy by combining analytical modeling with machine learning (ML). HYMELL models LLM execution through a three-level hierarchy: analytical estimation of primitive operations, ML prediction of higher-level components, and an end-to-end model that captures system-level overheads across both prefill and decode phases. The framework supports diverse architectures, including dense and mixture-of-experts (MoE) feed-forward networks (FFNs), as well as multi-head attention (MHA) and grouped-query attention (GQA) mechanisms. Evaluated on an NVIDIA H100 graphics processing unit (GPU), HYMELL achieves high predictive accuracy; notably, for LLaMA 3 8B, it attains less than 5% error for both prefill and decode phases. By predicting execution costs directly from architectural parameters, it enables fast, hardware-free design space exploration and energy-efficient optimization.

View source

Similar papers

Open access Aug 2026

Optimizing Latency and Energy Efficiency in Edge-Native Large Language Models (LLMs) for Autonomous Mobile Agents

This study introduces an edge-native framework for optimizing latency and energy efficiency in LLM-enabled autonomous mobile agents and shows decreased communication overhead, increased operational continuity, and faster response times without significantly lowering language comprehension or decision-making precision.

A. Rajalakshmi, D. Saveetha, S. V. Manikanthan et al. · 0 citations
Preprint Aug 2026

Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures

Results show that attention mechanism is the primary factor governing how decode energy scales with context length, and model size primarily determines absolute energy consumption, while batching reduces both energy per generated token and request latency by up to 87%.

Molka Chkir, Syed Muhammad Danish, Jos Höll et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Empowering Hybrid Attention Models on NPUs

HA-NPU is presented, the first system to enable efficient hybrid attention LLM inference on edge NPUs without modifying the underlying algorithms, and enhances execution efficiency by reorganizing the dataflow of the LA components across three levels.

Yin-Yuan Zhang, Da-Liang Xu, Xiao-Long Huang et al. · 0 citations
#machine learning Preprint Sep 2026

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajectories extend. While in practice, many queries do not require the capabilities of the largest available model, and routinely directing such queries to a high-capability model...

Muhammad Abdur Rab Siddiqui, Daniel Rojas, Chen Yang et al. · 0 citations
Review

Toward Sustainable On-Device Intelligence: A Survey on Energy-Efficient RAG Systems with Small Language Models

The three-way intersection of on-device AI inference optimization, retrieval-augmented generation, and Green AI / sustainability has not previously been drawn together into a unified system-design perspective, and this survey undertakes that cross-domain synthesis.

Zhiyuan Cheng, Longying Lai, Yue Liu et al. · 6 citations
Conference Open access 2026

DeepSeek-V3: Architecture and Optimizations-A Practical Review

The design of transformer-based Large Language Models (LLMs) is being radically changed through new architectures that are able to overcome scalability limitations of previous designs, including Mixture-of-Experts (MoE), Multi-Head Latent Attention (MLA), and Multi-Token Prediction (MTP). As an open-weighted model rele...

Yassine Zouhdi, B. Hdioud · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.