Preprint
Aug 2026
Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors
HYMELL is introduced, a hybrid three-level framework for estimating LLM inference latency and energy by combining analytical modeling with machine learning (ML), which enables fast, hardware-free design space exploration and energy-efficient optimization.
Saeid Shokoufa, Mohammad Erfan Sadeghi, M. Kamal et al.
· 0 citations