Extreme Learning Machine Ensemble for Deep Learning-based List Decoding in Error-prone Video Transmission Systems
Abstract
Deep learning–based no-reference image quality assessment (NR-IQA) has been shown to reliably select the highest-quality candidate from multiple list-decoded video reconstructions. However, deep transformer models are computationally intensive for candidate-by-candidate evaluation and tend to lose accuracy when the data distribution changes. This paper proposes a hybrid framework that combines a vision transformer backbone with an Extreme Learning Machine (ELM) ensemble to reduce computational cost and mitigate sensitivity to domain shifts by enabling fine-tuning without retraining the entire NR-IQA model for candidate ranking in list-decoding applications. Candidate reconstructions arising from an error-prone video list decoding scenario are ranked using PSNR relative to the intact reference frame, a pretrained NR-IQA model, and the proposed ELM ensemble. The top-ranked candidates from the ranking strategies are then compared. Experimental results across multiple video databases show that the proposed ELM ensemble achieves an average ranking accuracy of 92.37%, comparable to the pretrained NR-IQA model (93.69%) while incurring lower analytical complexity. Since both architectures share the same ViT backbone, the proposed method reduces the ranking-head complexity, per evaluated 224 × 224 input unit, from approximately 150 – 300 GFLOPs to 2.41 GFLOPs by replacing the Swin Transformer module with an ELM adaptation head. Including the shared backbone, ELM+ requires 125.41 GFLOPs for a single ViT pass versus approximately 273 – 423 GFLOPs for MANIQA+ ; because MANIQA+ requires 20 such passes per image (about 6.036 TFLOPs), the overall per-image reduction reaches about 48.1 ×.