This work proposes CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions, and achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse codecs and bitrates.
Abstract
Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and generative codecs that exploit the priors of video generation models. Yet the dominant metrics, LPIPS and DISTS, measure feature and texture similarity rather than content fidelity: a reconstruction that hallucinates a wrong face or blurs text into convincing strokes can still score well, even when a human rejects it instantly. To address this, we propose CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions. We optimize CodecArena with Facet-GRPO, a visual reinforcement learning scheme that aligns pairwise codec preferences while grounding the verdict in five fidelity facets: identity, objects, text, texture, and temporal consistency. Its facet-anchored reward uses automatically derived facet directions as weak anchors, rather than human per-facet labels, to prevent any single sub-score from dominating the holistic preference and to yield interpretable fine-grained quality judgments. To support training and evaluation in this underexplored regime, we construct two complementary resources: CodecArena-1K, a fully automatic preference dataset of 1,500 comparison groups built from traditional, neural, and generative codec reconstructions with fused vision-language and objective supervision; and CodecArena-Bench, a human-ranked benchmark with source-disjoint videos for fair out-of-domain evaluation. Extensive experiments demonstrate that CodecArena achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse codecs and bitrates, surpassing perceptual metrics and prior vision-language evaluators.
Under stringent bitrate constraints, existing video codecs struggle to balance source fidelity and perceptual realism. Distortion-oriented codecs often oversmooth details, while generative codecs risk introducing content and structural deviations and rely on codec-specific designs. This motivates a question: Can existi...
Yin-Huan Huang, Jing-Kai Ying, Pu Chen et al.· 0 citations
Current neural image compression (NIC) methods primarily focus on signal fidelity optimization. While perceptually optimized codecs can generate decoded images that better align with human visual preferences at equivalent bitrates, they raise authenticity concerns due to potential deviations from the original content....
Yu-Dong Mao, Pei-Lin Chen, Hao Luo et al.· IEEE Transactions on Image P...· 0 citations
This work proposes GVCCTurbo, a BPP-driven scheduler that separates expensive prior refreshes from codebook corrections, and supports BPP-to-compute scheduling as a controllable extension of sampler-length tuning, without requiring the allocated point to dominate every boundary point.
Ziyue Zeng, Ding-Jie Peng, Xun Su et al.· 0 citations
Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every system that generates them. Prior work has fr...
Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer et al.· 0 citations
Hybrid semantic video coding combines learned representations with standardized residual coding, but existing approaches either retrain an entire model for each group of pictures or use a fixed decoder that cannot adapt to local content. This paper introduces a context-aware framework that specializes a pretrained sema...
P. Samarathunga, Yasith Ganearachchi, Thanuj Fernando et al.· IEEE Access· 0 citations
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions sep...
Yin-Ming Huang, Shu-Yuan Tu, Xi Yan et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.