Current neural image compression (NIC) methods primarily focus on signal fidelity optimization. While perceptually optimized codecs can generate decoded images that better align with human visual preferences at equivalent bitrates, they raise authenticity concerns due to potential deviations from the original content. Therefore, achieving controllable decoding is crucial in various applications. This study presents a novel plug-and-play framework that leverages large vision-language model (LVLM) priors to balance fidelity and perception for existing NICs. Our approach consists of two key components: a scalable Low-Rank Adaptation scheme to controllably enhance the semantics of initially decoded images, and a two-stage agent-assisted decoding strategy with vision-language priors utilization. Specifically, the first stage extracts textual semantic information from an LVLM using decoded images enhanced by flexible fidelity-perception decoding, while the second stage effectively integrates semantic priors from LVLMs, further mitigating decoding semantic uncertainty and achieving higher-quality decoding. Extensive experiments on multiple benchmark datasets demonstrate that our method enables off-the-shelf NICs to achieve flexible control between optimal perceptual quality and signal fidelity.
Yudong Mao, Peilin Chen, Hao Luo et al.· IEEE Transactions on Image P...· 0 citations
Unsupervised Domain Adaptation (UDA) is essential for adapting object detection systems to diverse operational environments without requiring domain-specific labeled data. In parallel, real-time performance is crucial for deploying these detectors in intelligent vehicles. This paper introduces Uncertainty-Guided Adaptive Knowledge Distillation (UGAKD), a novel framework designed to enhance UDA while simultaneously reducing model size through targeted knowledge distillation. Given that adversarial learning is a common approach in UDA, often utilizing a domain classifier to identify domain-invariant features, UGAKD leverages the localization of these domain-invariant features to guide the distillation process. Furthermore, we propose a two-stage, difficulty-aware training scheme to facilitate learning, which emphasizes domain-invariant features to boost distillation efficacy. Experimental results across several challenging scenarios, including transitions from synthetic to real-world environments, varying weather conditions, and shifts between real and stylized domains, demonstrate that UGAKD effectively reduces model complexity while improving detection accuracy. Specifically, UGAKD decreases the number of parameters by over 37% and FLOPs by over 47%. Compared with the baseline fine-grained feature imitation method, UGAKD achieves an mAP improvement of 1.0% to 1.7%, highlighting its effectiveness in maintaining robust object detection across diverse settings and its suitability for applications that require both efficiency and adaptability.
Wei-Lun Tseng, Yi-Lun Wu, Yung-Hui Li et al.· ACM Transactions on Intellig...· 0 citations
The framework is the first, to the best of the authors’ knowledge, to integrate explainable permutation importance with dynamically reweighted TPE sampling for constrained electric-machine optimization, simultaneously enhancing feasibility, performance and parameter diversity.
Yi-Rong Lin, Chang-Yi Kao, Nien-Yi Jan et al.· Compel· 0 citations