Vision–Language Agent for Spatially Balanced Joint Source and Channel Coding
Abstract
Existing deep joint source-channel coding (JSCC) methods typically employ uniform resource allocation across spatial regions, which may lead to spatially imbalanced reconstruction quality under limited channel resources. To address this, this paper proposes AgentJSCC, a difficulty-aware transmission framework. Specifically, a vision-language model (VLM)-based transmission agent is introduced to predict patch-level reconstruction difficulty, generating a difficulty mask conditioned on the communication setting. Guided by this mask, a region-adaptive JSCC encoder dynamically allocates transmission rates, assigning higher resources to challenging regions before transmission. To further mitigate residual artifacts, we extend this framework to AgentJSCC+, which incorporates a two-step edge-guided denoising network where EdgeNet extracts multi-scale edge priors and DenoiseNet performs edge-guided RGB refinement. Experimental results on AWGN and Rayleigh fading channels demonstrate that AgentJSCC improves reconstruction fidelity over existing JSCC baselines. On DIV2K over the AWGN channel, AgentJSCC improves PSNR over the strongest baseline by 0.87 dB on average, while AgentJSCC+ further increases the gain to 1.77 dB PSNR and an 11.4% relative SSIM gain, demonstrating superior robustness across diverse SNR conditions and compression rates.