PLM-Based Two-Stage Semantic Encoding for Task-Oriented Semantic Communication on the Fly
Abstract
Semantic communication (SemCom) moves beyond bit-level transmission by conveying semantic features, and task-oriented communication (TOC) further targets the transmission of task-relevant features to improve communication efficiency. However, existing TOC schemes still rely on high-dimensional deep learning (DL)-based semantic source encoders for each data input, leading to high encoding complexity during communication. This paper proposes a novel paradigm named task-oriented SemCom on-the-fly (TO-SemComFly), which leverages pre-trained language model (PLM) to pre-compute semantic features offline, enabling fast and low-complexity encoding on the fly. TO-SemComFly adopts a PLM-based two-stage semantic encoder, which employs a PLM-assisted coarse (PAC) semantic encoder to retrieve raw semantic features from a PLM-based cache at the $1^{\mathrm {st}}$ stage and a lightweight fine (LF) semantic encoder at the $2^{\mathrm {nd}}$ stage to form coherent and channel-robust semantic features. To achieve low complexity, the PAC semantic encoder leverages Hierarchical Navigable Small World (HNSW) algorithm for graph-based efficient search. Moreover, the LF semantic encoder is designed as a simple encoder with only two layers. Simulation results demonstrate that TO-SemComFly enhances both transmission accuracy and efficiency. Compared with existing TOC schemes, TO-SemComFly achieves at least 63.4% and 51.7% accuracy improvements at low SNR, and 11.5% and 20.7% efficiency improvements in AWGN and Rayleigh fading channels, respectively.