Skip to content

Author

Jia-Wei Shao

We have 5 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference

Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-orient...

Lu-Ning Pang, Cheng Yuan, Jia-Wei Shao et al. · 0 citations
Preprint Sep 2026

Information Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality

Under the AI Flow framework, communication networks distribute intelligence across devices, edge servers, and clouds, and computation at the receiver becomes a resource that can substitute for transmitted bits. Generative video compression (GVC) embodies this exchange by sending compact tokens with ultra-low bitrate an...

Cheng Yuan, Jia-Wei Shao, Xue-Long Li · 0 citations
#artificial intelligence Preprint Sep 2026

RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse

Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk interactions. Existing me...

Ruo-Ling Qi, Yi-Rui Liu, Xuan'er Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs

Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundar...

Yi-Rui Liu, Ruo-Ling Qi, Xuan'er Wu et al. · 0 citations
Preprint Jul 2026

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenating their KV entries, and selectively recomputing a few tokens to res...

Yi-Rui Liu, Ruoling Qi, Long-Wen Wang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.