Skip to content

Author

Jia-Nan Sun

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

ParaCascade: A Parallel Cascading Framework Supporting Early Routing

Real-world inference tasks for large language models exhibit diverse difficulty levels. Existing LLM serving systems integrate models of different sizes and attempt to route tasks of appropriate difficulty to the most suitable model, aiming to reduce resource waste while guaranteeing service quality. Such systems usually adopt a cascading architecture, which performs inference sequentially from lightweight models to heavyweight models and validates outputs until a model that meets the task requirements is identified. However, when handling complex tasks, the cascading architecture inevitably processes unnecessary small models first, leading to cumulative latency and redundant resource consumption. This paper proposes ParaCascade, a parallel cascading framework that supports early routing. The core idea of ParaCascade is to bypass lightweight models and directly route difficult instances to heavyweight model tiers by pre-estimating task complexity, thus avoiding ineffective computation on lightweight models. In addition, ParaCascade adopts parallel prediction and model parallel inference strategies. At the cost of a slight increase in energy consumption, it significantly reduces the systemic latency caused by sequential processing, thereby improving the overall QoS. Extensive evaluations across diverse workloads on the MMLU-pro and MATH benchmarks show that ParaCascade significantly outperforms both single-model deployments and serial inference serving baselines. While maintaining answer quality, it achieves an inference speedup of 1.16× to 1.51×, demonstrating its superiority in efficient LLM serving systems.

Hao Wei, Lujia Yin, Chen Chen et al. · 0 citations
Conference Jul 2026

AdaRAG: Budget-Aware Adaptive Retrieval-Augmented Generation via Hierarchical Reinforcement Learning

Multi-turn retrieval-augmented generation (RAG) improves question answering by decomposing evidence seeking into iterative retrieval and reasoning steps. Existing multi-turn RAG methods usually optimize when and how to retrieve while fixing the number of retrieved documents per step. However, we discovered that this fixed-TopK design is suboptimal: single-hop questions tend to benefit from fewer retrieval rounds with larger per-round evidence sets, whereas multi-hop questions require more retrieval rounds with smaller evidence sets to support stepwise reasoning. To bridge this gap, we introduce AdaRAG, a budget-aware adaptive RAG framework that learns how to retrieve under a hard document budget, including how many retrieval rounds to perform, how many documents to retrieve in each round, and which retrieval source to use. AdaRAG implements this idea with a two-level policy architecture. ModeHead, a lightweight retrieval-mode classifier, selects passage retrieval, graph retrieval, or answer generation; TopkHead, a budget-aware document-allocation classifier, selects a legal TopK after query generation according to the remaining budget. These discrete policy heads are decoupled from language-model token generation, enabling direct reinforcement-learning optimization through hierarchical GRPO after supervised action-format learning. Our experiments across five QA benchmarks demonstrate AdaRAG's good generalization performance under constrained document budgets. In detailed comparisons on HotpotQA, it surpasses the strongest baselines by an average of 10.8 percentage points in Exact Match (EM) and F1 score.

Jia-Nan Sun, Miao Zhang, Chen Chen et al. · 0 citations
Conference Jul 2026

AMSche: Affinity-Aware Microservice Scheduling for Communication-Intensive Tasks

As computing resources in cloud environments become increasingly abundant, executing complex scientific workflows on large-scale cloud infrastructure has become a standard practice. However, communication-intensive workflows face two fundamental bottlenecks. First, the lack of physical topology awareness often forces high-frequency interacting microservices to be placed on geographically distant nodes, which generates excessive cross-node communication overhead, leads to network load imbalance, and increases latency. Second, the prohibitive online computation time of conventional iterative scheduling algorithms further degrades response speed, making them unsuitable for real-time scenarios. To address these bottlenecks, this paper proposes AMSche, a framework for microservice deployment and task scheduling that is aware of both position and topology. The framework comprises two core mechanisms. The first mechanism, position-aware service deployment, colocates high-frequency interacting services on the same physical node based on communication affinity, thereby compressing cross-node communication overhead at the physical level. The second mechanism, topology-aware task scheduling, leverages online topology feature similarity mapping to instantly reuse historical scheduling plans, achieving scheduling decisions at the millisecond level. Extensive experiments on real-world scientific workflow datasets demonstrate that AMSche achieves an average improvement of 16.19% to 39.15% over existing baseline methods in comprehensive metrics including response time, total communication volume, and network load balance.

Hao Wei, Hailiang Chen, Jia-Nan Sun et al. · 0 citations