Skip to content

Author

Zhongxiang Dai

We have 5 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling

Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential interaction under limited information. We...

Wang-Xuan Fan, Xiao-Yu Nie, Zhou-Tian Shi et al. · 0 citations
Book Open access Aug 2026

CES: Combinatorial Experts Selection via Contextual Linear Bandits

With the rapid advancement of large language models (LLMs), multi-agent systems have emerged as a promising alternative to scaling up a single model. Existing approaches ensemble multiple LLMs to improve response quality, but they often rely on static prior knowledge of model capabilities and prompts, and require exten...

Jinkun Xu, Minghan Wang, Zhiyong Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce \textbf{COBRA-Skills}, an efficient framework that formulates skill optimization as bud...

Ping-Chen Lu, Xiang-Yi Wang, Xiang Li et al. · 0 citations
Preprint Aug 2026

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispers...

Zi-Kun Qu, Min Zhang, Ming-Ze Kong et al. · 4 citations

Meta-Prompt Optimization for LLM-Based Sequential Decision Making

The EXPonential-weight algorithm for prompt Optimization} (EXPO) is proposed to automatically optimize the task description and meta-instruction in the meta-prompt for LLM-based agents and is extended to additionally optimize the exemplars (i.e., history of interactions) in the meta-prompt to further enhance the perfor...

Ming-Ze Kong, Zhiyong Wang, Yao Shu et al. · 7 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.