Skip to content

Author

Pei-Chun Hua

8 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Sep 2026

Unlocking Software-defined GPU Fabric Scheduling in the LLM Era

Large language model (LLM) systems increasingly rely on techniques such as prefill-decode disaggregation, KV-cache offloading, and computation-communication overlap. These optimizations often treat GPU interconnects as best-effort substrates, overlooking contention across shared PCIe, NVLink, and RDMA fabrics. We chara...

Dan-Yang Chen, Yu-Feng Gu, Yibo Huang et al. · 1 citation
#machine learning Preprint Sep 2026

Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG

Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain t...

Su-Ting Chen, Pei-Chun Hua, Yun-Ming Xiao · 0 citations
#artificial intelligence Preprint Sep 2026

Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers

This work proposes a controlled inject-and-remove cycle: deliberately inject a weaker backdoor and then unlearn it, which weakens detector-visible signatures and fools the backdoor detectors with an illusion of purification while preserving the malicious retrieval behavior.

Bei-Ning Xu, Pei-Chun Hua, Yun-Ming Xiao · 0 citations
#machine learning Preprint Sep 2026

Semantics Delivery Network: Rethinking Web Retrieval Infrastructure for LLM Agents

Large language models (LLMs) increasingly rely on external sources when answering questions that require proprietary information or up-to-date live web content, through both traditional single-shot retrieval-augmented generation (RAG) and multi-turn agentic RAG. Yet today's web infrastructure is still built for human c...

Pei-Chun Hua, Yun-Ming Xiao · 2 citations
#artificial intelligence Preprint Sep 2026

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

Retrieval-augmented generation (RAG) depends on dense retrieval: each document is stored as a learned vector, and a query is answered by finding its nearest neighbors in that vector space. Keeping one full-precision vector per document is the dominant index cost at corpus scale, so retrieval systems replace each vector...

Pei-Chun Hua, Yun-Ming Xiao · 1 citation
Preprint Aug 2026

Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills

Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as services while keeping the underlying packages hidden. Prior work focuses on prompt injection attacks that directly disclose these artifacts, and existing defenses accordingly aim to p...

Pei-Chun Hua, Haoxuan Xu, Meng-Yuan Li · 2 citations
#machine learning Preprint Sep 2026

Spruce: Scalable Private Outsourced Retrieval Using Compact Embeddings

Spruce learns compact binary codes that preserve candidates for full-precision reranking, replacing corpus-wide embedding scoring with efficient Hamming-distance computation under two-server multi-party computation under two-server multi-party computation (MPC).

Pei-Chun Hua, Yun-Ming Xiao · 2 citations
Preprint Aug 2026

Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale

This shortlist short-circuits full-corpus cryptographic search without sacrificing retrieval quality: with 200-500 candidates, it closely matches full-corpus retrieval across five zero-shot corpora spanning 25K to 5.4M documents.

Pei-Chun Hua, Dan-Yang Chen, Ju-Nan Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.