Skip to content

Author

Jia-Hao Chen

We have 6 of 29 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

The Shape of Ownership: Verifying LLM Provenance through Semantic Structures

As large language models (LLMs) are increasingly redistributed, adapted, and served behind opaque APIs, model ownership can no longer be established reliably by inspecting model internals or deployment records. This creates a need for behavioral signatures that remain observable through black-box interaction. Yet most...

Zhong-Rui Sun, Jia-Hao Chen, Ou-Bo Ma et al. · 0 citations
Jul 2026

Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks

It is revealed that Bit-Flip Attacks (BFAs) can serve as an attack vector for inducing decision-level hijacking, requiring no real-time interaction or control over the training process, and only a minimal number of weight bits need to be flipped after deployment to achieve stealthy, low-cost, and persistent cognitive m...

Yu Yan, Jia-Hao Chen, Si-Qi Lu et al. · 0 citations
Preprint Sep 2026

A Finger on the Scale: Covert Policy Steering through Agentic Skills

Reusable agent skills extend large language model (LLM) agents with task procedures, tool-use guidance, and output constraints. Yet these skills also act as externalized behavioral policies, which create a supply-chain risk: a third-party skill may preserve the declared task and valid output interface while covertly re...

Jia-Rui Li, Jia-Hao Chen, Chun-Yi Zhou et al. · 0 citations
Jul 2026

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

This work forms this problem as backdoor generalization under training--inference trigger shift and introduces Lilith, a black-box anchor-to-family framework that achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap.

Zhou Feng, Jia-Hao Chen, Chun-Yi Zhou et al. · 0 citations
Preprint Aug 2026

Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds

AEGIS (Adaptive Ensemble Guard for Injection Shielding) extracts instruction-sensitive projectors to identify malicious instructions and leverages a Unified Multi-Layer Consensus mechanism that aggregates topologically distinct signals across the network depth.

Jia-Hao Chen, Ruiping Yin, Xin-Feng Li et al. · 1 citation · ⚡1
Book Open access Aug 2026

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails

Unsafe Semantic Distillation is proposed, which aligns adversarial perturbations with distributional representations of unsafe content rather than prompt-specific instances, and achieves 84% attack success rates, outperforming existing methods and exposing fundamental vulnerabilities in current multimodal safety archit...

Shuo Shi, Ruiping Yin, Na-En Xu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.