Skip to content

Author

Hai-Bo Hu

We have 7 of 45 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models

Activation steering offers an inference-time defense for vision--language models (VLMs) by modifying intermediate representations without updating backbone parameters. However, protection on benchmark inputs may not persist across alternative expressions of the same harmful request. We investigate this gap using fixed...

Xin-Wei Zhang, Ao-Ting Hu, Hang-Cheng Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents

Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...

XiaoYu Xu, Zi Liang, Min-Xin Du et al. · 0 citations
Preprint Sep 2026

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

Text-to-image diffusion models enable data-efficient"mimicry"attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are br...

Hao-Yang Li, Ruo-Xi Sun, Qing-Qing Ye et al. · 0 citations
#machine learning Preprint Sep 2026

TraceGuard: Adaptive Multimodal Poison Filtering through Cross-Feature Rank Agreement

Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We theref...

Hao-Yang Li, Ya-Xin Xiao, Lin-Yan Dai et al. · 0 citations
Preprint Sep 2026

The Shape of Ownership: Verifying LLM Provenance through Semantic Structures

As large language models (LLMs) are increasingly redistributed, adapted, and served behind opaque APIs, model ownership can no longer be established reliably by inspecting model internals or deployment records. This creates a need for behavioral signatures that remain observable through black-box interaction. Yet most...

Zhong-Rui Sun, Jia-Hao Chen, Ou-Bo Ma et al. · 0 citations

SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning

SEBA is a sample-efficient framework for black-box adversarial attacks on visual RL agents that significantly reduces cumulative rewards, preserves visual fidelity, and greatly decreases environment interactions compared to prior black-box and white-box methods.

Tai-Ran Huang, Yu-Lin Jin, Jun-Xu Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?

This work proposes Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent-agnostic defense middleware that matches the detection performance of state-of-the-art defenses at lower token overhead than LLM-as-a-judge baselines.

Zi Liang, XiaoYu Xu, Yanyun Wang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.