Skip to content

Author

Ying-Guang Yang

We have 5 of 22 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

AquaMend: Minimal Re-probing and Conditional Rollback for Latent-Belief Failures in Embodied Agents

Physical changes or sensing errors can invalidate embodied agents'task-relevant beliefs. AquaMend compares re-probing, rollback, and supported continuation on a probe-belief-action graph under an expected-loss objective covering sensing, physical recovery, and uncorrected failures. A joint posterior guides a one-step p...

Yu-Fan Liu, Shang Luo, Yang Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

This work introduces HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations, and proves a detectability bound that converts a target error rate into an explicit probe-size budget, and delimit what probe rotation does and does not buy.

Rong-Xin Yang, Yang Liu, Shang Luo et al. · 1 citation
#natural language process... Preprint Aug 2026

Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit

Hindsight Memory-PRM exploits this audit trail twice: offline to train an operation-conditioned memory-utility critic, and online, where retrievals, citations, and one controlled deletion-and-reanswer per probe settle an intervention-calibrated entry-level presence credit.

H. Jia, Yang Liu, Ying-Guang Yang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

LoopHarness is presented, which restores a persistent, non-decaying safety state at the loop level at the loop level, and gives a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per...

Chenmin Wu, H. Jia, Yang Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.