ReflectRL is proposed, a lightweight plug-and-play framework that learns from Golden Negative Trajectories during on-policy training, and first uses these trajectories to elicit Reflective Reasoning, then applies Reflective-to-Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning...
Jin-He Bi, Chennan Zhou, Zeng-Jie Jin et al.· 2 citations
DriveCache is proposed, a training-free, action-aware controller that uses planned motion to allocate reuse across scenes and dynamic programming to place it across denoising steps under a calibrated response budget, which improves the overall fidelity-efficiency trade-off over evaluated cache methods.
Jianchun Yang, Jian Liang, Xian-Da Guo et al.· 0 citations
Federated Graph Learning (FGL) enables privacy-preserving GNN training over distributed graph data, yet dynamic task streams in Federated Graph Continual Learning (FGCL) inevitably lead to catastrophic forgetting. From a spectral perspective, this forgetting manifests as two fundamental challenges: high-frequency incon...
Hanyao Guo, Zihan Tan, Wen-Ke Huang et al.· Proceedings of the 32nd ACM...· 0 citations
It is suggested that LLMs differ not only in baseline risk disposition, but also in the risk signals they respond to and the flexibility with which they adjust, providing a behavioural basis for auditing risk-sensitive decision-making in interactive settings.
Xuan-Kun Rong, Wenke Huang, Bo Du et al.· arXiv.org· 0 citations
Federated Learning (FL) has demonstrated a promising future in privacy-friendly collaboration but it faces the data heterogeneity problem. Knowledge Distillation (KD) can serve as an effective method to address this issue. However, challenges arise from the unreliability of existing distillation methods in multi-domain...
Yue Yuan, Wenke Huang, Frank Wan et al.· Neural Information Processin...· 1 citation
TrustVLA is introduced, a mechanism-guided inference-time defense that adapts the Dirichlet evidence framework from trusted classification to monitor per-token, per-layer epistemic uncertainty in VLA policies, providing a retraining-free, mechanism-guided defense for visual-triggered VLA backdoors.
Pin-Han Fu, Xian-Da Guo, Xue-Tao Li et al.· arXiv.org· 0 citations
Switch-Reasoner is proposed, a GRPO-based framework that learns to adaptively select reasoning modes for MLLMs and introduces a dual-level regulation mechanism that balances the overall use of Thinking Mode and Direct Mode while providing sample-level supervision based on the relative benefit of the two choices.
Yiyang Fang, Pei Fu, Jinjie Li et al.· arXiv.org· 0 citations
DecisionQE is introduced, a questionnaire-based framework for measuring each model's persuasive and compliant tendencies across multiple decision scenarios, and the Werewolf game is used as an interactive testbed to study their effects on social influence and group outcomes under asymmetric information.
Wen-Wen He, Wen-Ke Huang, Wei Yang Bryan Lim et al.· 0 citations
The core of SSVAL is Visual Anchor Prompt Injection (VAPI), which introduces prompts that absorb rich knowledge from external VFMs during training, enabling them to serve as stable visual anchors that mitigate representation deviation during inference.
Qian-Long Yang, Bowen Ye, Xianda Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.