Skip to content

ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies

Sep 2026 · 0 citations · 18 references
Computer Science

TL;DR

ActSafeGuard is introduced, a differentiable and training-aligned safeguard layer for flow-matching based policies that integrates hard action feasibility into policy learning, not merely treating safety as an inference-time external component.

Abstract

Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step guarantees or correct unsafe actions only during inference, creating a mismatch between policy training and execution. We introduce ActSafeGuard, a differentiable and training-aligned safeguard layer for flow-matching based policies. ActSafeGuard integrates hard action feasibility into policy learning, not merely treating safety as an inference-time external component. Through an analytical ray-scaling operator design, ActSafeGuard enables boundary-aware gradients to guide the model to naturally learn constrained manifolds. Extensive experiments on multiple standard foundation backbones ($\pi_{0.5}$ and Fast-WAM) across various tasks demonstrate that ActSafeGuard consistently achieves a $100\%$ step safety rate while fully preserving or even boosting task success rates, providing a scalable and minimally invasive solution for safe embodied AI deployment.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models

Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existing fine-tuning methods provide limited safety guarantees. Current approaches primarily rely on Lagrangian optimization that enforces safety through soft penalties on expected cumulative cost, often re...

Manan Tayal, A. Nambi · 0 citations
#small language model Preprint Sep 2026

MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?

Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, but the model capacity actually required by the benchmark remains unclear. We introduce MINERVA (MINimal Efficient Robotic Vision-Action policy), a family of deliberately compact visuomotor policies designed...

Kohei Sendai, T. Matsushima, Yusuke Iwasawa · 3 citations · ⚡1
#small language model Preprint Aug 2026

SafeBranch: Branch-Pair Safety Alignment for Embodied Agents

SafeBranch is proposed, a framework that aligns an embodied actor on safety through branch pairs constructed from the actor's own unsafe rollouts via environment rollback, achieving roughly ten times more safe successes than the untrained baseline on the unseen-object variant.

Hyunse Lee, Jiwoo Jeong, Haneul Lee et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies

Every valid trajectory can teach what physically happened, while the actor is deployed only under the condition associated with relatively better actions, in \method, a direct world-action policy that combines outcome-agnostic predictive learning with outcome-aware policy improvement.

Bo-Tong Zhao, Fangjie Yu, Tim Yu et al. · 0 citations
Preprint Sep 2026

SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models

Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both unsafe and already-safe generations. This always-on intervention can unnecessarily perturb the model's original reasoning path and degrade g...

Cao-Yuan Ma, Tian Gu, Wen-Pu Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

Flow-matching Vision-Language-Action (VLA) models have emerged as a potential solution for generalist robot control, designed by combining a pretrained Vision-Language Model (VLM) backbone with an action expert that generates continuous robot actions. While these models exhibit impressive capabilities, due to their ver...

Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.