Low-cost fine-tuning of SmolVLA with smoothness regularization and chunk-aware auxiliary supervision for language-conditioned manipulation
Vision-Language-Action (VLA) models offer a unified interface for language-conditioned manipulation. This paper reports a negative result: a low-cost post-training recipe for SmolVLA with LoRA, a temporal smoothness regularizer, and chunk-aware auxiliary supervision from action-variation pseudo labels does not improv...