Low-cost fine-tuning of SmolVLA with smoothness regularization and chunk-aware auxiliary supervision for language-conditioned manipulation
Abstract
Vision-Language-Action (VLA) models offer a unified interface for language-conditioned manipulation. This paper reports a negative result: a low-cost post-training recipe for SmolVLA with LoRA, a temporal smoothness regularizer, and chunk-aware auxiliary supervision from action-variation pseudo labels does not improve manipulation success under a 30,000-step budget (approximately 1/67 of the original 2,000,000-step SmolVLA recipe). Because the auxiliary classifier is used only during training and the runtime chunk length remains fixed at k = 8, we denote the configuration SmolVLA+LoRA+S&C (smoothness and chunk-aware auxiliary losses) rather than “adaptive.” We evaluate four LIBERO suites with a unified protocol of 500 episodes per cell (50 episodes × 10 task variants). Over three training seeds, the matched fixed-chunk k = 8 baseline reaches a four-suite macro success rate of 10.9%±1.1, whereas SmolVLA+LoRA+S&C reaches 3.5%±1.4; the pre-specified primary comparison shows significant deficits on LIBERO-Spatial (−16.4 percentage points, exact sign-flip permutation p = 0.004) and LIBERO-Object (−17.5 pp, p = 0.002), with matched-pairs rank-biserial effect sizes of −0.96 and −1.0. Matched component ablations isolate the smoothness loss as the harmful term (smoothness-only collapses performance; chunk-classifier-only approximately matches the baseline), and executed-action logs show the regularizer triples trajectory jerk (1.69 vs. 0.59 RMS)—the opposite of its intent. A training-free inference-time temporal-ensembling baseline, by contrast, more than doubles success on the two suites where the baseline is competent (Spatial 36.2 vs. 19.0%; Object 56.2 vs. 24.0%) at a 2.3 × step-rate cost, and a 10–100K budget sweep shows the S&C deficit never closes. We conclude that, for the tested configuration, budgets, and benchmark, the added action-side losses are not competitive with a well-tuned fixed chunk, and inference-time ensembling is the stronger use of the same compute.