Skip to content
Open access

Low-cost fine-tuning of SmolVLA with smoothness regularization and chunk-aware auxiliary supervision for language-conditioned manipulation

Oct 2026 · Frontiers in Neurorobotics · 0 citations · 33 references

Abstract

Vision-Language-Action (VLA) models offer a unified interface for language-conditioned manipulation. This paper reports a negative result: a low-cost post-training recipe for SmolVLA with LoRA, a temporal smoothness regularizer, and chunk-aware auxiliary supervision from action-variation pseudo labels does not improve manipulation success under a 30,000-step budget (approximately 1/67 of the original 2,000,000-step SmolVLA recipe). Because the auxiliary classifier is used only during training and the runtime chunk length remains fixed at k = 8, we denote the configuration SmolVLA+LoRA+S&C (smoothness and chunk-aware auxiliary losses) rather than “adaptive.” We evaluate four LIBERO suites with a unified protocol of 500 episodes per cell (50 episodes × 10 task variants). Over three training seeds, the matched fixed-chunk k = 8 baseline reaches a four-suite macro success rate of 10.9%±1.1, whereas SmolVLA+LoRA+S&C reaches 3.5%±1.4; the pre-specified primary comparison shows significant deficits on LIBERO-Spatial (−16.4 percentage points, exact sign-flip permutation p = 0.004) and LIBERO-Object (−17.5 pp, p = 0.002), with matched-pairs rank-biserial effect sizes of −0.96 and −1.0. Matched component ablations isolate the smoothness loss as the harmful term (smoothness-only collapses performance; chunk-classifier-only approximately matches the baseline), and executed-action logs show the regularizer triples trajectory jerk (1.69 vs. 0.59 RMS)—the opposite of its intent. A training-free inference-time temporal-ensembling baseline, by contrast, more than doubles success on the two suites where the baseline is competent (Spatial 36.2 vs. 19.0%; Object 56.2 vs. 24.0%) at a 2.3 × step-rate cost, and a 10–100K budget sweep shows the S&C deficit never closes. We conclude that, for the tested configuration, budgets, and benchmark, the added action-side losses are not competitive with a well-tuned fixed chunk, and inference-time ensembling is the stronger use of the same compute.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.