Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) optimization offers a forward-only alternative with inferenc...
DASH-Q is proposed, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares, which outperform other PTQ baselines in ultra low-bit regime and improves zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines.
Jaemin Kim, Sungkyun Kim, Junyeol Lee et al.· EuroMLSys@EuroSys· 0 citations
Reported Speculative decoding (SD) speedups are difficult to compare because methods are commonly evaluated with different serving runtimes and configurations. We present SpecLLM, a pluggable evaluation framework that applies common scheduling, batching, KV-cache, CUDA Graph, and execution-backend policies across metho...
Sungkyun Kim, Jaemin Kim, Yeongpil Cho et al.· Electronics· 0 citations
NeuRIT is proposed, a Neuron-guided Robust Instruction-Tuning framework built on a localization-first perspective that mines context-aware neurons associated with relevant and irrelevant context processing, and uses them as anchors to selectively adapt both the identified neuron groups and the layers in which they conc...
Jae Lee, Jaemin Kim, Sumyeong Ahn et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.