Some attention heads learn similar patterns across inputs. Reusing these patterns could reduce training cost by avoiding repeated query-key score computation and softmax. Through controlled pretraining comparisons, we identify Selective Attention Freezing (SAF), which selects heads with low attention-pattern variance a...
Weixian Waylon Li, Yin-Tao Tai, Marcio Fonseca et al.· 0 citations
Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...
FinCAD is proposed, an inference-time adaptation of Context-Aware Decoding that attenuates contributions from memorised historical outcomes without retraining and raises the subset-averaged in-sample/out-of-sample Spearman correlation on an eleven-model leaderboard.