One Rate Is Not Enough: Adaptive Anisotropic Learning Rates for LoRA Fine-Tuning
Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR convention, applying a single global learning rate across every rank-one component of every adapter. We show that this convention overlooks substantial within-module he...