Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models
Reusable Latent Correction (RLC) is proposed, which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls.
Bo-Han Zhang, Li-Nan Yue, Weibo Gao et al.
· 0 citations