When a fact changes, how should a language model update the history stored in its key-value (KV) cache? Hiding the old record is cheap, but it may still contain needed details or answer questions about the past. We compare hiding whole records, hiding only replaced values, and deleting old text and recomputing the cach...
Chang-Hai Zhou, Yu-Hua Zhou, Shi-Yang Zhang et al.· 0 citations
As Large Language Models (LLMs) evolve, parallel reasoning has emerged as a vital inference paradigm that enhances robustness by concurrently exploring multiple thought trajectories. Unlike fragile sequential methods, parallel reasoning expands inference breadth to significantly improve problem-solving performance. T...
Zi-Qi Wang, Bo-Ye Niu, Zi-Peng Gao et al.· National Science Review· 0 citations
Model merging plays a crucial role in consolidating multiple specialized models into a single, unified model, especially in the era of large language models (LLMs). Recent research has primarily focused on developing strategies to enhance merging performance with the trained models, while the impact of training paradig...