From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
The results show that improving IF in LRMs can significantly enhance privacy, suggesting a promising direction for future privacy-aware LRMs, and introduces an SFT dataset that teaches models to follow general instructions throughout their reasoning process.
Haritz Puerto, Haonan Li, Xudong Han et al.
· 0 citations