Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models
This study empirically confirms the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.
Kelvin Zhang, Jing-Yu-Gin Chen, Yu-Fan Liu et al.
· 0 citations