Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reli...
Zhen-Ting Huang, Bo-Han Jiang, J. Liu et al.· 0 citations
Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offering either limited fine-grained control...
Knowledge editing provides an efficient way to update factual knowledge in large language models. However, malicious edits may introduce safety risks, making it necessary to reverse undesirable editing effects. Existing reversal methods for parameter-modifying edits mainly focus on global removal, which may also erase...
Wei-Feng Jiang, Rui-Rui Chen, Qian-Ren Mao et al.· 0 citations
We propose LLM-PeerReview, an unsupervised LLM Ensemble method that selects the most ideal response from multiple LLM-generated candidates for each query, harnessing the collective wisdom of multiple models with diverse strengths. LLM-PeerReview is built on a novel, peer-review-inspired framework that offers a transpar...
Zhijun Chen, Zeyu Ji, Qianren Mao et al.· arXiv.org· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.