It is found that embedding-based alignment metrics do not reliably indicate whether alignment will improve or degrade downstream performance, and that the compatibility between alignment and downstream objectives should be considered when designing/evaluating alignment methods.
Yana Veitsman, Yihong Liu, Hinrich Schütze· arXiv.org· 0 citations
It is argued that human participation may persist even with highly capable AI systems for three distinct reasons, and this perspective has important implications for the limits of automation and for the design, evaluation, and ethics of future AI systems.
Findings show that successful MoE-block restoration does not necessarily imply localization to a single expert, as it is shown that successful MoE-block restoration does not necessarily imply localization to a single expert.
Yuetian Lu, Ali Modarressi, Yihong Liu et al.· arXiv.org· 0 citations
This work proposes a pipeline for automatically generating step-by-step linguistic reasoning traces from Universal Dependencies treebanks, dictionaries, and grammar-rule banks and shows that linguistic reasoning traces are most effective as inference-time guidance in ICL, which substantially improve translation perform...