Analysis of Sycophancy Across Question-ing Styles A Comparative Study of Sycophancy Across Prompting Architectures and Large Language Models
This thesis presents a unified comparative analysis evaluating the robustness of three open-weights instruction-tuned models against a series of adversarial probing strategies spanning social, conversational, and analytical pressure, revealing that modern alignment strategies such as Reinforcement Learning from Human Feedback risk transforming state-of-the-art conversational agents into articulate echo chambers that validate human errors.
Antia Alonso Cancela, Tom Kouwenhoven, Michiel van der Meer
· 0 citations