Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong
Suramya R. Angdembay, Dikshant Aryal, Nick Rahimi
· 0 citations
2 papers indexed here
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Auditing black-box (behavioral) detection of unfaithful CoT against FaithCoT-Bench's human annotations, the authors find answer correctness structures the problem at every level.