Preprint
Aug 2026
UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
U N M ASK is presented, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation, and demonstrates that the discovery and validation stages generalize to reward model preference data.
Chidaksh Ravuru, Shashank Srivastava
· 0 citations