Can Revealed Preferences Clarify LLM Alignment and Steering?
An empirical pipeline is presented for estimating the implied preferences that an LLM's observed choices optimize: the model's probability distribution over unknowns is elicited along with the choice it would make for the decision task and a discrete choice model is fit to recover the cost function that best rationaliz...