Skip to content

Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

Sep 2026 · 0 citations
Computer Science

TL;DR

This study empirically confirms the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.

Abstract

As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes ("swing"), yet stubbornly ignore explicit instructions to correct ingrained biases ("rigidity"). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core. Second, how can these values be quantified? We propose the Prior-Environment-Cognition (PEC) framework. This model mathematically defines value expression as the joint outcome of inherent dispositions like parameter weights (Prior), external contexts such as user prompts (Environment), and internal reasoning processes like Chain-of-Thought (Cognition). Finally, how can LLMs'values be aligned toward a desired target? Using PEC diagnostics, we establish an adaptive"Alignment Prescription". Rather than blindly applying resource-intensive training, this method identifies the minimum effective intervention needed for each dimension, ranging from zero-cost prompts to targeted parameter updates. Extensive empirical validation confirms that our approach successfully verifies the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.

View source

Similar papers

#small language model Open access Sep 2026

Measuring structural value alignment in sixteen models: LLMs have human-like moral spaces

Evaluating whether large language models (LLMs) reason about morality in human-like ways requires more than measuring whether they produce the right outputs in isolated cases. Existing approaches – including scalar agreement, distributional analysis, rationale classification, and consistency testing – compare model and...

Ye-Xiang Tang · 0 citations
#natural language process... Preprint Aug 2026

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities, these datasets fundamentally conflate causal mech...

Davood Wadi, Mohsen Ghodrat, Matthew Philp · 2 citations
#machine learning Preprint Sep 2026

From Constitutions to Control: Interpretable Rewards for Aligning Language Models

Current approaches to aligning language models often make it hard to know what behavior is being rewarded or to change that reward in a targeted way. In particular, standard preference-based methods collapse multiple considerations into aggregate human judgments, obscuring what drives the resulting reward, while princi...

Johann D. Gaebler, C. Isley, Max Lamparth et al. · 0 citations
#natural language process... Preprint Aug 2026

Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments

A systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view is conducted, highlighting the risk of treating LLM judgments as faithful proxies for human belief updating and point to structural differences in how LLMs and huma...

Lin Chen, Yi-Tong Chen, Yong Li · 0 citations
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations
#natural language process... Preprint Sep 2026

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and thre...

Rong Wang, Kun Sun, Ya-Dong Guo · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.