Skip to content

Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

This work uses empirical studies to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment, and proposes a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models.

Abstract

Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, empirical studies are first used to reveal the existence of a massive, untapped performance headroom for personalized generation through test-time alignment. We demonstrate that personalized generation is uniquely suited for test-time scaling methods like Best-of-N (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While reward models could in principle exploit this headroom, they are poorly calibrated for personalization, and their billion-parameter scale makes scoring large candidate pools prohibitively expensive. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale multi-layer perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing N candidates. Extensive experiments on nine datasets spanning three personalized generation settings show that our personalized ranking model effectively exploits the discovered headroom, outperforming billion-parameter generalist reward models on every dataset, with under 0.4% of their parameters and four orders of magnitude lower scoring latency.

View source

Similar papers

Book Open access Sep 2026

GradSup: Gradient Superposition for Personalised and Scalable LLM Recommendation

Experiments conducted on movie, book and electronics benchmark datasets demonstrate that GradSup outperforms iterative fine-tuning to provide scalable and personalised recommendation that is consistently sustained above the frozen LLM backbone.

Kanishka Dandeniya, C. Dasanayaka, Daswin de Silva et al. · 0 citations
Preprint Aug 2026

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

This paper proposes BALIGN, a balanced data selection strategy that explicitly mitigates catastrophic forgetting while optimizing alignment efficacy, and identifies three key data-centric features that dictate parameter drift: the reference model's log-probability margin, the token length between chosen and rejected re...

Minsu Kim, Jian-Xun Lian, Xing Xie et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Generate to Explore, Select to Exploit: Aligning LLM-based Headline Generation with Personalized Recommendation

GESE (Generate to Explore, Select to Exploit), a framework operating at the system's presentation layer that decouples personalization into generative exploration and selective exploitation, validate that decoupling diversity-oriented generation from precision-oriented selection offers a robust blueprint for aligning g...

Yi Chen, Ru-Feng Cheng, Qiang Xie et al. · 0 citations
Open access Aug 2026

RosePO: Customized Preference Alignment in LLM-Based Recommendation

This work proposes RosePO, a framework to refine LLM-based recommendation through pairwise preference optimization with personalized smoothing, and incorporates a personalized smoothing factor predicted by a user oracle into the optimization objective.

Jia-Yi Liao, Xiang-Nan He, Ruo-Bing Xie et al. · 0 citations
#large language models Review Open access Sep 2026

PALRec: Large Language Model-Based Sequential Recommendation with Parameter-Preserving Augmentation

PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.

Hyunsoo Na, Minseok Gang, Sang-goo Lee et al. · 0 citations
Preprint Aug 2026

TSPORec: Token Selection via Preference Optimization for LLM-Based Sequential Recommendation

This work proposes a novel Token Selection approach for Preference Optimization in LLM-based sequential Recommendation, i.e., TSPORec, which accurately pinpoints informative tokens throughout the entire textual content to improve recommendation performance.

Wenqiao Zhu, Chao Xu, Haipang Wu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.