APeB: Benchmarking Personalization Ability of Large Language Model Agents
This work introduces personalized product search (PPS), a testbed for agentic personalization under raw queries and diverse histories, and constructs Agent Personalized Benchmark (APeB) from action logs, pairing underspecified intents with rich histories and user-viewed candidate items.
Gary Yang, Zi-Zhe Chen, Xinru Chen et al.
· 0 citations