Leveraging Large Language Models (LLMs) for generative recommendation has attracted significant research interest, where item tokenization is a critical step. It involves assigning item identifiers for LLMs to encode user history and generate the next item. Existing approaches leverage either token-sequence identifiers...
Xin-Yu Lin, Chuan-Bo Zhang, Yu-Fan Liu et al.· ACM Transactions on Recommen...· 0 citations
FinDeepIndicator is proposed, the first benchmark dedicated to evaluating Deep Research agents in end-to-end financial indicator construction, and covers fundamental, technical, and macroeconomic indicators organized into 21 fine-grained sub-categories.
Chaoqun Yang, Fengbin Zhu, Xinyu Lin et al.· 0 citations
DASH is a decision-aware user simulator that jointly generates thinking traces and predicts behavioral actions from heterogeneous cross-domain histories and tailors a rubric-based reward model that evaluates thinking traces along form, content, and logic for RL training.