PersonaMem-v3 is introduced, a real-world-grounded benchmark and evaluation harness for omni-platform personal intelligence that evaluates whether AI agents can infer holistic user understanding from cross-platform evidence, personalize responses, rerank recommendations on social media, follow user steering through nat...
Bo-Wen Jiang, Yuan Yuan, Zhuo-Qun Hao et al.· 0 citations
A hypergraph-based paired failure attribution (HPFA) framework that attributes the failure root cause by comparing the hyperedges of the targeted failure reasoning path against a reference successful path is proposed, and the trained attributor consistently improves reasoning accuracy at test time.
Runchuan Zhu, Hong Dung Lai, Bowen Jiang et al.· 1 citation
MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Xiaomin Li, Yuexing Hao, Jian Hou et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.