Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Xiaomin Li, Yuexing Hao, Jian Hou et al.· 0 citations
As data volumes and analytical demands grow, traditional data science workflows struggle to meet the need for efficiency, scalability, and reliability. The rapid advancement of large language models (LLMs) has opened new possibilities for AI-powered agents to augment or automate end-to-end data science pipelines—from data exploration and cleaning to modeling, evaluation, and deployment. This emerging paradigm, termed the AI Data Scientist, has gained significant attention in research and industry, yet discussions remain fragmented regarding its integration, evaluation, and real-world impact. This workshop seeks to consolidate these efforts by providing an interdisciplinary forum for presenting cutting-edge research, sharing deployment experiences, and showcasing real-world systems. The workshop will feature invited talks, paper presentations, a demo track, and a panel discussion, aiming to foster community-building and guide responsible development in this rapidly evolving field.
Hao Liu, M. Zitnik, Yong Li et al.· Proceedings of the 32nd ACM...· 0 citations
Machine learning (ML) has seen promising developments in materials science, yet its efficacy largely depends on detailed crystal structural data, which are often complex and hard to obtain, limiting their applicability in real-world material synthesis processes. An alternative, using compositional descriptors, offers a simpler approach by indicating the elemental ratios of compounds without detailed structural insights. However, accurately representing materials solely with compositional descriptors presents challenges due to polymorphism, where a single composition can correspond to various structural arrangements, creating ambiguities in its representation. To this end, we introduce PCRL, a novel approach that employs probabilistic modeling of composition to capture the diverse polymorphs from available structural information. Extensive evaluations on sixteen datasets demonstrate the effectiveness of PCRL in learning compositional representation, and analysis on model uncertainty highlights its potential applicability of PCRL in material discovery.
Namkyeong Lee, Heewoong Noh, Gyoung S. Na et al.· Proceedings of the 32nd ACM...· 0 citations