Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data. First, we devise an irregularity-aware multi-scale encoder to capture sparse clinical evidence at diverse temporal scales. Then, we propose a temporal evidence distiller to integrate representations across these scales and compress them into a small number of LLM-compatible tokens. Moreover, we introduce a progressive alignment strategy that sequentially aligns the irregular trajectories with the LLM's textual embedding space. To facilitate training, we construct 30,000 clinical time series paired with multi-scale descriptions, together with 41,000 instruction-tuning instances spanning 11 tasks. Using a 4-billion-parameter LLM backbone, ClinPRISM achieves state-of-the-art performance on the held-out evaluation benchmark while using only 16 time-series tokens and achieving an average inference latency of 0.15 seconds per question.
Frank Nie, Ethan B. Liu, Yuan Zhu et al.· 1 citation
As data volumes and analytical demands grow, traditional data science workflows struggle to meet the need for efficiency, scalability, and reliability. The rapid advancement of large language models (LLMs) has opened new possibilities for AI-powered agents to augment or automate end-to-end data science pipelines—from data exploration and cleaning to modeling, evaluation, and deployment. This emerging paradigm, termed the AI Data Scientist, has gained significant attention in research and industry, yet discussions remain fragmented regarding its integration, evaluation, and real-world impact. This workshop seeks to consolidate these efforts by providing an interdisciplinary forum for presenting cutting-edge research, sharing deployment experiences, and showcasing real-world systems. The workshop will feature invited talks, paper presentations, a demo track, and a panel discussion, aiming to foster community-building and guide responsible development in this rapidly evolving field.
Hao Liu, M. Zitnik, Yong Li et al.· Proceedings of the 32nd ACM...· 0 citations
CLINLENS is introduced, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms, which exposes a substantial gap between runnable submissions and correct clinical analyses.
Yuan Zhu, Ethan B. Liu, Frank Nie et al.· 0 citations
CLIR-Bench is introduced, a benchmark for irregular clinical time series QA constructed from de-identified ICU records through a principled four-stage pipeline, enabling evaluation of both answer accuracy and evidence use.
Frank Nie, Ethan B. Liu, Yuan Zhu et al.· 2 citations