Skip to content

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Sep 2026 · 0 citations · 21 references
Computer Science

TL;DR

WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.

Abstract

Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versus cross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt a dual-grounding framework that combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-source LLMs demonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.

View source

Similar papers

Review Open access Aug 2026

Wearable-Derived Digital Biomarkers in Preventive and Personalized Medicine: Promise, Evidence, and Barriers to Clinical Translation

Wearable technologies now permit near-continuous measurement of physiological and behavioral parameters under free-living conditions. Combined with advances in artificial intelligence (AI), these devices support the development of wearable-derived digital biomarkers that may shift healthcare from a reactive to a preven...

Damilola Alabi, Anyebe Daniel Ameh, D. Okon · 0 citations
Review Open access Aug 2026

WEARABLE HEALTH TECHNOLOGIES AND THE MEDICALIZATION OF EVERYDAY LIFE: A SOCIO-MEDICAL REVIEW

Consumer wearable technologies have evolved from simple activity trackers into multisensor systems that continuously measure physiological and behavioural signals, generating algorithmic assessments of cardiac rhythm, sleep, recovery, stress, and metabolic responses. This trajectory increasingly blurs the boundary betw...

M. Sierant, Malwina Kwaśniewska, Kamila Zmysłowska et al. · 0 citations
Review Sep 2026

Digital Biomarkers and Wearable Bioelectronics Across Neurological, Cardiopulmonary, and Mental Health Conditions: Clinical Evidence and Barriers to Adoption

A narrative synthesis of clinical, materials-science, and computational evidence on wearable-derived digital biomarkers across neurological, cardiopulmonary, inflammatory, wound-care, and mental-health conditions is conducted, drawing on peer-reviewed studies identified through the source literature and organized accor...

Unknown authors · 0 citations
#machine learning Preprint Sep 2026

HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care

As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients and clinicians to understand health trends, detect safety-critical events, and make informed decisions. While large language models (LLMs) show promise for transforming...

Yu-Chen Niu, Ya-Nan Ma, Srinivasan Nandakumar et al. · 0 citations
#machine learning Preprint Aug 2026

Learning Human Health and Diseases from 24-hour Wrist Movement

These findings establish 24-hour wrist movement as a rich and scalable source of health information, with the potential to support passive health monitoring and disease prediction at population scale, and establish Sensori, a self-supervised foundation model that learns general-purpose health representations directly f...

Yong Wang, D. McGagh, K. Broomberg et al. · 0 citations
#federated learning Review Open access Sep 2026

Decision-centered wearable biosensors for personalized rehabilitation: integrating multimodal monitoring, artificial intelligence, and closed-loop intervention

Personalized rehabilitation requires repeated assessment of movement, physiological tolerance, fatigue, and adherence, yet conventional clinic-based measurements provide only intermittent snapshots. The convergence of wearable biosensors with artificial intelligence, edge-cloud computing, and connected data infrastruct...

Fu-Jin Jia, Xin-Zhu Li, Hui Song et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.