Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text...
Guang-Zhi Xiong, Xin-Yuan Zhang, Xiao Yang et al.· 0 citations
WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.
Ji Soo Lee, Xi-Lun Chen, Pierce Chuang et al.· 0 citations
A structured meta-rubric framework that captures the grading criteria at authoring time, and fixed mechanical rules compile it into a flat checklist of binary, machine-gradable checks that an LLM judge scores reliably at evaluation time is instantiated.
Xilun Chen, Zhaleh Feizollahi, Ross Goodwin et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.