Collecting health narratives at scale: A multi-study evaluation of speech-to-text in health surveys
Abstract
Patient narratives can reveal aspects of health that standardized measures miss, yet collecting open-ended responses at population scale remains difficult. We evaluated whether speech-to-text (STT) could enable richer responses in online health surveys without the resource demands of interviews. Across four surveys involving people with post-COVID-19 condition, healthy adults, people who engage in sex work, and older adults, participants could answer open-ended questions by typing or using STT. We compared response length, lexical diversity, stop-word proportion, and total unique content words, alongside STT uptake and participant feedback. Quantitative comparisons were made using a mixed-effects model including modality (STT vs typed) as fixed effect and a random intercept per participant. STT uptake varied markedly across populations (0.7-80.0%). Spoken responses were estimated to contain more characters than typed responses in two studies, by 245.0 and 891.8 characters on average (both p<0.001), and contained more unique content words, despite lower lexical diversity and higher stop-word proportions. Participants generally valued STT for its convenience and spontaneity, but uptake was constrained by context, privacy, interface design and individual preferences. STT may enable population-scale collection of richer health narratives, provided implementation is tailored to the study population and research context.