Skip to content
Book Open access

Enabling AI to Find Multimodal Signals of Storytelling Cross-Culturally: A Multi-Step Approach

Oct 2026 · 0 citations · 14 references

Abstract

Storytelling is a fundamental human practice through which people share experiences and make sense of the world. But while conversational storytelling itself has been extensively studied, little is known about the multimodal cues that signal the emergence of storytelling in spontaneous interaction and whether these cues are shared across cultures. The EU-co-funded project Storytel addresses this gap by building a corpus of naturally occurring conversations which feature stretches of embedded storytelling, and which is collected across several European countries. Students from four different universities record videos of multiparty interactions with their smartphones capturing speech, sign and other bodily actions, such as gestures, gaze, posture, facial expressions and practical actions. By building on manually identified storytelling sequences, the project investigates whether machine learning methods can automatically detect such sequences and determine their interactional structure. By using a customized processing pipeline, we extract lexical, acoustic, and visual features and integrate them into an interactive visualization tool with fully anonymized videos, enabling users to explore multimodal features. We discuss the challenges of manually identifying storytelling sequences, developing the feature extraction pipeline, and applying comprehensive image and voice anonymization. The paper combines a cross-cultural multimodal case study with the evaluation of AI-based methods for analyzing culturally situated interaction, contributing to multimodal communication research and supporting the development of AI systems capable of recognizing storytelling in human interaction.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.