Skip to content
Open access

Event-Driven Multimodal Sensing and Computing for Context-Aware Home Monitoring Using Stereo Vision and Dietary Event Anchoring

Jul 2026 · Italian National Conference on Sensors · Vol 26, pp. 4803 · 0 citations · 46 references
Medicine

TL;DR

An event-driven multimodal sensing and computing framework for context-aware home monitoring using stereo vision and dietary event anchoring is proposed and demonstrates the system-level feasibility of transforming irregular domestic observations into structured, temporally indexed, and privacy-aware multimodal behavioural records for future home monitoring applications.

Abstract

Real-world home monitoring requires sensing systems that can capture daily behaviour without continuous raw-video retention or excessive user burden. However, domestic environments present irregular activity timing, fragmented human presence, asynchronous multimodal events, and privacy-sensitive data management. This study proposes an event-driven multimodal sensing and computing framework for context-aware home monitoring using stereo vision and dietary event anchoring. The framework integrates stereo RGB-based three-dimensional human motion sensing, dining-zone-triggered meal image acquisition, runtime event orchestration, timestamp-based cross-modal synchronization, privacy-aware local storage, and large-language-model-assisted dietary context interpretation. Instead of continuously recording all sensor streams, the system activates and organizes sensing through human presence detection, debounce logic, cooldown-based session control, and dining-zone occupancy events. Meal-related events are used as contextual anchors to associate motion sessions and dietary observations into synchronized behavioural episodes. The prototype was deployed for 11 consecutive days in a real kitchen–dining environment, with the stabilized real-time monitoring phase evaluated from 11 to 14 February 2026. During this phase, the system generated 26 event-driven motion sessions and 51,165 captured pose frames, of which 25,925 were valid. Sustained active sessions accounted for 30.8% of all sessions but contributed 81.5% of captured pose frames, indicating that event-driven orchestration concentrated motion data within behaviourally meaningful activity windows. Eight meal-related records were obtained, seven of which overlapped with motion sessions, resulting in 87.5% meal-event overlap coverage. Structured pose outputs required approximately 550 kB/min, corresponding to about 33 MB/h of recorded pose data. LLM-assisted meal-image interpretation achieved a mean absolute percentage error of 25.44%, supporting its use for coarse dietary-context description rather than precise nutritional quantification. However, this result is interpreted only as evidence for coarse dietary-context description and not as validation of a precise nutritional or clinical dietary assessment method. These results demonstrate the system-level feasibility of transforming irregular domestic observations into structured, temporally indexed, and privacy-aware multimodal behavioural records for future home monitoring applications.

Read PDF

Similar papers

Open access Sep 2026

A Real-World Smart Home Dataset Integrating Sleep, Environmental, Physiological, and Ambient Sensing for Homecare Research

Modern homecare research increasingly relies on multimodal sensing technologies in smart homes to monitor daily routines, sleep, environmental conditions, and physiological activities over long periods. However, publicly available datasets often lack real-world longitudinal tracking, multimodal integration, and detaile...

Raja Omman Zafar, Yves Rybarczyk, Sandra Saade et al. · 0 citations
Review Open access Aug 2026

From Inertial to Ambient: A Systematic Review of Sensor Modalities and Data Acquisition Strategies in Human Activity Recognition

This systematic review provides a comprehensive examination of the full spectrum of sensor modalities employed in HAR research spanning wearable inertial sensors, physiological sensors, vision-based sensors, depth cameras, ambient sensing systems, and emerging multimodal fusion frameworks tracing, technical characteris...

O. A. Ayegbusi, M. Onyesolu · 0 citations
Open access 2026

Edge-Intelligent Wearable IoT for Real-Time Stress Monitoring and Indoor Localization: A TinyML-Enabled Adaptive RPL Approach

An edge-intelligent wearable IoT system that integrates photoplethysmography-based sensing, edge machine learning (TinyML), adaptive networking based on the Routing Protocol for Low-Power and Lossy Networks (RPL), and zone-aware indoor tracking is proposed, representing a validated proof-of-concept toward preventive he...

Hariprasath Madhalingam, Naganathan Meyyappan Ramesh, Aadhil Ahamed Jaffarullah et al. · 0 citations
Open access Sep 2026

Smart Furniture-Embedded Bed-Interaction and Multimodal Room Sensing with Sleep-Aware Hierarchical Fusion for Control-Oriented Bedroom State Recognition

Highlights What are the main findings? A bed-centered multimodal sensing framework was developed by integrating furniture-embedded bed-interaction sensing with bedside IEQ, room-context, and time features to recognize eight control-oriented sleep-related bedroom states. The proposed SA-HMF model achieved strong subject...

Xin-Ying Ru, Yu-Mei Bai, Mao-Xu Wang et al. · 0 citations
Open access Sep 2026

Adaptive TinyML with Shift-Aware Routing for Human Activity Recognition Under Distribution Shifts

Human activity recognition (HAR) on wearable and mobile devices requires models that can operate with limited computational resources while remaining reliable under changing real-world sensing conditions. However, most TinyML-based HAR systems follow the same computational path for every input, regardless of whether an...

Hilal Akarkamçı, Mahmut Kılıçaslan · 0 citations
Open access Aug 2026

A Real-Time Communication Framework for Distributed Wearable Human Activity Recognition

A distributed HAR framework in which five wearable sensors are associated with local embedded nodes that perform acquisition, windowing, preprocessing, and convolutional neural network–long short-term memory (CNN–LSTM) inference.

Jhonathan L. Rivas-Caicedo, Laura Saldaña-Aristizábal, Kevin Niño-Tejada et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.