Toward Multimodal Detection of Body-Focused Repetitive Behaviors Using Vision and Wearable Sensors
Abstract
Body-Focused Repetitive Behaviors (BFRBs), such as hair pulling, skin picking, and nail biting, affect an estimated 3-5% of the population and are difficult to self-monitor because they often occur outside conscious awareness. Habit Reversal Training depends on timely recognition of these behaviors, motivating accessible real-time detection systems. This paper proposes a behavior-centric semantic representation for BFRB detection: a shared 16-dimensional feature space, organized around four behavioral primitives (proximity, contact, motion, and rotation), in which both wrist-sensor and webcam-derived observations can be expressed. Using the Child Mind Institute “Detect Behavior with Sensor Data” dataset (574,945 readings, 81 participants, 18 gesture classes), the best multi-class classifier (XGBoost) achieved a macro-F1 of 0.694 under stratified cross-validation, dropping to 0.610 (95% CI: 0.585-0.636) under stricter Leave-One-Subject-Out (LOSO) validation. Within the shared space, a binary BFRB classifier achieved an ROC AUC of 0.837, and a within-sensor late-fusion experiment combining a proximity/contact branch with a motion-dynamics branch improved macro-F1 from 0.487 and 0.563 to 0.658, indicating that complementary behavioral cues integrate effectively in the shared representation. A preliminary webcam deployment demonstrates real-time feasibility with sub-200 ms latency. Genuine synchronized camera-plus-wearable fusion remains future work.