Skip to content
Book Open access

MCL-SED: A Multimodal CASED-LLaMA Framework for Student Engagement Detection in Virtual Classroom

Oct 2026 · 0 citations

Abstract

Measuring student engagement in virtual classroom remains a major challenge for online education. Aligned with the ICMI 2026 theme of Context and Cultural Awareness for Multimodal Interaction, we address engagement understanding by jointly modeling student behavioral cues from isolated student video feeds and synchronized acoustic streams through a unified multimodal representation. We introduce MCL-SED, built by modifying Emotion-LLaMA-v2 [11] into a high-capacity end-to-end multimodal language model, to encode spatiotemporal visual and acoustic features into a causal language-model backbone. Selective low-rank adapters (LoRA), structured patch concatenation and dropout regularization reduce identity-aware memorization. Dedicated classification and regression heads replace the text-generation objective to jointly predict binary engagement labels and continuous scores. Under a strict student-identity-disjoint evaluation protocol on the CASED dataset, our model achieves a Macro F1 of 0.52 on the classification track and a primary regression score of − 0.78, securing 2nd place on the CASED Challenge leaderboard. Code is available at: https://github.com/mohitcodehub/MCL-SED.git.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.