In non-native English classrooms, multimedia resources often produce a cognitive gap of high input but low output: visual contexts and target-language forms fail to form stable semantic connections. To address this problem, this study proposes a three-stage multimedia teaching strategy of context anchoring, modal verification, and output transfer. First, a micro-context video library is constructed using word-vector semantic association, with target vocabulary anchoring the semantic field while 12-second clips control cognitive load. Second, a semi-structured virtual dialogue agent dynamically triggers recast and clarification feedback according to learners’ spoken output, grammar breaks, and pragmatic deviations. Third, digital narrative reconstruction tasks require learners to convert video-based semantic information into written storyboards, promoting cross-modal syntactic transfer. Experiments show that the semantic priming effect size increases to 145.2 ms, speech-flow break density decreases to 12.3, and average clause nesting depth rises to 1.89. The framework turns multimedia into a cognitive scaffold and is compatible with networked multimedia transmission, wireless interactive classrooms, and electromagnetic-safe digital learning spaces.
Juan Wu, Yingying Du, Xiaohui Zhang et al.· Advanced Electromagnetics· 0 citations
Physics problems in textbooks are typically presented as static diagrams accompanied by brief textual descriptions, requiring learners to infer dynamic physical behaviors through mental visualization. This process often imposes high cognitive demands and limits learners'ability to form accurate mental models. In this paper, we present \textbf{LivePhys}, a framework that enables a \emph{Scan-to-Play} paradigm for mechanics learning by transforming static textbook physics problems into executable, interactive simulations. LivePhys decouples multimodal perception from physics-aware reasoning and deterministic simulation. Given a problem diagram and its accompanying text, LivePhys performs text extraction, geometric segmentation, and cross-modal grounding to construct a structured, physics-aware intermediate representation. A multimodal large language model is then used as a reasoning controller to infer entities, parameters, and constraints, which are executed by a physics engine to generate spatially consistent and interactive simulations that allow learners to explore and manipulate problem conditions dynamically. Our evaluation results show that LivePhys significantly outperforms general-purpose multimodal models in simulation executability, spatial accuracy, and interaction fidelity. In addition, a user study demonstrates that interacting with LivePhys-generated simulations reduces learners'perceived cognitive load compared to static textbook materials.
Xiaowei Dai, Ziyu Luo, Xiangwen Zhang et al.· 0 citations