StellAR: A Multimodal Generative AI Framework for Context-Aware Augmented Reality in Stem Education
Abstract
Although Augmented Reality (AR) has shown great promise for use in education, the reliance on 3D modeling expertise to develop AR educational content has hindered its widespread adoption and implementation in STEM, or science, technology, engineering, and mathematics classrooms. This paper introduces a context-aware immersive learning framework, called StellAR, which automatically transforms static educational PDF files and textbooks into interactive AR learning sessions, without any programming or 3D modeling skills from educators. At the heart of StellAR is a Multimodal Retrieval-Augmented Generation (RAG) pipeline driven by Gemini 1.5 Pro, which validates semantic domains and extracts keywords from uploaded documents. A hybrid Asset pipeline that routes the requests according to the cosine similarity between the concepts, and fetches pre-compiled GLB assets from a vector store (cos 0.85) otherwise fetches a new concept through a generative path using 2D reference images and a ShapeVAE-based implicit Signed Distance Function (SDF) representation using Hunyuan3D 2.1, a flow-matching Diffusion Transformer for geometry-consistent 3D model generation. An Intelligent Tutoring System (ITS) is embedded in the system and provides adaptive audio explanations in real time, starting right at the time of context extraction, even though 3D assets are loaded asynchronously - making passive AR viewing into active guided problem-solving. With a “Custom Classroom” module, teachers can select AR lessons without technical expertise. The system has been tested on the retrieval path with 50 retrieval tests and has achieved an AR plane detection accuracy of 93.2%, a semantic analysis latency of below 200 ms and Time-to-Interaction of less than 2.2 seconds. StellAR is proof that generative AI can become the solution that eliminates the content creation barrier to scale in AR for STEM learning.