A teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students is developed, which provides broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.
Abstract
Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene and thus reducing false alarms, yet their large model size makes deployment on embedded fire sensors impractical. In this paper, we study how domain-specialized VLMs can be compressed for fully on-device deployment without losing the safety-critical behavior required for fire detection. We develop a teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students. Experiments across multiple VLM families and model scales show that compact students preserve most of their teachers'fire-understanding capability. We further deploy the distilled models on our commercial Detectium fire detection sensor and jointly evaluate reasoning accuracy, latency, and memory usage. The results show that compression and deployment affect not only accuracy but also model failure modes, with Qwen2.5-0.5B providing the strongest overall deployment trade-off. Our findings provide broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.
Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware ev...
Peng-Fei Li, Naufal Suryanto, Si-Cheng Zhang et al.· 0 citations
Sub-billion-parameter Vision-Language Models are increasingly viable for on-device deployment, yet compact model size alone does not guarantee usability. On a phone, such a model can still require more than two minutes to produce its first token. On-device usability depends on three axes: accuracy, efficiency, and beha...
Khurram Azeem Hashmi, M. Zolfaghari, Changdae Park et al.· 0 citations
Industrial anomaly detection (IAD) requires reliable identification and precise localization of subtle defects, yet most existing methods depend on manually tuned decision thresholds and large collections of defect-free samples, limiting scalability in real-world production. To address these constraints, we present gro...
We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-...
Chang-Di Yang, Feng-Quan Jiao, Hao-Chih Lin et al.· 0 citations
This paper benchmarks the performance trade-offs among fully onboard, cloud-based, and split-computing architectures for lightweight VLMs using SmolVLM-256M as a representative lightweight VLM and shows that no deployment strategy is universally optimal.
Zoha Azimi, Reza Farahani, S. Dustdar et al.· 0 citations
SPARK is introduced, a two-stage framework for targeted KV-memory repair that reduces multimodal attack success while preserving general capability and suggests that multimodal jailbreak behavior can be substantially mitigated by selectively repairing intervention-relevant KV subspaces at prefill.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.