Skip to content

Distilling Vision-Language Models for On-Device Fire Understanding

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

A teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students is developed, which provides broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.

Abstract

Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene and thus reducing false alarms, yet their large model size makes deployment on embedded fire sensors impractical. In this paper, we study how domain-specialized VLMs can be compressed for fully on-device deployment without losing the safety-critical behavior required for fire detection. We develop a teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students. Experiments across multiple VLM families and model scales show that compact students preserve most of their teachers'fire-understanding capability. We further deploy the distilled models on our commercial Detectium fire detection sensor and jointly evaluate reasoning accuracy, latency, and memory usage. The results show that compression and deployment affect not only accuracy but also model failure modes, with Qwen2.5-0.5B providing the strongest overall deployment trade-off. Our findings provide broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware ev...

Peng-Fei Li, Naufal Suryanto, Si-Cheng Zhang et al. · 0 citations
Preprint Sep 2026

VisionPsy-Nano: Improving Accuracy, Efficiency, and Reliability in On-Device Vision-Language Models

Sub-billion-parameter Vision-Language Models are increasingly viable for on-device deployment, yet compact model size alone does not guarantee usability. On a phone, such a model can still require more than two minutes to produce its first token. On-device usability depends on three axes: accuracy, efficiency, and beha...

Khurram Azeem Hashmi, M. Zolfaghari, Changdae Park et al. · 0 citations
Sep 2026

GRPO-Anomaly: Reinforcement Fine-Tuning Vision--Language Model for Industrial Defects Detection and Reasoning

Industrial anomaly detection (IAD) requires reliable identification and precise localization of subtle defects, yet most existing methods depend on manually tuned decision thresholds and large collections of defect-free samples, limiting scalability in real-world production. To address these constraints, we present gro...

Yue-Ning Li, Hong-Liang Wang, Peng-Fei Xiu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence

We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-...

Chang-Di Yang, Feng-Quan Jiao, Hao-Chih Lin et al. · 0 citations
Preprint Sep 2026

Cloud, Edge, or Split? Profiling Onboard and Split Vision-Language Model Deployment for Drone AI

This paper benchmarks the performance trade-offs among fully onboard, cloud-based, and split-computing architectures for lightweight VLMs using SmolVLM-256M as a representative lightweight VLM and shows that no deployment strategy is universally optimal.

Zoha Azimi, Reza Farahani, S. Dustdar et al. · 0 citations
Preprint Sep 2026

SPARK: Representation-Level KV Memory Alignment for Safer Vision-Language Models

SPARK is introduced, a two-stage framework for targeted KV-memory repair that reduces multimodal attack success while preserving general capability and suggests that multimodal jailbreak behavior can be substantially mitigated by selectively repairing intervention-relevant KV subspaces at prefill.

Mohd. Azfar, I. Khan · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.