Large-scale video diffusion models (V-DMs) have achieved remarkable text-to-video generation quality, yet their massive computational complexity makes deployment costly. Post-Training Quantization (PTQ) offers an appealing route to accelerate inference without retraining, but existing diffusion PTQ methods remain fragi...
Wei-Lun Feng, Chuan-Guang Yang, Haotong Qin et al.· IEEE Transactions on Pattern...· 3 citations
This work presents an energy-proportional, context-aware vision IoT node that addresses this challenge through a heterogeneous multimodal dual-camera architecture, enabling always-on visual monitoring in a place-and-forget scenario through autonomous edge intelligence.
Julian Moosmann, P. Mayer, Luca Benini et al.· 0 citations
Experiments on captioning, question answering, and unseen-action generalization show that mmMind consistently outperforms existing radar-language baselines, while ablations confirm the importance of pose-guided pretraining.
Duo Zhang, Zhe-Hui Yin, Zhi-Yun Yao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.