Skip to content
Conference

Unified Accident Anticipation and Explainable Report Generation via Self-Supervised Vision-Conditioned Language Modeling

Jul 2026 · International Conference on Ubiquitous and Future Networks · pp. 599-604 · 0 citations · 14 references

Abstract

Accident anticipation from dashcam videos is essential for advanced driver-assistance systems, yet existing methods often provide limited interpretability beyond a risk score. We propose a unified multi-task framework that jointly performs frame-level incident anticipation, accident category recognition, and structured natural-language reporting. Our approach uses a pretrained self-supervised video backbone to extract temporally aligned representations, enabling estimation of a continuous risk trajectory over time and robust incident-type classification. To generate human-readable explanations efficiently, we condition a pretrained Phi causal language model via a lightweight visionconditioned logit adapter, avoiding costly cross-modal token fusion while maintaining grounding in video context. To further support spatial interpretability and qualitative auditing, we augment the dataset representation with frame-level risky-object bounding boxes and generate object-aware explanation media, including motion saliency overlays, bbox-guided heatmaps, and focus visualizations. On DADA-2000, our method achieves 81.85% AUC and 4.16 s TTA, improving over the strongest prior baseline in our comparison by +9.58 AUC points (from 72.27% to 81.85%) and +0.33 s TTA (from 3.83 s to 4.16 s). These results demonstrate that combining foundation video representations with efficient language conditioning yields both stronger early-warning performance and more interpretable, structured reasoning for driving safety analysis.

View source