In-vehicle network intrusion detection via traffic-to-image encoding and transfer learning-based visual transformers
Abstract
In complex communication environments, in-vehicle networks face severe security challenges. To address the limited multi-scenario coverage, imbalanced data distribution, and high training cost of existing in-vehicle network intrusion-detection methods, an intelligent intrusion detection system based on transfer learning and optimized visual Transformer models is proposed. The proposed approach encodes controller area network (CAN) messages and external network traffic features into a unified image representation, thereby facilitating the extraction of structural and textural differences among diverse attack patterns in the visual space. For model construction, three advanced visual Transformer architectures, namely DeiT-Small, Swin-Tiny, and Vision Transformer (ViT-S), are selected as base learners and initialized with ImageNet pretrained weights to enhance feature representation under limited data conditions. Furthermore, Bayesian optimization is employed to automatically search key hyperparameters during the transfer learning process, thereby improving training stability and generalization performance. On this basis, Majority Voting and Confidence Averaging are adopted as ensemble learning strategies to fuse predictions from multiple models, which further improves the robustness of intrusion detection. Experimental results on two representative datasets, Car-Hacking and the CICIDS2017 dataset, covering in-vehicle networks and vehicle external communication scenarios, demonstrate that the proposed method consistently outperforms multiple mainstream approaches in terms of Precision, Recall, F1-score, and Accuracy, while maintaining favorable computational efficiency. These results support the effectiveness and practical applicability of the proposed method in complex in-vehicle network intrusion detection scenarios.