Skip to content
Open access

An improved Inception Transformer with Bi-Level Routing Attention and multi-scale feature collaborative enhancement for Chinese ancient architecture image classification

Sep 2026 · npj Heritage Science · 0 citations

Abstract

Chinese ancient architecture embodies China’s cultural heritage but remains challenging to classify because of spatially dispersed key components, high inter-class similarity, and multi-scale visual characteristics. To address these challenges, we propose BDM-ViT, an improved Vision Transformer model based on the Inception Transformer. Bi-Level Routing Attention enhances the long-range semantic interactions among highly relevant semantic regions through region-level sparse routing, while a Dual-Dimension Feature Collaborative Enhancement module jointly refines channel and spatial features. A Multi-Kernel Feed-Forward Network captures fine-grained textures and global structurals, and a Joint Discriminative Loss improves classification through boundary optimization, hard-sample learning, and intra-class feature constraints. We also construct the Chinese Ancient Architecture Dataset(CAAD), containing 18 architectural categories. Comparative experiments show that BDM-ViT significantly outperforms state-of-the-art convolutional neural networks and Vision Transformer models, achieving 95.46% accuracy with superior precision, recall, and F1-score. Ablation experiments and visualization analyses further verify the effectiveness and complementarity of each proposed module.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.