Skip to content

Author

Tuong-Lan Le van

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

A Deformable Hybrid Transformer With Multi-Axis Strip Attention for Geometric-Aware Medical Image Diagnosis

Accurate diagnosis in medical imaging is often hampered by two intrinsic factors: the complex, anomalous geometric distortion of anatomical structures and the extreme size variability of pathological lesions. Existing convolutional neural networks (CNNs) struggle with global context, while Vision Transformer (ViT) models are limited by quadratic computational costs or reliance on window-based attention mechanisms that disrupt semantic continuity. To address these limitations, we propose MaxStripViT, a novel hybrid architecture that effectively integrates distortion modeling with multi-axis strip attention mechanisms and Local Position-Aware blocks. Our method introduces three key contributions: 1) geometric-adaptive stem (GAS) leverages learnable offsets via Deformable Convolutions (DCNv2) to dynamically align the sampling grid with irregular organ boundaries at the earliest feature extraction stage. It effectively mitigates background noise; 2) position-aware local block is proposed to enhance the Mobile Inverted Bottleneck (MBConv) with Coordinate Attention. This mechanism explicitly models long-range dependencies along spatial axes, improving the precise localization of subtle lesions; and 3) novel sequential multi-axis strip attention mechanism is proposed to replace shift window-based self-attention. The proposed method is experimented on three dataset benchmarks of CT and ChestXray imaging for multi-class lung disease such as CheXtImageNet, IQ-OTH/NCCD, and ChestXray-Image dataset. The results demonstrated that MaxStripViT achieves robust performance and outperforms standard ViT models, and improves performance compared to the most advanced hybrid models currently available, including Swin Transformer, MaxViT, CSWin, ConvNeXt, ConvNeXtv2, and EfficientNet-B7, in terms of classification accuracy and computational efficiency. The proposed method provides a robust solution for health scenarios requiring both geometric flexibility and multi-scale interpretability.

Thanh-An Pham, Tuong-Lan Le van, Van-Dung Hoang · 0 citations