Skip to content

Author

Lourdu Mahimai Doss

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Cross-Modal Transformer Networks for Unified Intelligence in Multi-Sensor IoT Systems

The emergence of multi-sensor Internet of Things (IoT) deployments has led to an immediate demand to have unified intelligence models that can take advantage of heterogeneous data modalities, such as time-series sensor readings, visual data, audio streams as well as contextual metadata, to form incoherent understanding and decision-making. The conventional methods treat each of the modalities separately, which does not reflect the deep cross-modal associations that are necessary to understand a scene holistically and act appropriately in a context. The proposed paper suggests a new cross-modal transformer architecture network operating on the principle of joint representation learning of various sensor modalities in order to achieve the unification of intelligence in multi-sensor IoT systems. The framework presented is based on modality-specific encoders and then a set of cross-modal attention mechanisms which allow two way flow of information between sensor streams to capture the highly complex inter-modal dependencies without paired training data. The new hierarchical fusion approach is a combination of the local cross-modal interactions and global contextual reasoning where the model can adaptively weigh the contributions of the sensors depending on the environmental conditions and tasks. The architecture requires built-in adaptive modality gating mechanisms, which ensures that performance does not suffer even in case of a failure of individual sensors or in case of poor quality. Cross-modal evaluation on three multi- sensor IoT benchmark datasets indicates that the proposed cross-modal transformer is 96.8% accurate in unified perception tasks, which is 18.4% and 11.2% better than single-modality evaluation and other traditional fusion methods, respectively. The framework exhibits high levels of robustness to sensor dropouts of up to 40 percent and is also able to generalize to unobservable sensor layouts. The results have made cross-modal transformer networks a disruptive paradigm of unified intelligence in heterogeneous multi-sensor IoT systems.

Selvin Pradeep Kumar S, Lourdu Mahimai Doss, L. S et al. · 0 citations