Skip to content
Open access

When does heterogeneous structured evidence help? A unified benchmark for structured-data modeling

2026 · AI Plus · 0 citations

Abstract

Real-world entities are often characterized by heterogeneous structured information, including static attributes, temporal dynamics, and relational dependencies. However, existing benchmarks provide limited understanding of whether multimodal models can effectively exploit complementary information across these modalities. In this work, we introduce Trimodal Structured-Data Benchmark (TSDBench), an entity-centric benchmark for evaluating multimodal structured-data learning. To establish a reliable foundation for evaluation, we first verify the multimodal complementarity of benchmark tasks from both empirical and theoretical perspectives: input-ablation experiments demonstrate widespread predictive gains from incorporating additional modalities, while Flow-Partial Information Decomposition (Flow-PID) analysis quantifies the unique, redundant, and synergistic information contributed by different modalities. Furthermore, TSDBench adopts a target-rotation design that enables systematic analysis of cross-modal interactions by allowing different static features or time series signals to alternately serve as prediction targets. Extensive experiments reveal that multimodal gains are highly task-dependent and asymmetric, varying with the target modality and dataset characteristics. Although the proposed fusion models achieve strong overall performance, Flow-PID analysis exposes a substantial gap between the complementary information available in multimodal data and what current architectures can effectively exploit. These findings highlight the need for more adaptive multimodal learning methods capable of selectively leveraging heterogeneous structured evidence.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.