The results suggest that global temporal context and explicit player-position cues are useful for generating more detailed soccer commentaries, while also revealing remaining trade-offs in precise timestamp localization.
Chen-Yi Xu, Yi-Hao Wu, Li-Qi Yan et al.· Scientific Reports· 0 citations
Sentiment analysis of multimodal social media data is of great importance, not only for recognizing objective information but also for capturing subjective emotional states. While single-modal sentiment analysis has achieved notable progress, existing multimodal approaches still face two key challenges: (1) inadequate...
Guo-Guo Ye, Qi-Qi Chen, Li-Qi Yan et al.· Big Data and Cognitive Compu...· 0 citations
Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow representational capacities of lightweight text encoders when processing lengthy, terminology-dense clinica...
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening ge...
Yihao Wu, Chen-Yi Xu, Li-Qi Yan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.