Recent visual document retrieval (VDR) systems such as ColPali use multi-vector page embeddings, in which patch-level vectors enable fine-grained evidence matching but incur substantial index storage and MaxSim scoring overhead. Post-hoc merging offers a practical route to efficient VDR by reducing this cost without re...
Recent advances in artificial intelligence (AI) are reshaping smart glasses from egocentric capture and display devices into platforms for wearable intelligence. Smart glasses increasingly serve as wearable AI systems that connect first-person observation with real-time assistance under strict form-factor constraints....
Xu Yuan, Yi Wang, Zhuohang Jiang et al.· 0 citations
Large language model (LLM)-empowered recommender systems have emerged as a promising paradigm for generative recommendation, leveraging their strong semantic reasoning and generative capacity to model complex, diverse user preferences. However, most existing approaches rely on an autoregressive paradigm that is subopti...
Chengyi Liu, Yong-Qi Zhou, Junwei Pan et al.· arXiv.org· 0 citations
Fashion is a knowledge-intensive domain in which effective decision-making depends on integrating multiple types of knowledge. Although Large Language Models (LLMs) have transformed many areas, their application in fashion remains limited by hallucinations and weak domain specialization. Knowledge Graph (KG)-based Retr...
Yujuan Ding, Linyin Luo, Shijie Wang et al.· 0 citations
A three-stage investigation framework is established to empirically study the inference design space of multimodal pLMs across three representative pLMs and four fundamental tasks, and to derive conclusions about base models that differ from prior consensus.
Yi Zhou, Qiping Wang, Yunqing Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.