Vision-language retrieval with CLIP-style dual encoders achieves strong cross-modal performance, yet practical accuracy often hinges on localized semantic distinctions where top-ranked near misses differ from the true match by a single critical detail. Hard-sample mining can select confusable candidates but cannot cons...
Siyi Liu, Xiao-Rong Zhu, En-Jun Du et al.· 0 citations
Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning essential for practical deployment. Existing pruning methods retain tokens by attention saliency, yet we show that saliency is systematically misaligned with ranking contrib...
Siyi Liu, Han-Jun Yang, Chen-Chen Zhang et al.· 0 citations
Multi-objective ranking serves as the backbone of industrial information retrieval, requiring a holistic assessment of documents across dimensions such as Relevance, Authority, and Recency. The prevailing industry paradigm relies on ensembles of specialized BERT-based models, which are costly to maintain and fundamenta...
De-Zhi Ye, Junwei Hu, Xiaoyang Chen et al.· Proceedings of the 32nd ACM...· 0 citations
Real-world image search queries are multimodal and compositional: ``find this shirt in pink''specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits...
Enjun Du, Siyi Liu, Zi-Rong Chen et al.· 1 citation
These findings suggest that, within the pointwise scoring paradigm, routing continuous relevance semantics through discrete text constrains ranking signal resolution reveals a bottleneck that is stable and difficult to overcome under current standard methods, rather than an easily resolvable training bias.
Xiaoyang Chen, Jie Liu, Haijin Liang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.