Large Language Models are increasingly used for embedding extraction. In fact, there are many approaches that try to optimize the embedding representations that these models can learn, exploiting the knowledge gained from extensive pre-training and large parameter counts. However, most works currently focus on the English language and textual input only, reflecting the trend of current Large Language Model training corpora. Recently, several Large Vision-Language Models, which are Large Language Models capable of processing multimodal signals in input, have been released. Yet, their training procedure still remains predominantly based on English data. This limitation also affects the evaluation step, with embedding benchmarks that provide limited coverage for low-resource languages. To address these challenges, we adapt a Large Vision-Language Embedding Model trained on English multimodal tasks to support multilingual inputs. Furthermore, we introduce a new benchmark to evaluate the multilingual and multimodal capabilities of embedding models.
CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models, and argues that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Malvina Nissim, Danilo Croce, V. Patti et al.· Italian Journal of Computati...· 0 citations