Data and Knowledge Dual-Driven Text Embedding Approaches: A Survey
Abstract
Text embedding has emerged as a pivotal technique in natural language processing, facilitating the effective understanding and processing of textual information by machines. With the continuous advancement of data-driven methods like large language models (LLMs), text embeddings have become richer and of higher quality. However, researchers have identified limitations in purely data-driven methods, which may lack interpretability and logical consistency. Conversely, purely knowledge-driven methods require extensive manual effort from experts to design rules, leading to low efficiency. To overcome these challenges, researchers have explored integrating data-driven and knowledge-driven methods, termed Data and Knowledge Dual-driven Text Embedding (DKDTE). In this paper, we introduce a novel taxonomy categorizing existing text embedding methods into three primary categories, namely, knowledge-driven approaches, data-driven approaches, and dual-driven approaches. We provide formal definitions of text embeddings with distinctions in input granularity, a dedicated overview of application tasks, evaluation benchmarks (including MTEB, BEIR, and AIR-Bench), and real-world applications. We offer a comprehensive comparison of representative methods across categories and discuss the latest advances including LLM-based embedding methods, instruction-tuned embedding paradigms, and multimodal knowledge integration. We also identify promising future research directions, including debiasing, exploring diverse knowledge sources, and data contamination mitigation.