Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning
High-quality representations are essential for a wide range of downstream tasks. Dedicated embedding models are explicitly optimized for representation learning, yet their training data are often more limited in scale and diversity than the massive corpora used to pretrain modern large language models and multimodal la...