Unsupervised Cross-Modal Hashing Algorithms for Web Multimedia Retrieval
Abstract
With the Web witnessing a rapid surge in multimodal data, it’s becoming increasingly vital to develop efficient and budget-friendly cross-modal (CM) retrieval techniques to elevate the user experience in Web applications. Traditional hashing methods, however, often neglect the valuable semantic information hidden within the text descriptions that come with Web images. Moreover, they tend to lean heavily on supervised learning, which poses a challenge when it comes to adapting to real-world Web scenarios where annotations are often in short supply. To tackle this, this research introduces an unsupervised CM hashing algorithm for Web multimedia retrieval. By mining the semantic structure of text associated with Web images and utilizing a deep network to achieve semantic transfer from text to vision, a unified and efficient hashing learning framework is constructed. Experiments indicate that the introduced approach achieves mAP values of 0.3370 and 0.6990 with 16-bit hash codes (HCs). When the HC length is increased to 128 bits, the mAP increases to 0.3632 and 0.7575, representing an absolute improvement of 13.0% and 4.04% compared to the best performing baseline method. Further analysis shows that the semantic transfer mechanism significantly improves the semantic representation ability of the HCs. Even in a semi-supervised setting using only 20% of labeled data, the retrieval mAP can still reach 0.8871. The method requires only a portion of image-text pairs during the training phase and supports pure image queries during the retrieval phase, achieving millisecond-level response times under Hamming distance calculation. This provides an efficient and practical solution for Web-scale multimedia retrieval.