Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Multimodal Context-Enriched Visual Representation Learning for Enhanced Vision–Language Image Captioning

Image captioning models can produce rapid Sentences, without visual relationships, or insert non-existing plausible objects. A common cause is to compress image evidence into visual symbols that carry a weak neighbourhood context. The multimodal context-enhanced visual representation learning framework (MCVRL) addresses this error mode by adding local neighbourhood descriptors, global scene tokens, prefix-conditioned visual doors and adaptive contextual corrections be-fore caption decoding. The encoder is trained with cross-entropy and contrast terms for image–text alignment. MSCOCO 2014’s Karpathy test classification results show that BLEU-4, METEOR, CIDEr and SPICE are more powerful captioning bases. The best configuration received a score of 1.352 of the CIDEr compared to 1.308 of BLIP-2 in the same evaluation protocol. The results of the ablation show that the most important contribution is the visual feature enriched by the context, followed by crossmodal gating and adaptive contextual attention. Qualitative examples show that objects with hallucinations are fewer and that spatial relationships are better recovered.

E. Divya, Johnson Kolluri, Kiran Siripuri · 0 citations