Skip to content

Author

Hisham Cholakkal

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Preprint Oct 2026

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

This work presents Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter, and is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average.

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly et al. · 0 citations
Aug 2026

Robust Audio-Visual Question Answering with Missing Modality in Training and Testing.

Audio-Visual Question Answering (AVQA) requires reasoning over temporally evolving audio and visual signals to answer natural-language questions about dynamic scenes. Most existing methods assume that both modalities are available during training and testing. In practice, however, an audio or visual stream may be unava...

Jin-Xing Zhou, Zhangbin Li, Di Hu et al. · 0 citations
#computer vision Preprint Sep 2026

Small yet Assistive: Spatially-Aware Post-Training for Low Vision

An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs...

Rishabh C. Choudhary, S. Raj, Umesh Goyal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.