Skip to content

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Gaussian-JEPA: Joint-Embedding Predictive Learning for 3D Gaussian Splats

3D Gaussian Splatting (3DGS) represents 3D content with anisotropic primitives that jointly encode geometry and appearance. Fixed-budget encoders consume sampled observations of Gaussian assets, so the same object may be observed through different primitive realizations. Existing self-supervised methods mainly reconstr...

Bin Ren, Qi Ma, Yue Li et al. · 1 citation
Aug 2026

Robust Audio-Visual Question Answering with Missing Modality in Training and Testing.

An AVQA-specific two-stage framework that adapts established cross-modal reconstruction and dependency-modeling principles to the supervision constraints of TM-AVQA is developed, a setting in which modality-complete, audio-missing, and visual-missing samples may occur during both training and testing.

Jin-Xing Zhou, Zhangbin Li, Di Hu et al. · 0 citations
Preprint Sep 2026

Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings enable efficient, reusable search but can miss the transient actions, state changes, and subtle constr...

Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar et al. · 0 citations
Preprint Aug 2026

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

This work introduces CVPD (Contrastive Counterfactual Visual Process Distillation), which is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs, and proposes a three-gate Counterfactual Criterion that identifies visual blind spots where zooming into a region ch...

Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.