Skip to content

Author

Bernard Ghanem

9 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Preprint Sep 2026

Advancing Video-Text Pretraining with Multi-View Captions

This work proposes a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage, and introduces a granularity-aware text representation with separate CLS tokens for summary and detailed views.

F. M. Thoker, Renaud Vandeghen, Karen Sanchez et al. · 0 citations
Jul 2026

EgoPlay: Event-Triggered Video Editing for Egocentric Streams

We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form"when X happens, do Y,"EgoPlay infers whether and w...

Jinjie Mai, G. Qian, W. Menapace et al. · 0 citations
Preprint Sep 2026

ContextFlow: In-Context Flow Matching for Robot Manipulation

Although highly effective in vision and language domains, applying in-context learning to robotics remains challenging. Existing autoregressive in-context imitation methods discretize continuous actions and exacerbate the accumulation of early prediction errors through next-token prediction, limiting their generalizati...

Jian Ding, Xian-Jie Dai, Roei Herzig et al. · 0 citations
Review Jul 2026

SoccerNet 2026 Challenges Results

Each task and its evaluation protocol is described, the challenge leaderboards are presented, and the leading submissions are summarized, with the aim of documenting the current state of each task as measured on held-out challenge data.

A. Cioppa, Silvio Giancola, Haakan Ardo et al. · 1 citation
Preprint Aug 2026

B-MIM: Biased Masked Image Modeling for Generalizable Segmentation of Fine-Grained Anatomical Structures

B-MIM is introduced, a modification of the iBOT objective that stochastically reduces global semantic alignment to prioritize local patch reconstruction and suggests that reducing global semantic pressure during pretraining enhances generalization to intricate anatomical structures.

S. González, Karen Sanchez, J. M. Saavedra et al. · 0 citations
Jul 2026

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP

This work studies an inference-time substitution of the row-wise softmax in the final visual self-attention layers with the $\alpha$-entmax transform, applied across both the standard query-key attention and self-correlation variants.

Fatima Zohra, Chen Zhao, Shuming Liu et al. · 0 citations

with Pretrained Generative Model for Self-Supervised Learning

GenView is presented, a controllable framework that augments the diversity of positive views leveraging the power of pretrained generative models while preserving semantics, and an adaptive view generation method that dynamically adjusts the noise level in sampling to ensure the preservation of essential semantic meani...

Xiao-Jie Li, Yibo Yang, Xiang-Tai Li et al. · 0 citations
Jul 2026

HyperGS: Fast and Generalizable Gaussian Video Representation

This work proposes HyperGS, a feedforward, optimization-free approach that directly predicts Gaussian representations from any video in a single forward pass, speeding up encoding and decoding by orders of magnitude while generalizing to out-of-distribution videos at higher resolutions.

Fatima Zohra, Chen Zhao, Shuming Liu et al. · 0 citations
Preprint Aug 2026

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

DiSCO is proposed, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals, and can be readily applied to any text-to-image system without necessitating any changes to the model itself.

Tong Zhang, M. Alfarra, Carlos Hinojosa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.