Skip to content
Preprint

Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

An injectively aligned leave-one-cell-out protocol over support x operation grids and SO-OPF, a readout that factors cell energy into support salience and a competitive operation posterior, and a readout that factors cell energy into support salience and a competitive operation posterior are introduced.

Abstract

Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can reward multiple operations that reuse the same predicted slot. We call this failure operation laundering. We introduce an injectively aligned leave-one-cell-out protocol over support x operation grids and SO-OPF, a readout that factors cell energy into support salience and a competitive operation posterior. This formulation separates two questions that aggregate scores conflate: whether the carrier composes held-out bindings when the grid is known, and whether that grid can be recovered from flat cell labels. With frozen DINOv3 features, known factorial assignment reaches 0.874 injective accuracy on Shapes3D-Extended and 0.799 on globally image-disjoint COCO; learning the assignment from flat labels reaches 0.769 and 0.762, respectively. Under matched-axis-aware supervision on Shapes3D, the factored carrier improves learned-assignment accuracy from 0.653 to 0.841 over a dense carrier and eliminates its laundering gap. SigLIP2 replicates the COCO separation. A rebuilt MuJoCo substrate exposes a boundary: learned-assignment accuracy is 0.569 with DINOv3 and 0.484 with SigLIP2, with substantial slot collapse. Thus factored readout and injective evaluation recover held-out bindings on two substrates while exposing, rather than hiding, a renderer-specific failure boundary; they do not establish universal recovery from flat labels.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools

Sharing a deep encoder does not, by itself, fix the central confound of task-comparison scores. We show that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, w...

Kentaro Oda · 0 citations
Preprint Sep 2026

UnfoldCRF: Structured Mask Refinement with Image-Conditioned Latent Regions

Learned mask refiners improve segmentation accuracy, but it is hard to tell how much of the improvement comes from explicit structure rather than from extra capacity, and whether it holds up when the mask generator or its error distribution changes. UnfoldCRF treats refinement as inference in a conditional random field...

Chun-Ming He, Rihan Zhang, Lei Xu et al. · 0 citations
Preprint Aug 2026

Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception

This work proposes MCR-GRPO, a marginal contribution assignment framework that derives box-level credit directly from each sampled response, preserving GRPO's response-level comparison while enabling box-aware optimization of structured multi-object grounding.

Xinheng Han, Jianfeng Wang, Yu Chen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

What Changed? Drift Detection with Real, Virtual, and Incomparable Diagnosis

Sharing a deep encoder does not, by itself, fix the central confound of task-comparison scores. We show that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, w...

Kentaro Oda · 0 citations
Preprint Aug 2026

GramLoop: Training-Free Gram-Gated Replay for Robust Dense Prediction

GramLoop is introduced, a training-free framework that replays a short transformer window and controls each replay through final-layer cosine-Gram consistency, which improves object detection and semantic segmentation under corruptions, perturbations, and natural shifts.

Yang Chen, Can-Yu Shen, Xin-Zhe Rao et al. · 0 citations
Preprint Aug 2026

Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for Reinforcement Learning of Vision-Language Models

Noise-Contrastive GRPO is introduced, which injects scale-calibrated Gaussian noise into the last hidden layer of the prompt-encoding pass for half of each rollout group, branching those rollouts from a displaced departure state and integrates into a standard RLVR pipeline as a ~50-line change to the inference engine.

Michael M. Jerge, Joseph Pelczar, J. Downes · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.