Sep 2026· Lecture Notes in Computer Science· pp. 188-206· 0 citations· 62 references
Computer Science
TL;DR
This work presents FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder--decoder, achieving state-of-the-art results on major benchmarks, including Sintel, KITTI-2015, and Spring.
Abstract
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefined heuristics, which can limit its expressivity and lead to more complex pipelines and additional computational cost. We present FreeFlow, a hierarchical transformer built without any flow-specific components, using instead a single feed-forward encoder--decoder. FreeFlow combines three attention variants: window attention for local processing, shifted-window attention for cross-window information exchange, and a global attention operating at a reduced resolution. The resulting architecture scales naturally with model capacity, enabling a consistent accuracy gain from small to large variants. Despite the absence of standard inductive biases, FreeFlow achieves state-of-the-art results on major benchmarks, including Sintel (0.68/1.48 EPE on Clean/Final), KITTI-2015 (3.23 Fl-all), and Spring (3.192 1px), while remaining memory efficient at 1080p inference.
A training-free looping framework that repeatedly applies selected transformer layers inside each denoising call is introduced, which improves primary and auxiliary quality metrics with competitive quality--efficiency trade-offs across two Scale-RAE model scales.
Yuan-Yi Yan, Xin-Zhe Rao, Can-Yu Shen et al.· 0 citations
DiD is introduced, a label-free conversion method that exclusively trains the linear-attention backbone by aligning detector-facing interface tensors with those of a frozen Softmax teacher, and substantially outperforms established baselines and matches supervised, fully trained linear models.
Huai-Yuan Qin, Gabriel James Goenawan, Zihang Lin et al.· 1 citation
Energy-Guided Flow Matching is introduced that explicitly models a coarse-to-fine generative trajectory by moving endpoint that evolves smoothly from low-frequency image to clean image and requires no adaptation of the backbone and training data.
Accurate stereo matching remains challenging in ill-posed regions such as fine structures, reflective, or transparent objects, where appearance cues are often ambiguous or unreliable. To tackle this, we propose PhasorNet, a lightweight yet powerful framework that boosts geometric discrimination via frequency-domain cue...
Md Raqib Khan, S. Vipparthi, Subrahmanyam Murala· 0 citations
Data augmentation is fundamental to training modern deep vision and multimodal models. While individual methods, such as RandAug, CutMix, Mixup, RandErase, and DropPath, offer strong regularization effects, their combined use has saturated in performance due to overlapping functionalities, and aggressive pixel-level ma...
Hyesong Choi, Daeun Kim, Song Park et al.· 0 citations
Flow-based image editing (FlowEdit) enables inversion-free semantic changes through the difference between source and target velocities. In this paper, we observe that FlowEdit's default classifier-free guidance (CFG) configuration, with asymmetric source and target scales, causes substantial background leakage. Matchi...
Zhe-Yuan Zhan, Can Wang, Jia-Wei Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.