Contact-rich manipulation remains challenging because successful control depends on physical interaction cues that are often weakly observable from vision alone. Recent tactile world action models jointly model future visual observations and tactile signals to guide action generation, but how such futures should be structured for effective use by the action expert remains underexplored. Directly studying this question with learned world action models is difficult because end-to-end behavior entangles physically invalid visual futures, unreliable predictions, inaccurate or cross-modally inconsistent tactile forecasts, and an unreadable future-to-action interface. To make this interface independently studyable, we introduce Oracle Visuo-Tactile Foresight (OVTF), a controlled framework that supplies paired RGB and tactile futures from successful trajectories verified in simulation. By fixing the future provider, OVTF isolates the interface and asks a cleaner question: if the future is successful and physically executable, what representation allows the action expert to absorb its benefit? Within OVTF, we propose Asymmetric Phase-Local Future Memory (AFM), in which visual memory reads future vision, each tactile memory jointly attends to its own tactile stream and phase-aligned future vision, and cross-tactile access is blocked. We compare AFM with Modality-Isolated Future Memory (IFM), which removes visual-to-tactile access and processes each future modality independently. Across seven tasks on the UniVTAC simulation benchmark, AFM achieves 32.0% average success, compared with 23.7% for IFM and 14.9% for UniVTAC-ACT. This controlled comparison shows that selective phase-aligned visual-tactile routing provides a more actionable future-to-action bridge than complete modality isolation.
Zihang Yao, Chaoyue Ding, Yingying Yu Brigham Young University et al.· 0 citations
Deterministic multiscale gas flow simulations have long suffered from the curse of dimensionality: the number of discrete velocities increases dramatically with the velocity space dimension and the Mach number, exhausting available memory and computational resources. To address this issue, this paper proposes a memory-efficient deterministic method based on an ensemble-of-subproblems strategy using stochastic discrete velocities. This strategy transforms the originally computationally expensive problem into a series of independently and efficiently solvable subproblems. To be concrete, the proposed method replaces the conventional large deterministic velocity set with multiple small random velocity sets. Each random set defines a subproblem, which is solved by a deterministic multiscale numerical scheme that computes macroscopic moments via Monte Carlo integration. The final flow field is obtained by arithmetic averaging over all sub-problems. In this work, we employ the discrete unified gas kinetic scheme (DUGKS) for spatial discretization and term the resulting method SDV-DUGKS. To validate the proposed method, several numerical test cases are conducted, including (a) the one-dimensional shock structure, (b) the two-dimensional cavity flow, and (c) supersonic flow around a square cylinder. The results of the one-dimensional shock structure confirm the feasibility of the proposed method. The two-dimensional cases demonstrate that, compared to its deterministic counterpart, the proposed method saves more than 80% of memory usage while maintaining comparable accuracy. These results indicate that the proposed method markedly reduces memory demand for multiscale flow simulations and exhibits strong potential to alleviate the curse of dimensionality that currently hinders deterministic multiscale numerical schemes from being applied to engineering problems.
Shuyang Zhang, Weidong Li, Ming Fang et al.· 0 citations
An object-centric 3D representation alignment framework built upon $\pi_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training, which enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time.
Zongbo Liu, Shan Jie, Xiaoquan Sun et al.· 0 citations