The introduction of Markov-Chain JEPA (MCJEPA), which replaces the latent predictor by a learned transition matrix, and Markov-Chain JEPA (MCJEPA), which replaces the latent predictor by a learned transition matrix, give temporal JEPA a principled state-space interpretation.
Abstract
A hidden Markov model (HMM) combines three roles: inference of a hidden-state belief from observations, propagation through a Markov transition, and emission back to observation space. We show that full, time-indexed Predictive Information Bottleneck VJEPA (PIB-VJEPA) exposes the same computational structure: a stochastic context encoder plays the role of an amortized filtering distribution, a probabilistic predictor defines latent-state dynamics, and a decoder, inverse target encoder, or induced implicit conditional supplies the emission direction. We distinguish 4 progressively stronger levels of correspondence and give sufficient conditions for exact sequence-level HMM equivalence. To make the connection concrete, we introduce Markov-Chain JEPA (MCJEPA), which replaces the latent predictor by a learned transition matrix; in the finite time-homogeneous case, matrix powers guarantee exact multi-horizon Chapman--Kolmogorov consistency. Conditioned discrete-state transitions, continuous-state Markov kernels, and continuous-time dynamics extend this construction, while deterministic temporal JEPA appears as a degenerate Dirac-kernel special case. We further interpret predictive information-bottleneck learning as seeking a compact predictive state: compression promotes minimality, while residual predictability tests sufficiency. Controlled experiments support transition composition, the filtering interpretation, predictive Markovization in a known synthetic process, and the distinction between JEPA latent prediction and HMM-style sequence learning. Together, these results give temporal JEPA a principled state-space interpretation.
We present a three-tier progression of Hidden Markov Models: maximum-likelihood (Baum-Welch), variational Bayesian (VBEM), and a neural variational extension (Neural VBEM), that model each customer's transaction history as a trajectory through a small number of latent behavioural regimes, one of which is empirically id...
Joseph Uririoghene Obukofe, A. O'Hare, Chioma Sandra Dike· 0 citations
This work develops a filtering and optimal-control framework for partially observable stochastic systems in which each observation identifies a class of a finite measurable partition of the hidden state space, and proposes class-dependent finite-dimensional approximations capable of preserving both continuous component...
Saul Díaz-Infante Velasco, Yofre H. García, J. Minjárez‐Sosa· 0 citations
An information-theoretic framework for generalization in next-token prediction under temporally dependent data is developed, revealing the roles of context length, model complexity, sample size, margin, and temporal mixing and yielding informative guarantees for deterministic algorithms over continuous hypothesis space...
Recurrent models must preserve information that changes future behavior while suppressing hidden-state error. These objectives conflict: contraction improves stability, but contraction along a future-distinguishing direction destroys memory. We formalize this boundary through the predictive quotient of a recurrent stat...
SLP-ProbHard is introduced, a cross-family, representation-centered framework for probabilistic hard-constrained learning when explicit structural parameterizations are available, and how feasible coordinates and maps affect stochastic dimension, dependence, calibration, expressiveness, and computation is studied.
It is concluded that exploration is better described as a dynamic state than as a fixed trait, and that modeling it as a switching process is both more accurate and useful.