Beyond Textual Chain-of-Thought: JEPA-Conditioned Latent Reasoning for Large Audio Language Models
JELAR, a Joint-Embedding Predictive Architecture (JEPA)-based latent reasoning framework that conditions latent reasoning supervision on acoustic representations learned from raw waveforms, demonstrates the effectiveness of JEPA-conditioned latent reasoning as an alternative to explicit textual CoT supervision.