This work proposes Explainable Probing of Cross-Domain Sparse Embeddings (EXPOSE), a framework that uses Sparse Autoencoders (SAEs) as an explainable bottleneck to identify and suppress domain-specific components in VFM embeddings.
Anja Witte, M. Lennartz, Jan Baumbach et al.· 0 citations
This work presents a deep learning-based web system for automatic identification of Bangladeshi mango varieties and integrated the model into a Streamlit web application that enables users to upload a mango image and receive a predicted variety with class probabilities.
Video-level AUC here is thus a composite of event evidence and pre-event source cues, a shared source of discrimination that can obscure differences between representations.
This work proposes several architectural changes to the BSF, including a Tournament Top-K selection rule that significantly reduces feature splitting, and extends the block paradigm to the crosscoder.
This work formalizes probabilistic alignment as a distributional criterion for world models and introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics, and introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors.
Yuandong Pu, Le Zhuo, Sayak Paul et al.· 0 citations
LiveVVT is introduced, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation, and a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference.
Yushe Cao, Shikun Feng, Ru-Xiang Duan et al.· 0 citations
A hybrid GAN-guided diffusion framework that uses a pretrained Wasserstein GAN with gradient penalty (WGAN-GP) as a feature prior for conditional diffusion-based image restoration that consistently improves the quality of both degraded and low-resolution images.
Saif Ahmed, Ashadullah Galib, S. R. R. Antu et al.· 0 citations
This work introduces Vis-Poison, a novel visual knowledge poisoning attack where the poisoned image itself is the attacker-controlled payload, without manipulating captions, summaries, metadata, or other associated text.
Ru-Jin Liang, Zhongpu Chen, Yuhao Lei et al.· 0 citations
PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning, is introduced, and the observed accuracy spread across geometry families is larger than across presentation formats.
RecoverFly is proposed, a failure-aware RL post-training framework for end-to-end UAV-VLA policies that adapts token-level RL for stable optimization of grammar-constrained autoregressive UAV actions, revisits unresolved failure cases to strengthen corrective learning and sample utilization, and combines a two-stage long-tail scene curriculum with reference-policy regularization to improve scene adaptation while preserving acquired capabilities.
Boxiong Wang, Hui Kang, Geng Sun et al.· 0 citations
The proposed Barycentric Rational Forecasting with Chebyshev Enhancement (BRACE) maintains a local sliding window to cache sparse historical features and leverages adapted Chebyshev weights to formulate a barycentric rational function, directly aggregating these raw features to ensure numerical stability.
Jinlong Yang, Jinke Wu, Lizilin et al.· 0 citations
GHR-VLM, a visual grounded hybrid reasoning framework for zero-shot transit-bus video analytics, is proposed, motivated by the observation that explicit visual grounding can improve VLM reasoning by converting long surveillance streams into compact, passenger-centered spatiotemporal evidence.
Kaicong Huang, Weiheng Oh, Jack M. Reilly et al.· 0 citations
The visionary PhysioNet platform launched 25 years ago, based on a system developed at MIT in the 1970s. It has become one of the most comprehensive biomedical and clinical data repositories in existence.
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.
MIT News · Artificial Intelligence· news.mit.eduJun 30, 2026
Computer scientist Phillip Isola cuts through the hype to explain how AI agents work and what the future might hold for this rapidly advancing technology.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.