Skip to content

Category

computer vision

817 papers

#machine learning Preprint Aug 2026

EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders

This work proposes Explainable Probing of Cross-Domain Sparse Embeddings (EXPOSE), a framework that uses Sparse Autoencoders (SAEs) as an explainable bottleneck to identify and suppress domain-specific components in VFM embeddings.

Anja Witte, M. Lennartz, Jan Baumbach et al. · 0 citations
#machine learning Preprint Aug 2026

Empowering Local Agriculture: A Deep Learning-Powered Web System for Identifying Bangladeshi Mango Varieties

This work presents a deep learning-based web system for automatic identification of Bangladeshi mango varieties and integrated the model into a Streamlit web application that enables users to upload a mango image and receive a predicted variety with class probabilities.

Monowar Islam, Safaruzzaman Shovo · 0 citations
#machine learning Preprint Aug 2026

A Deeper Analysis of Block-Sparse Featurizers

This work proposes several architectural changes to the BSF, including a Tournament Top-K selection rule that significantly reduces feature splitting, and extends the block paradigm to the crosscoder.

Alexandru-Iulius Jerpelea, Amith Ananthram · 0 citations
#artificial intelligence Preprint Aug 2026

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

This work formalizes probabilistic alignment as a distributional criterion for world models and introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics, and introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors.

Yuandong Pu, Le Zhuo, Sayak Paul et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT is introduced, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation, and a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference.

Yushe Cao, Shikun Feng, Ru-Xiang Duan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

A hybrid GAN-guided diffusion framework that uses a pretrained Wasserstein GAN with gradient penalty (WGAN-GP) as a feature prior for conditional diffusion-based image restoration that consistently improves the quality of both degraded and low-resolution images.

Saif Ahmed, Ashadullah Galib, S. R. R. Antu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation

RecoverFly is proposed, a failure-aware RL post-training framework for end-to-end UAV-VLA policies that adapts token-level RL for stable optimization of grammar-constrained autoregressive UAV actions, revisits unresolved failure cases to strengthen corrective learning and sample utilization, and combines a two-stage long-tail scene curriculum with reference-policy regularization to improve scene adaptation while preserving acquired capabilities.

Boxiong Wang, Hui Kang, Geng Sun et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference

The proposed Barycentric Rational Forecasting with Chebyshev Enhancement (BRACE) maintains a local sliding window to cache sparse historical features and leverages adapted Chebyshev weights to formulate a barycentric rational function, directly aggregating these raw features to ensure numerical stability.

Jinlong Yang, Jinke Wu, Lizilin et al. · 0 citations
#artificial intelligence Preprint Jul 2026

GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning

GHR-VLM, a visual grounded hybrid reasoning framework for zero-shot transit-bus video analytics, is proposed, motivated by the observation that explicit visual grounding can improve VLM reasoning by converting long surveillance streams into compact, passenger-centered spatiotemporal evidence.

Kaicong Huang, Weiheng Oh, Jack M. Reilly et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.