Skip to content

Category

computer vision

817 papers

SDGBiasBench: Benchmarking and Mitigating Vision-Language Models' Biases in Sustainable Development Goals

This work proposes CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors that yields significant gains on the proposed benchmark, which can foster the development of more fair and reliable AI systems for sustainable development.

Zihang Lin, Huaiyuan Qin, Mu Yang et al. · 0 citations

Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence

This paper proposes a camera-agnostic, one-shot, post-training pruning method for 3D Gaussian splats that relies solely on attribute-derived neighbourhood descriptors, and introduces a hybrid descriptor framework that captures structural and appearance consistency directly from the splat representation.

Peter O. Fasogbon, Ugurcan Budak, P. R. Alface et al. · 0 citations

Aligning Agentic World Models via Knowledgeable Experience Learning

WorldMind is introduced, a framework that autonomously constructs a symbolic World Knowledge Repository by synthesizing environmental feedback that unifies Process Experience to enforce physical feasibility via prediction errors and Goal Experience to guide task optimality through successful trajectories.

Baochang Ren, Yunzhi Yao, Rui Sun et al. · 3 citations · ⚡1
#artificial intelligence Conference Oct 2025

Riverbank Erosion Analysis in Bangladesh Using Spatiotemporal Segmentation

Riverbank erosion is a serious environmental problem in Bangladesh, causing land loss, damage to infrastructure, and displacement of local communities. Manual analysis of satellite images is often slow and difficult to apply consistently across large river networks. This study uses a parameter-efficient adaptation of the Segment Anything Model (SAM) to detect and measure riverbank erosion from historical Google Earth images. A dataset of 500 image pairs from 2003 to 2025 was prepared from erosion-prone areas, including Mokterer Char, Kedarpur, and Chowhali Upazila, with pixel-level labels for river, stable land, and eroded regions. During training, the ViT-H image encoder and prompt encoder were kept frozen, while only the lightweight mask decoder was fine-tuned for riverine segmentation. The adapted model achieved an erosion-class IoU of 0.867 and an F1-score of 0.928 on the primary held-out test set. Evaluation on unseen riverbank regions also showed that the model could generalize to new geographic areas, although detecting accreted land from RGB-only images remained difficult. The estimated erosion area differed from the ground truth by only 0.17%, showing that the model can produce reliable area measurements. Overall, this study demonstrates that adapted foundation segmentation models can support faster and more consistent riverbank erosion monitoring, with future scope for using multi-modal remote sensing data in broader environmental assessment.

M. Rafat, Akif Islam, Mohd Ruhul Ameen et al. · 0 citations
#artificial intelligence Preprint Sep 2025

OceanGym: A Benchmark Environment for Underwater Embodied Agents

OceanGym is introduced, the first comprehensive benchmark for ocean underwater embodied agents, designed to advance AI in one of the most demanding real-world environments, and reveals substantial gaps between state-of-the-art MLLM-driven agents and human experts.

Yida Xue, Mingjun Mao, Xiangyuan Ru et al. · 0 citations
#artificial intelligence Preprint Open access Aug 2026

Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection

Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. State-of-the-art VLMs still struggle to handle complex queries involving descriptive attributes and relational clauses. To address this problem, we propose restructuring linguistic representations according to the hierarchical relations within sentences for language-based object detection. A key insight is that textual tokens should be disentangled into core components-objects, attributes, and relations-and aggregated into hierarchically structured sentence-level representations. Building on this principle, we introduce the TaSe (Talk in Pieces, See in Whole) framework with three main contributions: (1) a hierarchical synthetic captioning dataset spanning three tiers from category names to descriptive sentences; (2) the three-component disentanglement module guided by a novel disentanglement loss function, transforms text embeddings into subspace compositions; and (3) aggregating disentangled components into hierarchically structured embeddings guided by the proposed hierarchical objectives. Experimental results under the OmniLabel benchmark show a 24% performance improvement, demonstrating the importance of linguistic compositionality.

Sojung An, Kwanyong Park, Yong Jae Lee et al. · 0 citations

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4

Long Story Short: Story-level Video Understanding from 20K Short Films

This work proposes Short-Films 20K (SF20K), the largest publicly available movie dataset, and accompanies this dataset with SF20K-Test, a manual, open-ended question answering benchmark, showing that instruction tuning on the large-scale dataset substantially improves model performance.

Ridouane Ghermi, Xi Wang, Vicky Kalogeiton et al. · 11 citations · ⚡1
#artificial intelligence Review Apr 2023

Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey

This survey reviews representative Transformer-based autonomous driving models and organizes them by task role, sensing configuration, and architectural design and analyzes how efficiency constraints reshaping model design choices in practice affects deployability, robustness, and safety.

J. Zhong, Zheng Liu, Xiangshan Chen · 21 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.