This work proposes CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors that yields significant gains on the proposed benchmark, which can foster the development of more fair and reliable AI systems for sustainable development.
Zihang Lin, Huaiyuan Qin, Mu Yang et al.· arXiv.org· 0 citations
A novel Dual Self-Consistency Reinforcement Learning optimization paradigm is introduced, which utilizes Round-Trip Verification to penalize degenerate code and boost overall self-consistency in TikZ code.
Juekai Lin, Yun Zhu, Honglin Lin et al.· arXiv.org· 5 citations
This paper proposes a camera-agnostic, one-shot, post-training pruning method for 3D Gaussian splats that relies solely on attribute-derived neighbourhood descriptors, and introduces a hybrid descriptor framework that captures structural and appearance consistency directly from the splat representation.
Peter O. Fasogbon, Ugurcan Budak, P. R. Alface et al.· arXiv.org· 0 citations
This work proposes a cognitive-inspired, multi-agent framework that operationalizes Conceptual Blending Theory (CBT) through a novel Schema Grammar ("G"), providing a rigorous foundation for cross-domain logic re-instantiation.
Yu Xu, Yuxin Zhang, Juan Cao et al.· arXiv.org· 4 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
WorldMind is introduced, a framework that autonomously constructs a symbolic World Knowledge Repository by synthesizing environmental feedback that unifies Process Experience to enforce physical feasibility via prediction errors and Goal Experience to guide task optimality through successful trajectories.
Baochang Ren, Yunzhi Yao, Rui Sun et al.· arXiv.org· 3 citations· ⚡1
Riverbank erosion is a serious environmental problem in Bangladesh, causing land loss, damage to infrastructure, and displacement of local communities. Manual analysis of satellite images is often slow and difficult to apply consistently across large river networks. This study uses a parameter-efficient adaptation of the Segment Anything Model (SAM) to detect and measure riverbank erosion from historical Google Earth images. A dataset of 500 image pairs from 2003 to 2025 was prepared from erosion-prone areas, including Mokterer Char, Kedarpur, and Chowhali Upazila, with pixel-level labels for river, stable land, and eroded regions. During training, the ViT-H image encoder and prompt encoder were kept frozen, while only the lightweight mask decoder was fine-tuned for riverine segmentation. The adapted model achieved an erosion-class IoU of 0.867 and an F1-score of 0.928 on the primary held-out test set. Evaluation on unseen riverbank regions also showed that the model could generalize to new geographic areas, although detecting accreted land from RGB-only images remained difficult. The estimated erosion area differed from the ground truth by only 0.17%, showing that the model can produce reliable area measurements. Overall, this study demonstrates that adapted foundation segmentation models can support faster and more consistent riverbank erosion monitoring, with future scope for using multi-modal remote sensing data in broader environmental assessment.
M. Rafat, Akif Islam, Mohd Ruhul Ameen et al.· 2026 International Conferenc...· 0 citations
OceanGym is introduced, the first comprehensive benchmark for ocean underwater embodied agents, designed to advance AI in one of the most demanding real-world environments, and reveals substantial gaps between state-of-the-art MLLM-driven agents and human experts.
Yida Xue, Mingjun Mao, Xiangyuan Ru et al.· 0 citations
Vision-language models (VLMs) have advanced multimodal perception, demonstrated by open-vocabulary object detection with simple language queries. State-of-the-art VLMs still struggle to handle complex queries involving descriptive attributes and relational clauses. To address this problem, we propose restructuring linguistic representations according to the hierarchical relations within sentences for language-based object detection. A key insight is that textual tokens should be disentangled into core components-objects, attributes, and relations-and aggregated into hierarchically structured sentence-level representations. Building on this principle, we introduce the TaSe (Talk in Pieces, See in Whole) framework with three main contributions: (1) a hierarchical synthetic captioning dataset spanning three tiers from category names to descriptive sentences; (2) the three-component disentanglement module guided by a novel disentanglement loss function, transforms text embeddings into subspace compositions; and (3) aggregating disentangled components into hierarchically structured embeddings guided by the proposed hierarchical objectives. Experimental results under the OmniLabel benchmark show a 24% performance improvement, demonstrating the importance of linguistic compositionality.
Sojung An, Kwanyong Park, Yong Jae Lee et al.· 0 citations
Evaluating nine closed-source model routes from Anthropic, Google, and OpenAI on TallyBench and CompareBench reveals strong overall performance but persistent failures in counting, spatial reasoning, geometric comparison, and temporal ordering.
Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.
Jinhe Bi, Yifan Wang, Danqi Yan et al.· arXiv.org· 73 citations· ⚡4
This work proposes Short-Films 20K (SF20K), the largest publicly available movie dataset, and accompanies this dataset with SF20K-Test, a manual, open-ended question answering benchmark, showing that instruction tuning on the large-scale dataset substantially improves model performance.
Ridouane Ghermi, Xi Wang, Vicky Kalogeiton et al.· International Journal of Com...· 11 citations· ⚡1
This survey reviews representative Transformer-based autonomous driving models and organizes them by task role, sensing configuration, and architectural design and analyzes how efficiency constraints reshaping model design choices in practice affects deployability, robustness, and safety.
The visionary PhysioNet platform launched 25 years ago, based on a system developed at MIT in the 1970s. It has become one of the most comprehensive biomedical and clinical data repositories in existence.
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.
MIT News · Artificial Intelligence· news.mit.eduJun 30, 2026
Computer scientist Phillip Isola cuts through the hype to explain how AI agents work and what the future might hold for this rapidly advancing technology.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.