This paper introduces diagnostic word-ablation metrics to quantify this phenomenon and proposes a data-centric solution that can alleviate the observed overcorrection in stance-aware argument retrieval and demonstrates that, for sufficiently powerful models, this approach can alleviate the observed overcorrection.
Angelo Sparacino, Francesca Toni, Adam Dejl· 0 citations
The results reveal that text restoration of missing text caused by physical lacunae in damaged ancient manuscripts using language models cannot be fully automated, but it can serve as a useful tool to assist paleographers in their work.
Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano et al.· 0 citations
Slicing a byte-level BPE tokenizer allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head, yielding a falsifiable prediction for future work.
Twin Worlds (TW), a framework for improving reliability in knowledge-intensive reasoning through equivariance-based abstention, is proposed, which identifies when answers are not reliably grounded in the provided evidence and outperforms uncertainty- and sufficiency-based baselines.
Vy Nguyen, Ziqi Xu, Jeffrey Chan et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work proposes LandingAgent, a three-phase agentic framework that profiles the target, constructs a reference-guided wireframe, and refines the page through critique-guided polishing, and evaluates it against direct prompting on faithfulness, conciseness, readability, aesthetics, and structural diversity.
In-Chang Baek, Hyeongseok Lee, Yearim Kim et al.· 0 citations
This work introduces OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer, and shows that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior methods.
It is shown that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution, and that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution.
Adarsh Sudheer, David Li, Omar Elbanna et al.· 0 citations
(k)-SwordStamp is designed: semantic watermarks with order-robust detection over sub-sentence units, reducing sensitivity to attacker-chosen structure at a small quality cost.
Abdulrahman Diaa, Jonathan Petit, Florian Kerschbaum· 0 citations
The results suggest that language models internally encode whether retrieved evidence is sufficient to support answering, and that this signal can be decoded reliably for RAG triage.
Syed Mahbubul Huq, Chris Child, Tillman Weyde et al.· 0 citations
This work develops a trajectory-level speculative framework that constructs draft denoising trajectories via confidence-stratified tree exploration and verifies them through blockwise parallel evaluation with bidirectional attention masking, and introduces inter-block speculation, exploiting diffusion models'bidirectional structure to perform cross-block lookahead.
Tian-Xiang Pan, Baitao Gong, Mo Guang et al.· 0 citations
It is demonstrated that source-precision auditing alone does not rule out quantization-triggered behavior and that the final deployed configuration must be included in behavioral certification for trustworthy edge AI.
Jacopo Dardini, Claudio Stanzione, G. Colò et al.· 0 citations
A Bayesian framework that defines constitutions as prior distributions over evaluation criteria and rubrics as conditional instantiations is introduced, and a taxonomy of rubric-guided RL along the prior-posterior axis is presented, covering constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and their agentic and multimodal extensions.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.