Skip to content

Category

natural language processing

3,231 papers

#artificial intelligence Preprint Aug 2026

Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

Slicing a byte-level BPE tokenizer allows one language model to operate at several vocabulary sizes, use a control token to indicate the active size, and be deployed at any trained size by slicing its embedding and output head, yielding a falsifiable prediction for future work.

Christos Koutsiaris · 0 citations
#artificial intelligence Preprint Aug 2026

Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning

Twin Worlds (TW), a framework for improving reliability in knowledge-intensive reasoning through equivariance-based abstention, is proposed, which identifies when answers are not reliably grounded in the provided evidence and outperforms uncertainty- and sufficiency-based baselines.

Vy Nguyen, Ziqi Xu, Jeffrey Chan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages

This work proposes LandingAgent, a three-phase agentic framework that profiles the target, constructs a reference-guided wireframe, and refines the page through critique-guided polishing, and evaluates it against direct prompting on faithfulness, conciseness, readability, aesthetics, and structural diversity.

In-Chang Baek, Hyeongseok Lee, Yearim Kim et al. · 0 citations
#artificial intelligence Preprint Aug 2026

OpenStamp: A Watermark for Open-Source Language Models

This work introduces OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer, and shows that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior methods.

Miroojin Bakshi, Saksham Rastogi, Danish Pruthi · 0 citations
#artificial intelligence Preprint Aug 2026

Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

It is shown that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution, and that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution.

Adarsh Sudheer, David Li, Omar Elbanna et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Trajectory-Level Speculative Decoding for Diffusion Language Models

This work develops a trajectory-level speculative framework that constructs draft denoising trajectories via confidence-stratified tree exploration and verifies them through blockwise parallel evaluation with bidirectional attention masking, and introduces inter-block speculation, exploiting diffusion models'bidirectional structure to perform cross-block lookahead.

Tian-Xiang Pan, Baitao Gong, Mo Guang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap

It is demonstrated that source-precision auditing alone does not rule out quantization-triggered behavior and that the final deployed configuration must be included in behavioral certification for trustworthy edge AI.

Jacopo Dardini, Claudio Stanzione, G. Colò et al. · 0 citations
#artificial intelligence Review Aug 2026

A Survey on Rubric-Guided Reinforcement Learning for Language Models

A Bayesian framework that defines constitutions as prior distributions over evaluation criteria and rubrics as conditional instantiations is introduced, and a taxonomy of rubric-guided RL along the prior-posterior axis is presented, covering constitutional AI, instance-specific rubrics, process-level supervision, self-evolving rubrics, and their agentic and multimodal extensions.

Zifei Shan, Fang-Ning Shao · 0 citations
#artificial intelligence Preprint Aug 2026

XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries inside the reasoning chain. We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence. Each instance is modeled as an evidence-dependency graph whose question, bridge evidence, answer-bearing evidence, and distractors have explicit language assignments. The audited resource contains 15,661 training and 7,405 validation instances, with sentence-level support supervision and supplied distractors. In validation, 99.81% of items cross the question-to-gold-evidence language interface and 95.60% use gold paragraphs in different languages. Across three reader artifacts, full question-evidence mismatch is associated with 10.25 to 15.79 lower Unicode-aware answer F1 than partial alignment, and different-script evidence with deficits of 11.98 to 23.70 points; the corresponding adapted-selector contrasts are 1.71 and 1.78 points. Under this supplied-candidate design, the evaluated readers therefore show substantially larger condition-associated deficits than the selector. XHotpotQA provides role-aware diagnostics, modular evaluation, and an audited test bed for knowledge-based systems that must integrate evidence across languages.

Iman Barati, A. Ghafouri, B. Minaei-Bidgoli · 0 citations
#artificial intelligence Preprint Aug 2026

Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection

A systematic comparison of retrieval strategies for candidate generation under a shared LLM-based selection stage, combining sparse retrieval (BM25), Web KB search, and a state-of-the-art trained dense retriever with several open- and closed-source LLMs is presented.

Fina Polat, Daniel Daza, Pengyu Zhang et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.