Skip to content

Category

natural language processing

3,089 papers

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.

Jinhe Bi, Yifan Wang, Danqi Yan et al. · 73 citations · ⚡4

Long Story Short: Story-level Video Understanding from 20K Short Films

This work proposes Short-Films 20K (SF20K), the largest publicly available movie dataset, and accompanies this dataset with SF20K-Test, a manual, open-ended question answering benchmark, showing that instruction tuning on the large-scale dataset substantially improves model performance.

Ridouane Ghermi, Xi Wang, Vicky Kalogeiton et al. · 11 citations · ⚡1
#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations of such models.

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#artificial intelligence Preprint Jul 2026

Set-shifting Behavioral Test for Harnessed Agents

This work borrows the notion of set-shifting from cognitive psychology to study how well LLM agents adapt to hidden reliability shifts, and introduces a suite of measures to quantify agent behavior after reliability shifts.

Zi-Hao Ye · 0 citations
#artificial intelligence Preprint Aug 2026

NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry

The Natural Language to AlphaGeometry Benchmark (NL2AGBench), which evaluates LLMs in translating English geometry problems into AlphaGeometry-compatible formal representations and introduces an error taxonomy distinguishing syntax and logic errors, which yield measurable improvements across multiple model families.

Samuel Xiao, Judy Song, Rory Hu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL

This work instantiates 17 paradigm-level configurations across five recurring modules of the ICL text-to-SQL pipeline under a single controlled implementation, and reveals that execution-feedback refinement is the only paradigm whose benefit holds universally at consistently low cost, while most other modules help only under backbone-dependent conditions.

Jia-Yan Lin, Yu-Jia Liu, Zi-Jin Hong et al. · 0 citations
#artificial intelligence Preprint Aug 2026

When Linguistic and Internal Confidence Diverge in Large Language Models

Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls, and support a lossy-channel view of linguistic confidence.

Hefan Zhang, Bing-Quan Zhang, Ming Cheng et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

How a stack behaves, and the quantities a defender cares about diverge: coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence.

Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani et al. · 0 citations
#artificial intelligence Preprint Aug 2026

VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

Verifier-Informed Student-to-Teacher Adaptation (VISTA), which preserves the standard OPSD student update while using outcome-verified rollouts to adapt the teacher toward the student distribution and demonstrates the value of student supervision from outcome-verified rollouts.

Ze-Wen Ding, Ze-Zhong Wu, Zhou Tao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

A Probabilistic Interpretation of KV Cache Eviction

This paper formalizes the problem of KV eviction and proves that it is computationally hard, and shows that this probabilistic version of KV eviction coupled with decode time correction is more robust to different tasks compared to existing eviction methods and achieves competitive performance at the same compression budget.

Renato Lui Geh, Alexander K. Chen, Daniel Mingyi Israel et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.