Empirically, PRISM reduces the end-to-end time for data selection and model tuning to just 30% of conventional pipelines, and achieves this efficiency while simultaneously enhancing performance, surpassing models fine-tuned on the full dataset across eight multimodal and three language understanding benchmarks.
Jinhe Bi, Yifan Wang, Danqi Yan et al.· arXiv.org· 73 citations· ⚡4
This work proposes Short-Films 20K (SF20K), the largest publicly available movie dataset, and accompanies this dataset with SF20K-Test, a manual, open-ended question answering benchmark, showing that instruction tuning on the large-scale dataset substantially improves model performance.
Ridouane Ghermi, Xi Wang, Vicky Kalogeiton et al.· International Journal of Com...· 11 citations· ⚡1
GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations of such models.
Xiaotian Zhang, Chun-yan Li, Yi Zong et al.· arXiv.org· 216 citations· ⚡17
This work borrows the notion of set-shifting from cognitive psychology to study how well LLM agents adapt to hidden reliability shifts, and introduces a suite of measures to quantify agent behavior after reliability shifts.
Zi-Hao Ye· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The Natural Language to AlphaGeometry Benchmark (NL2AGBench), which evaluates LLMs in translating English geometry problems into AlphaGeometry-compatible formal representations and introduces an error taxonomy distinguishing syntax and logic errors, which yield measurable improvements across multiple model families.
Samuel Xiao, Judy Song, Rory Hu et al.· 0 citations
One model passed the authors' fidelity check without ever opening the datasheet: a structured-output constraint had silently disabled tool use, and the model answered anyway, with fabricated source text, only the per-tool trace exposed it.
This work instantiates 17 paradigm-level configurations across five recurring modules of the ICL text-to-SQL pipeline under a single controlled implementation, and reveals that execution-feedback refinement is the only paradigm whose benefit holds universally at consistently low cost, while most other modules help only under backbone-dependent conditions.
Jia-Yan Lin, Yu-Jia Liu, Zi-Jin Hong et al.· 0 citations
Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls, and support a lossy-channel view of linguistic confidence.
Hefan Zhang, Bing-Quan Zhang, Ming Cheng et al.· 0 citations
BanglaMed-QA, a robust QA system specifically designed for the Bangla medical domain, is introduced and supervised machine learning models in which SVM is found to be the best model to categorize questions are adopted.
Rowzatul Zannat, Abdullah Al Shafi, K. M. Azharul Hasan et al.· 0 citations
How a stack behaves, and the quantities a defender cares about diverge: coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence.
Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani et al.· 0 citations
Verifier-Informed Student-to-Teacher Adaptation (VISTA), which preserves the standard OPSD student update while using outcome-verified rollouts to adapt the teacher toward the student distribution and demonstrates the value of student supervision from outcome-verified rollouts.
Ze-Wen Ding, Ze-Zhong Wu, Zhou Tao et al.· 0 citations
This paper formalizes the problem of KV eviction and proves that it is computationally hard, and shows that this probabilistic version of KV eviction coupled with decode time correction is more robust to different tasks compared to existing eviction methods and achieves competitive performance at the same compression budget.
Renato Lui Geh, Alexander K. Chen, Daniel Mingyi Israel et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.