A lightweight training framework that learns a single Behavior-Equivalent Token that substantially reduces inference cost and frees nearly the entire context window for user inputs and model outputs.
Jiancheng Dong, Pengyue Jia, Jingyu Peng et al.· 2 citations
This work formulates student reasoning as a multi-step sequential decision problem and introduces Monte Carlo Tree Search (MCTS) to explore optimal correction paths and proposes a dual evaluation protocol centered on solution accuracy and correct-step retention, offering a comprehensive measure of educational applicability.
Biaojie Zeng, Min Zhang, Juan Zhou et al.· arXiv.org· 0 citations
It is found that standard benchmarks do not adequately capture the strengths of the dataset, but expert judgment shows that SQPsych makes LLMs significantly better at therapist roleplaying.
Doan Nam Long Vu, Rui Tan, L. Mary Moench et al.· arXiv.org· 5 citations· ⚡2
This work introduces pruning laws, simple and interpretable scaling relations that connect a pruned LLM's post-pruning performance to its unpruned performance and pruning ratio, and demonstrates that the functional form transfers across dense and mixture-of-experts architectures, pruning methods, and unseen models in zero-shot and one-shot setups.
BEACON is an LLM-driven framework for cross-source CTI knowledge graph construction that constructs and releases two human-annotated datasets from 34 sources and outperforms all baselines by at least 23% and 9%, respectively.
CamoDocs is proposed, a poisoning attack that avoids direct query inclusion by camouflaging adversarial documents among benign content, and shows that erasure-heavy clustering defenses such as TrustRAG can reduce ASR, but only with substantial utility drops on retrieval-dependent benchmarks such as NeoQA.
Jaewon Jung, Haizhong Zheng, Hongsun Jang et al.· 0 citations
This work studies ViT attention heads and finds they differentiate into object- and background-specialist roles, a pattern most pronounced under full attention, and proposes SHS-Index to quantify this specialization, showing that it distinguishes full-attention from chunk-window ViTs, and finds that it strongly tracks downstream benchmark performance.
Chenyu He, Lei Li, Shi-Cheng Li et al.· 0 citations
AIM is proposed, a two-stage method that anchors an identity-forgetting target with a universal visual prompt and then matches the vision encoder to that target under a Fisher-based constraint, which achieves competitive identity forgetting while preserving non-deleted identities, prior knowledge, and visual perception on the same images.
Wonjun Lee, Jaehyuk Jang, Kangwook Ko et al.· 1 citation
Evaluation results demonstrate that the synthetic dataset constructed is the most effective approach for improving LVLM performance on reading vertically written Japanese text.
SILICA is an open instrument that tests three doubts of large language model agents: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read.
This work forms sector-targeted CTI dissemination as a multilabel classification problem, leveraging deep field knowledge of CTI structures and sector-specific threat patterns, and applies BERT, a transformer-based model, to automate the mapping of CTI events to sectors.
Fajar Wijitrisnanto, A. Abuadbba, Yan-Song Gao et al.· 0 citations
This work presents the first fine-grained cross-lingual analysis of prosody using multilingual dubbing data across English-German, English-Spanish, and English-French language pairs and reveals inherent cross-lingual correlations in prosodic structure between certain languages.
Haopeng Xie, Ismail Rasim Ulgen, Sofia Son et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.