This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: security through analyzing, detecting, and designing backdoor attacks in NLP and VLMs, and efficiency through advanced multimodal representation methods tailored for clinical and medical imaging applications.
NE-BERT, a domain-specific multilingual encoder model trained on approximately 8.3 million sentences spanning 9 Northeast Indian languages and 2 anchor languages, addresses critical vocabulary fragmentation issues in extremely low-resource languages such as Pnar and Kokborok through aggressive upsampling strategies.
Badal Nyalang· Proceedings of the Second Wo...· 2 citations
Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment using only a small set of contrastive prompts. We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted. To hinder this process, we introduce a weight-editing method that obscures the refusal signal by applying rank-$k$ updates to residual stream writer matrices while replacing refusal-inducing activations with random aliases and correcting downstream reader matrices to preserve the model's original behavior. On Llama-3-8B, AMRA improves post-abliteration refusal scores by $2.16$ points over the undefended baseline with less than $0.5$ percentage points of MMLU degradation. On Gemma-2-9B, it improves the post-abliteration refusal by $14.70$ points over the baseline while keeping harmful output rates similar to the baseline, albeit at a greater utility cost.
This work makes two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.
Songeun Chae, Min Kim, Donghoon Jung et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work shows how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained.
This work introduces Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time.
This work proposes SuTRA (Structurally-Unified Tokenization with Root Awareness), a morphology-aware algorithm that preserves akshara indivisibility and penalizes merges crossing morphological boundaries and releases a new morphological segmentation dataset for Hindi, Marathi, and Gujarati.
Vaibhav Rathore, Siddhant Gole, Dadhichi Telwadkar et al.· 0 citations
Analyzing a large corpus of publicly released post-training trajectories, it is found that across different tasks, the agent's training strategy is locked in at the very beginning, and the entire remaining budget is spent on local adjustments within the selected strategy.
J. Lim, Xinyuan Huang, Hao Peng et al.· 0 citations
An adaptive memory and reflection (AMR) agentic system, a multi-agent framework in which specialized agents use dedicated memory and reflection-based feedback to retrieve relevant prior cases and improve subsequent reasoning.
P. Murugesan, Luoxiao Yang, Xueli Chen et al.· 0 citations
EvalCEGAR is a pool of small Python operators that each flag a candidate for one named defect, or abstain, and vote, and borrows counterexample-guided abstraction refinement from program verification to score agents against a reliable automatic metric.
Xing Zhang, Ya Cui, Guanghui Wang et al.· 0 citations
BudgetDoc is introduced, the first multimodal benchmark providing explicit supervision for model-budget-performance trade-offs across three document tasks, and DRB (Document-Reasoning Balancer), an approx.
This work presents ComponentBench, a benchmark and diagnostic pipeline for component-level evaluation of computer-use agents on modern web UIs, and introduces a scalable pipeline for auditing realized structural difficulty after implementation and synthesizing structured failure analyses across tasks and component families.
Tianchen Guan, Xinlei Lin, Royce Cheng-Yue et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.