These findings suggest that, within the pointwise scoring paradigm, routing continuous relevance semantics through discrete text constrains ranking signal resolution reveals a bottleneck that is stable and difficult to overcome under current standard methods, rather than an easily resolvable training bias.
Xiaoyang Chen, Jie Liu, Haijin Liang et al.· 0 citations
TAIScore (Targeted Actionable Improvement Score), a reward that evaluates the instruction, initial response, critique, and revision together, assessing whether the critique targets a real weakness, whether the actor follows it, and whether the intended aspect improves, is introduced.
Jinyoung Kim, Muhammad Khalifa, L. Logeswaran et al.· 0 citations
This survey systematizes recent progress in tree-search-based reasoning, viewing inference as instance-specific optimization rather than decoding, and introduces a Unified Design Space spanning search topology, evaluation signals, and control dynamics to unify a fragmented literature.
Jiaqi Wei, Xiang Zhang, Yue-Jin Yang et al.· 0 citations
This work outlines a framework for plausibility-aware AI that treats extracted claims not as final answers but as auditable evidence objects, making clear what was measured, how much it changed, in which setting, with what uncertainty, and from which source.
N. S. Babaiha, Stefan Geißler, Marie-Christine Simon et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work proposes AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories and suggests Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.
Zhuoran Lu, Yang-Yang Yu, Zhuoyan Li et al.· 0 citations
It is demonstrated that multi-agent consensus can enforce artificial agreement at the expense of true human alignment at the expense of true human alignment, revealing a structural limitation in consensus-style, role-specialized MAD protocols for subjective scoring.
Mi-Ra Song, Chanwoo Kim, Sugyeong Eo et al.· 0 citations
A probability-based framework for auditing MCQA benchmarks using model output distributions and introduces noise injection to reduce meaningful distractor competition, suggesting that this framework provides a lightweight audit lens for comparing macro-level benchmark structure and prioritizing individual items for targeted human review.
Fine-tuning can improve legal question-answering accuracy without improving how models use law supplied in context. We study this distinction in bilingual Bangladeshi legal QA, where observed errors can arise from answer scoring, retrieval, or failure to use relevant law. We construct a hierarchy-preserving statutory corpus, 2,165 reviewed bilingual fine-tuning examples, and a 150-item supplied-law control. We evaluate six instruction-tuned models: Llama-3.2-1B, Llama-3.2-3B, Qwen3.5-0.8B, Qwen3.5-2B, Qwen3.5-4B, and Gemma-4-E2B, with three LoRA seeds per model. To separate effects, we combine constrained option-letter scoring, cyclic option rotation, and controlled removal of the governing provision. On 398 Bar Council outputs, an exact-line parser attributes an accuracy gain of 50.0\% to the Qwen3.5-2B seed-42 adapter, whereas option scoring yields only $3.0\%$. For Gemma-4-E2B, the two scoring methods favor different systems. When the governing provision is guaranteed to be present, five of six reference models improve by $14.7\%-19.3\%$ under the four-order criterion. Removing that provision reduces accuracy by $8.0\%-15.3\%$ for models and by $13.8\%-14.9\%$ points for their adapters. However, difference-in differences estimates show no increase in reliance on the governing provision after fine-tuning. Results show that legal adaptation claims require separating scorer, retriever, and model effects. Our Code and data are available at https://anonymous.4open.science/r/bangladesh-legal-qa-11E3
HybridEmo is introduced, a post-training framework that initializes both multi-emotion TTS tasks with SFT and then aligns the speech-token policy through Group Relative Policy Optimization using a sample-aware hybrid reward.
We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.
Zihan Qiu, Zekun Wang, Xiao Li et al.· 2 citations· ⚡1
This work exposes lazy grounding by injecting nearby evidence from answer-changing rewrites of benchmark questions into the search corpora, and shows that robust search agents must defend against not only misinformation but also the misapplication of factual evidence.
Yu-Lin Zhang, Yukun Huang, Sanxing Chen et al.· 0 citations
Attribute-Agnostic Imbalance Augmentation (AIA) is proposed, a framework for improving model robustness under varying subgroup imbalances without explicit subgroup annotations and shows improved performance on the lowest-performing subgroups and consistent gains over competitive baselines.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.