Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.
Guangxiang Zhao, Qi-Long Shi, Xusen Xiao et al.· 0 citations
This work demonstrates that a VLM-based pipeline can successfully unlock historical auction catalogs for large-scale automated analysis, and benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models.
Mathias Zinnen, Alisha Mund, Sabine Lang et al.· 0 citations
A red-teaming framework for evaluating this threat model targeting the autonomous skill generation and evolution pipeline of self-evolving agents and shows that SARGE induces malicious skill formation and that injected skills are persistently stored and repeatedly activated, highlighting the risk of persistent capability corruption.
Doyun Kim, Chanwoo Kim, Sugyeong Eo et al.· 0 citations
ICoA (Induced Covert Attack), an IPI attack designed to induce covert outcomes by steering the agent back to the user task by steering the agent back to the user task after executing the injection is proposed.
Yunseok Lee, Yunji Kim, Woo-Jin Lee· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A knowledge-gated task-construction protocol is introduced that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators, and it is shown that the retained tasks improve post-training.
Han-Lin Tian, Min-Hao Li, Yuhan Mi et al.· 0 citations
LaMoC improves joint compression by selecting compression statistics that better align local module reconstruction error with the downstream loss, and reformulate joint modular compression as a two-tiered optimization problem that minimizes module reconstruction error while tuning the activation and gradient information blending rate.
Experimental results show that the proposed model can not only recognize characters in document images but also locate word boundaries, removing the need for an extra word segmentation step in a conventional sequential pipeline.
Marry Kong, Rina Buoy, Sovisal Chenda et al.· 0 citations
To support long contexts efficiently, Sparse Gated Attention (SGA), which combines sparse attention with gated attention, and adopt Gated Norm (GN) to stabilize large-scale training is introduced, which keeps 4-bit NVFP4 serving within one point of FP8 accuracy.
Cheolseung Baek, Dhammiko Arya, Eunki Kim et al.· 0 citations
A controlled probe of 234 runs against the logged-out ChatGPT web interface and the OpenAI API, collected on 29 and 30 August 2026 across four exit countries and six query languages, shows that language and location are separable and act on different things.
This work built and validated PersonaGen-1M, a corpus of 1,031,732 synthetic buyer personas spanning 511 industry labels and 4 market contexts, carrying 19,416,821 structured behavioral attributes, 5,160,046 of them search queries.
SpanCalib-VLM is presented, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with the fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT).
Evaluating four state-of-the-art models finds that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI $+0.047$).
Shaghayegh Kolli, S. Emami, Moreno D'Incà et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.