Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users'data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user interface (GUI) actions such as tapping, swiping, and typing, which often form long, interface-dependent sequences, cannot directly access device capabilities, and make execution boundaries difficult to define. We present PalmClaw, an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device. PalmClaw exposes device capabilities as device tools with explicit arguments, structured results, and clearly defined execution boundaries. This design enables agents to use mobile capabilities directly while keeping each action explicit and controlled. Experiments show an 11.5% relative improvement in task success and a 94.9% reduction in completion time over the strongest baseline, with lower setup burden and traces illustrating how execution boundaries are applied. Code is available at https://github.com/ModalityDance/PalmClaw.
Hongru Cai, Yongqi Li, Ran Wei et al.· 0 citations
The results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad, which cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.
Jiatong Li, Yuxuan Ren, Weida Wang et al.· arXiv.org· 1 citation
This work demonstrates a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on synthetic reasoning tasks and indicates that frontier models already perform consequential computation with no interpretable trace in their output tokens.
Vatsal Baherwani, Tom Goldstein, Ashwinee Panda· arXiv.org· 4 citations· ⚡2
A pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work is introduced.
V. Ravikumar, Sina Ahmadi, L. Jäger et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
HALO (Hallucination-Aware Layered Oversight) is presented, an assurance architecture which treats hallucination as a containable failure mode rather than an eliminable one and detail each layer, give particular attention to evidence-based confidence (which verifies extractions against the source document rather than trusting the model's self-reported certainty).
Bogdan Raduta, Horia Velicu, Alexandru Preda et al.· arXiv.org· 0 citations
A unified phoneme-based TTS-to-ASR augmentation pipeline built around a multilingual TTS model trained from scratch using the F5-TTS architecture with language-ID conditioning is presented and phoneme-frequency-guided selection (PFGS) is proposed, which ranks candidate sentences using phoneme frequencies estimated from real ASR training labels.
Zhen Wang, Tian-Rui Wu, Rong-Qi Han et al.· 0 citations
This formulation enables a systematic study of key self-improvement factors through the proposed Evo-Harness, and provides a principled understanding of how LLM agents can effectively learn on the fly.
Tianxin Wei, Zhan Shi, Minhua Lin et al.· 3 citations
SeMoCo, a semantic-first motion codec, is introduced together with a dual-axis motion generator for language-conditioned motion generation and $\Omega$-MotionVerse, a large-scale, multi-source human-motion dataset unified under the SOMA representation is constructed.
Tianlv Huang, Hetian Guo, Zi-Yi Cai et al.· 0 citations
This work presents a security-oriented framework for risk identification, evaluation, and mitigation in a multi-agent GIS system while maintaining adaptability to broader agentic architectures.
K. Gao, Pranavi Kotta, Linlin Xu et al.· IEEE Geoscience and Remote S...· 0 citations
MemoryCard is a video-memory-based augmentation framework that organizes long videos into self-contained Memory Cards, each corresponding to a distinct topic or event, and consistently improves long-video QA performance under comparable visual-token budgets.
Qing Yang, Pengcheng Huang, Xinze Li et al.· arXiv.org· 0 citations
A trust-aware post-routing framework is proposed that reweights clients using returned-evidence feedback, including retrieval relevance, profile consistency, and cross-client agreement, and online experiments show that it suppresses persistent hijacking over recurring queries and transfers to a learned neural router.
Mean Cumulative Drift (MCD), an embedding-based measure of content retention across three representation spaces, and Multi-Generation GenEval (MGG), extending GenEval's object-level compliance scoring across generations are proposed, to quantify drift.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.