Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a particular local transition. We formalize this distinction...
Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate...
Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degradation, characteristic structures in TFRs align predominantly along the frequency or time axis, making it challenging f...
Hanbyeol Park, Jungho Choo, Hyerim Bae· 0 citations
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile div...
Mao-Kai Qin, Chuan Qin, Qi Zhang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no...
Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit analysis across six model architectures spanning four familie...
Arjun Pillai, Christian Hoang, Anjelo Laroza· 0 citations
Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding...
Jia-Qi Zhu, Nai-Li Xing, He-Xiang Pan et al.· 0 citations
FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments, shows that individually capable agents do not automatically form safe financial systems, highlighting system-level evaluation and interaction design as central problems for financial AI safety.
A clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory) that improves session quality and achieves a longitudinal mean of 2.29 on cross-session continuity, intervention progression, and personalization.
Deng-Du Jiang, Shuo Zhang, Wei-Wei Liao et al.· 0 citations
Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games, offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.
Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni et al.· 0 citations
JevSoup is proposed, a training-free framework separating System One expert routing from System Two execution using only the input and expert descriptions, which retains the leading expert's update, projects the second onto the orthogonal complement of the first update's row space, and combines them with equal weights.
Xiu-Ying Wang, Jia-Hua Cheng, Shuo-Tian Li et al.· 0 citations
We introduce Acacia, a graph foundation model, trained on the web graph. Acacia (i) supports arbitrary feature dimensionalities and semantics without additional training, (ii) supports a wide range of tasks, including node classification, link prediction, node clustering, and graph generation, without additional traini...
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.