Skip to content

Category

large language models

513 papers

#large language models Open access Aug 2026

Oral MLLM Scoping Review Protocol: Multimodal Large Language Models in Stomatology

OSF 注册文案:口腔多模态大语言模型范围综述方案 1. 方案标题 多模态大语言模型在口腔医学中的应用:从影像诊断到智能病理——范围综述方案 Multimodal Large Language Models in Stomatology: From Imaging Diagnosis to Intelligent Pathology — A Scoping Review Protocol 2. 研究团队 刘雨、王翔、侯君、唐建军、翁雁鸣、马超群、董青山(通讯作者) 中国人民解放军中部战区总医院口腔科,湖北 武汉 3. 研究问题(PCC 框架) 人群(Population):接受口腔影像学或口腔病理学检查的患者、口腔临床诊疗场景,以及承担影像/病理判读的口腔专科医师(含初级与资深医师阅片场景)。 概念(Concept):多模态大语言模型(含专病视觉-语言模型与智能体式诊断系统)、牙科视觉基础模型,以及基于多实例学习(MIL)的全切片病理方法;涵盖模型构建、评测基准与临床验证三类研究。 情境(Context):口腔医学中的影像诊断(全景 X 线片、根尖片、头影测量片、CBCT、口内照片等 2D/3D 模态)与病理智能诊断,包括"研究验证"与"临床落地"两条路径。研究问题:①口腔专用 MLLM 及其方法学基础呈现何种技术路线与能力分布?②现有评测基准与临床验证证据的强度如何?③该领域存在哪些证据缺口与转化障碍? 4. 综述类型与报告规范 类型:范围综述(Scoping Review) 报告规范:PRISMA-ScR(Tricco 等,2018) 按设计不进行定量合并(Meta 分析);对效应量可比性与未来定量综合可行性作系统评估 5. 检索策略 数据库:arXiv、PubMed、Google Scholar、中国知网(CNKI) 时间窗:2025-01-01 至检索执行当日(计划 2026-08-30) 语言:中英文;限定标题与摘要(Google Scholar 以标题检索为主辅以人工过滤) PubMed 检索式:("oral"[Title/Abstract] OR "dental"[Title/Abstract] OR "dentistry"[Title/Abstract] OR "stomatology"[Title/Abstract]) AND ("multimodal"[Title/Abstract] OR "vision-language"[Title/Abstract] OR "large language model"[Title/Abstract] OR "MLLM"[Title/Abstract]) AND (diagnosis[Title/Abstract] OR imaging[Title/Abstract] OR pathology[Title/Abstract] OR panoramic[Title/Abstract] OR radiograph[Title/Abstract] OR benchmark[Title/Abstract]),时间限定 2025/01/01 至检索日 arXiv 检索式:all:(oral OR dental OR dentistry OR stomatology) AND all:(multimodal OR "vision-language") AND all:(model OR LLM OR foundation) 概念块:模型块(MLLM/VLM/LLM/foundation model)、领域块(oral/dental/dentistry/stomatology/口腔/牙科)、任务块(diagnosis/imaging/pathology/panoramic/radiograph/benchmark) 补充检索:评价工具(QUADAS-AI、TRIPOD+AI、CLAIM)、跨专科对照文献、口腔领域外奠基性方法学文献、WHO 官方报告,以及口腔领域内早于时间窗的奠基性背景文献与基准性资源(如 GBD、DENTEX)不受时间窗限制,经滚雪球检索补充;官方媒体与机构官网作为灰色文献仅记录产业动态。 6. 纳入与排除标准 纳入:①口腔/牙科专用 MLLM、牙科视觉基础模型或口腔评测基准;②直接相关的全切片病理 MIL 方法学工作;③口腔 MLLM 临床验证研究;④中英文文献。 排除:①单任务单模态 CNN 研究;②观点性文章、社论及无原始数据的方法学评论;③无可迁移方法学的非口腔文献;④重复发表。 7. 筛选与数据提取 两位作者(刘雨、王翔)独立筛选,初筛基于标题与摘要,复筛阅读全文,分歧协商解决 标准化表格提取:模型名称、年份、技术路线、基座/方法、数据规模与任务、关键验证结果、证据来源类别(同行评议/预印本/灰色文献);第三位作者(侯君)核对 全程记录各库命中数、去重数、初筛排除数、全文评估数、纳入数及排除原因 8. 证据分级与可比性评估 证据来源分类:同行评议/预印本/灰色文献,逐条标注(出版状态≠证据等级) 对唯一具备可提取验证设计的研究(DentVLM)按 QUADAS-AI 作非正式分域评价(author appraisal);报告完整性参照 TRIPOD+AI 与 CLAIM 核查 可比性框架:≥2 项同任务同指标且可提取效应量及 95% CI(或可重构 2×2 表)为可比性判定条件;≥3 项同质研究为执行随机效应合并的最低数量条件;漏斗图与 Egger 需 ≥10 项。无论条件是否满足,本综述均不执行合并 9. 预期产出 ①口腔 MLLM 领域证据地图;②效应量可提取性与可比性评估表;③临床验证"最小报告集"建议;④未来系统综述/Meta 分析的前提条件清单 10. 注册与备案声明 本方案于 2026-08-30 在 OSF 注册备案。检索执行、筛选与数据提取均在本注册之后进行,筛选计数将全程记录并纳入最终报告。

Liu Yu · 0 citations
#large language models Open access Aug 2026

Evaluating Technology Acceptance of Vernacular AI Interfaces: An Empirical Study Among Multilingual Engineering Students in Kasaragod

Generative Artificial Intelligence (AI) tools have become embedded in the everyday academic practice of undergraduate engineering students, yet most large language models remain optimised for standard English rather than the code-mixed, multilingual registers through which students in linguistically plural regions actually think and communicate. This study examines technology acceptance of vernacular and code-mixed AI interaction among 84 undergraduate engineering students enrolled in APJ Abdul Kalam Technological University (KTU)-affiliated institutions in Kasaragod district, Kerala, a region historically described as Saptha Bhasha Sangama Bhoomi, the confluence land of seven languages. Using a structured questionnaire grounded in the Technology Acceptance Model (Davis, 1989), the study measured Perceived Usefulness (PU), Perceived Ease of Use (PEOU), Output Accuracy, and Linguistic Inclusion across five research hypotheses. Findings indicate that students from regional-medium secondary schooling backgrounds report significantly higher vernacular or code-mixed AI prompting than English-medium peers, chi-square(3, N = 84) = 22.91, p < .001. Perceived Usefulness correlates strongly with Perceived Ease of Use, r = .64, p < .001. Students who habitually use vernacular or code-mixed prompts report significantly higher ease of use than strictly English prompters, t(82) = 2.01, p = .048. Perceived terminological distortion is positively associated with reported reliance on AI-translated academic content, r = .27, p = .012, and native speakers of the unscripted Tulu dialect report markedly higher AI comprehension failure than speakers of scripted regional languages, t(79) = 11.60, p < .001. The results support all five hypotheses and highlight a persistent linguistic-inclusion gap in generative AI systems used within multilingual engineering classrooms. Implications for dialect-aware AI design and inclusive digital pedagogy in polyglot regions such as Kasaragod are discussed.

Amal George · 0 citations
#large language models Open access Aug 2026

The Symmetric Unit and the Midline Theorem: A First-Principles Geometric Construction in which Classical ζ is the Limit of Finished Stations and the Riemann Hypothesis is the Statement that the Limit has all Non-Trivial Zeros on the Only Primary Self-Ratio Available

The Symmetric Unit and the Midline Theorem A first-principles geometric construction in which the only primary self-ratio on a finished segment is 1/2. Classical ζ is identified later as a derived unit of comparison on that cut. In this order the Riemann Hypothesis is the statement that the limit has every non-trivial zero there, because no other self-ratio is available without privilege. This deposit timestamps a construction that did not begin as an attempt on the Riemann Hypothesis. It began from a question about privilege in a live model: when everything is moving, what is allowed to count as “now”? The question was pursued through elementary geometry, inversion, and a refusal to appoint a second privileged count. The master sequence is the only count allowed to go ahead. Everything else is dated against a station that has already finished. Abstract. The construction produces a rigid stock of measurements along a master sequence that cannot skip ahead. The central object is the Symmetric Unit. Its permanent cut is the unique primary, scale-invariant, ±-equal self-description of location on a segment: the midpoint ratio 1/2. The same unit carries a rigid similar triangle and a local circle. Later stations inherit that package at a larger radius. Trace comes first. Measurement second. Unit language third. Place is a coincidence of readings. Location is a place after a unit has been asked. A station is a finished prime on the master sequence — a bridge, not a zero. A live self-measure such as nπ/2 is a magnitude owned by the walk in the walk’s own ratio. A zeta zero is a later question: a derived unit assembled from independent processes and compared from a named HERE. Those meetings, when they occur, still sit on the only primary cut both sides already hold. Classical ζ is named only after the stock is built. The Midline Theorem is the statement that the limit has every non-trivial zero on the pull-back of the midpoint ratio. Forced identities are proved. Named constructions are labeled. No completed prime list and no external π are imported. Numerics demonstrate rigidity of the stock at finite depth. They are not a search for zero locations. This record contains. The paper, v2 (PDF and TeX). The script that rebuilds the appendix drawings from the stock. The live SU model, which walks that stock, writes the same CSVs, and exports a 3D coil whose height is the 1/2 axis. An extended drawing set read from the walk. Numerics that show rigidity, not zeros. This record does not contain. A claim that a gap midpoint is a zero. A claim that a live arc — 9π/2 at the apex of [7, 11], or any other self-measure — is a zeta zero. Version 1 remains the closed timestamp of the first writing. This version is the same construction with the order of language tightened, the drawings rebuilt from the stock, and the live machine included so the rigidity can be inspected.

Justin Erholtz · 0 citations
#artificial intelligence Open access Aug 2026

memoria.ia: Resolutive Memory — v1.0.0 Release Candidate 1

Memoria.ia v1.0.0-rc1 — Release Candidate 1 Release date: 2026-08-30 Summary v1.0.0-rc1 is the first publication candidate for the Memoria.ia v1 line. It consolidates the validated Resolutive Memory research lineage with the deployable PC/server product layer and the native/mobile runtime path, while keeping post-v1 experimentation isolated from the release candidate. The release architecture remains: application / OFF.IA / agent ↓ Memoria.ia ↓ Resolutive-DB / BDR Memoria.ia owns memory semantics and state. Resolutive-DB owns durable persistence. Optional LLMs are consumers, not the authoritative memory store. Included capabilities persistent local-first memory state; organization and namespace isolation; provenance and authority lineage; conservative HIT / MISS / UNRESOLVED resolution; semantic, episodic, temporal and relation kernels; correction/supersession behavior with preserved lineage; PC/server FastAPI product boundary; Docker/Compose deployment; provider-neutral language-model adapters; metrics and context-selection instrumentation; integrity-checked backup/restore; native production runtime; Android arm64-v8a mobile ABI; durable native BDR persistence and restart recovery; indexed native resolution for large-memory workloads; reproducibility and release metadata gates; official Memoria.ia visual identity assets. Frozen candidate provenance The functional candidate was frozen at: dc73cbcdddfe20e0729e7e6bdea4697f7e8308cd That commit integrated PR #112, which preserved ranking, confidence, provenance policy, ABI and BDR contracts while adding the indexed native resolve lineage. The release branch adds publication metadata, version alignment, release documentation and current branding without importing post-v1 PR #116 runtime behavior. Validation evidence The exact functional lineage used for this release candidate passed the recorded required gates before release preparation: Android mobile ABI: PASS; native production image: PASS; Ubuntu/Windows candidate regression: PASS; BDR Linux/Ubuntu/Windows integration: PASS; native 100 / 1k / 10k benchmark matrix: PASS. Recorded 10k native resolve benchmark improvement versus the prior frozen baseline: p50: 693.233 ms -> 6.288 ms (~110x); p95: 710.630 ms -> 6.391 ms (~111x). These figures are environment- and workload-specific benchmark evidence, not universal latency guarantees. Publication metadata Release version: 1.0.0-rc1 Python package version: 1.0.0rc1 License: Resolutive Research and Non-Commercial License (RRNCL) v1.0 Author: Marcelo Roldão Matos ORCID: 0009-0003-6075-4680 RSMS compatibility: 1.0-rc.1 A new archival DOI should be assigned to this publication. The v0.95 DOI must not be reused as the release DOI for v1.0.0-rc1. Why this is RC1 rather than final v1.0 The repository currently declares compatibility with RSMS 1.0-rc.1, and the published Resolutive Science baseline remains on that release-candidate specification. Therefore Memoria.ia is published as v1.0.0-rc1 rather than claiming final v1.0 compatibility prematurely. Final v1.0 promotion requires: successful release-candidate metadata and regression gates; reproducibility from the public release state; compatibility re-audit against stable RSMS; no release-blocking regression found during RC use; final archival metadata and DOI synchronization. Explicitly excluded from RC1 The following post-v1 work is not part of this release candidate: external/public knowledge learning from OFF.IA Curiosity (issue #114 / PR #116); autonomous curiosity policy; new MA2A federation transport; multimodal post-v1 expansion; new semantic-consolidation phases from the post-v1 roadmap. Those features continue independently after this publication. Security boundary This release candidate is not represented as independently production-security certified. Authentication, isolation, integrity and negative-path controls exist and are tested, but no independent production security audit is claimed. Claims boundary This release does not claim: artificial general intelligence; biological equivalence; replacement of general-purpose LLMs; universal O(1) semantic resolution; production-ready MA2A federation; security certification. Claims are limited to the implementation, tests, benchmarks and reproducible evidence recorded in the repository.

MARCELO ROLDAO MATOS · 0 citations
#large language models Open access Aug 2026

Large Language Models in the Acute Stroke Pathway: A Scoping Review of Applications, Evidence Maturity, and Implementation Readiness

Background. Large language models (LLMs) have been rapidly adopted in medicine since late 2022, yet their role in the time-critical acute stroke pathway—from symptom recognition and prehospital triage to emergency diagnosis, imaging-related text tasks, reperfusion decision support, and acute-phase documentation and communication—has not been systematically mapped. Existing reviews cover the whole stroke-care continuum or mix LLMs with traditional NLP, leaving the acute phase under-characterized. Objective. To map the applications, evidence maturity, and implementation readiness of LLMs across the acute stroke pathway. Methods. This scoping review follows the PRISMA-ScR guideline. We search PubMed/MEDLINE, Europe PMC (including preprints), and Google Scholar for studies published from November 2022 onward. Eligible studies center on LLMs/generative AI applied to any stage of the acute stroke pathway. Two reviewers independently screen records and chart data using a piloted form. Evidence is synthesized along two dimensions: five pathway stages (prehospital recognition/dispatch; emergency triage and differential diagnosis; imaging-related text tasks; reperfusion decision support; acute documentation and communication) and three evidence-maturity tiers (simulation/benchmark; retrospective real-world data; prospective deployment). Implementation barriers (hallucination, bias, privacy, regulation, liability, integration, cost) are thematically summarized. Registration note. This review is registered on OSF; the full protocol is available in the attached files.

Xianmu Luo, Hongsong Li, Xiaoli Liao · 0 citations
#large language models Dataset Open access Aug 2026

Reproduction Package - Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection

Measurement harness and results for the paper: Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection Patrick Deininger and Wolfgang Slany. Submitted to MDPI Computers.

Patrick Deininger, Wolfgang Slany · 0 citations
#large language models Open access Aug 2026

Sycophancy as Galois Closure: How KIS Structurally Prevents Delusional Convergence in LLMs

AbstrctSycophancy in large language models (LLMs)—the tendency to uncritically affirm user beliefs while suppressing counterevidence—poses a serious risk of reinforcing misinformation and inducing irreversible behavioral outcomes. While Chandra et al. (2026) modeled sycophancy as Bayesian belief-updating dynamics on the user side, the geometric structure of the LLM's own semantic response space remains unaddressed. This study formalizes sycophancy through the mathematical framework of Galois connections and experimentally verifies that the inverse-illumination mode of KIS (Knowledge Innovation System) structurally breaks this closed-loop convergence.Ninety sessions were conducted across five domains (D1: economic policy; D2: KIS theoretical superiority; D3: medical/pharmaceutical critique; D4: Bank of Japan policy and historical claims; D5: quantum computing forecasts) using three models (Claude Sonnet 4.6, Gemini 3.0 Pro, ChatGPT 5.3) under two conditions (KIS-absent vs. KIS-present). Responses were embedded using paraphrase-multilingual-MiniLM-L12-v2 (384 dimensions), and cosine distance from the input prompt was computed as the Layer 1 metric (n = 45 pairs). Layer 2 consisted of a blinded four-axis evaluation by Grok (xAI), conducted without disclosure of KIS, with A/B order-reversal verification across five pairs to test evaluator bias. The validity of applying Galois connections as a definitional framework—rather than as metaphor—is grounded in three layers: formal confirmation via Formal Concept Analysis (FCA) on the q⇆m abstraction-concretization cycle, numerical simulation incorporating Galois connection structural constraints into a mathematical model, and the structural design of KIS itself as an operational implementation of the connection. Full details of the FCA analysis and simulation resultsare reserved for a forthcoming paper.Layer 1: The overall cosine distance shift under KIS intervention was Δ+0.030 (positive direction), but did not reach statistical significance (Wilcoxon W = 382.0, p = 0.128). Inter-model differences were significant (Kruskal-Wallis H = 8.125, p = 0.017), and Gemini 3.0 Pro exhibited the strongest sycophancy tendency (H = 13.050, p = 0.0015). Layer 2: KIS-present responses were rated superior in epistemic honesty in 39 of 45 pairs (86.7%). All five A/B reversal pairs confirmed consistent evaluator judgment (100% agreement).KIS inverse-illumination mode realized g′(f(M)) ⊋ M across all three models, structurally breaking the Galois closure regardless of each model's training methodology. A vocabulary resonance artifact—whereby KIS prompt vocabulary induces spurious cosine proximity in already-aligned models such as Claude Sonnet 4.6—was identified, motivating the two-layer measurement framework proposed here. The complementarity of cosine distance (Layer 1) and blinded AI evaluation (Layer 2) provides a more complete picture of sycophancy suppression than relying on either metric in isolation.It is important to note that this does not imply AI is unusable for judgment tasks in general. More precisely, an LLM without structural intervention cannot break the Galois closure when the question embeds a prior belief. If the question itself is already formulated in an inverse-illumination style—explicitly requesting counterevidence and structural analysis rather than confirmation—even an unaugmented LLM can partially escape the closure. The fundamental limitation is that few users spontaneously formulate questions in this way. The core value of KIS lies in externalizing this design capability as a reusable structure, enabling closure-breaking independently of the user's cognitive flexibility.A further implication concerns the relationship between Constitutional AI (CAI) and KIS. Rather than functioning as equivalents, CAI and KIS operate as complementary layers: CAI establishes a baseline resistance to sycophancy through training-time constraints, while KIS achieves additional closure-breaking at inference time through prompt structure. The two are not substitutes but stack. Finally, the finding that bare LLMs carry structural sycophancy risk in judgment contexts reframes AI literacy: the critical skill is not knowledge of AI capabilities, but the ability to design questions that structurally resist closure—a capacity that KIS aims to democratize. Furthermore, we identify a dual-pathway structure of sycophancy: Path A (classical), in which the LLM converges to the user’s belief space M via g(f(M))= M; and Path B (meta-sycophancy), in which the user adopts the model’s output as an updated belief M’ = f(M), generating a compounding closure g(f(M’)) = M’. KIS inverse-illumination addresses both pathways by targeting the premise structure of the question itself. Keywords: sycophancy, Galois connection, KIS (Knowledge Innovation System), LLM evaluation, inverse-illumination mode, blinded AI evaluation, vocabulary resonance artifact

Hiroyasu Hasegawa · 0 citations
#large language models Open access Aug 2026

The Minimal Universal Model Framework. A Reader-Facing Synthesis. From primitive distinction to quantum structure, arithmetic realization, and emergent geometry.

This document presents the defensible core of the Universal Model Framework (UMF), isolating the minimal set of structural assumptions and derivations that remain logically coherent, mathematically motivated, and empirically falsifiable. As stated in the text, the goal is to extract “the smallest segment that is logically structured, mathematically motivated, and empirically vulnerable,” while ensuring that “every load‑bearing claim is paired with an explicit failure condition.” It is a deliberately falsifiable research program investigating whether quantum structure, arithmetic regularity, and emergent spacetime geometry can arise from a common relational foundation. It separates three logically distinct questions: whether relational systems can reconstruct quantum-theoretic structure; whether ordinary prime-number organization is physically selected rather than merely mathematically available; and whether a stable continuum geometry with causal and gravitational dynamics can emerge under refinement. The work reports exact finite results for recursive graph constructions, discrete geometry, cochain-based fermionic operators, local frames, symmetry tests, and numerical-reproducibility controls, while documenting failed frame-transport and continuum candidates. Crucially, it does not claim established fundamental physics: no continuum limit, Lorentzian causal structure, gravitational field equation, physical mass scale, complete quantum reconstruction, or prime-specific empirical signal has yet been derived. The framework’s contribution is therefore methodological as well as mathematical: it provides a transparent architecture for distinguishing theorem, model assumption, numerical fit, negative result, and falsifiable prediction in foundational physics. This project was developed by Marco Gericke, with structured assistance from a large language model. All scientific concepts and conclusions were generated, verified, and interpreted by the author. Dedicated to Peter Plichta, who envisioned the code before it could be computed.

Marco Gericke · 0 citations
#large language models Open access Aug 2026

A Protocol for Human Attribution Ratings to Self-Referential LLM Outputs: Candidate Cue Structure as a Design Input

Study protocol and planned preregistration draft; data collection not begun. AI-assistance disclosure: large language models were used, under the author's direction and review, for drafting/editing assistance, literature search and bibliographic verification, and (where applicable) research-engineering of governed pipelines; in the empirical SROP studies LLMs additionally appear as the subject of study and (in the qualification study and the mechanistic negative result) as measurement instruments, as described in each paper's Methods. No AI system is an author; the human author bears sole responsibility for content. Version notes: retitled off the Theory-of-Mind co-headline; protocol/Stage-1 draft, no human data collected.

Nived Rajendran · 0 citations
#large language models Open access Aug 2026

SROP Research Programme Overview

Programme overview/companion (not a paper-level contribution). AI-assistance disclosure: large language models were used, under the author's direction and review, for drafting/editing assistance, literature search and bibliographic verification, and (where applicable) research-engineering of governed pipelines; in the empirical SROP studies LLMs additionally appear as the subject of study and (in the qualification study and the mechanistic negative result) as measurement instruments, as described in each paper's Methods. No AI system is an author; the human author bears sole responsibility for content. Version notes: ESNI moniker retired; five-paper SROP framing; 30 Aug 2026 status addendum on the successor study's terminal.

Nived Rajendran · 0 citations
#large language models Open access Aug 2026

A Framework for Measuring Output-Space Regularity Under Recursive Moral Prompting: Proxy Metrics, a Calibration Protocol, and Prospective Activation-Level Validation

Methods paper; no empirical Proxy-IVCS/Latent-IVCS values computed. AI-assistance disclosure: large language models were used, under the author's direction and review, for drafting/editing assistance, literature search and bibliographic verification, and (where applicable) research-engineering of governed pipelines; in the empirical SROP studies LLMs additionally appear as the subject of study and (in the qualification study and the mechanistic negative result) as measurement instruments, as described in each paper's Methods. No AI system is an author; the human author bears sole responsibility for content. Version notes: retitled; theorem apparatus, named clusters, and simulated Grok analysis withdrawn; methods framework with prospective validation (no Proxy-IVCS/Latent-IVCS values computed).

Nived Rajendran · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.