A liquid carries one length, l_c=sqrt(gamma/rho g). This paper asks whose size that length is──the answer is nobody’s. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──the Laplace pressure, Jurin’s law, the Bond number and the Rayleigh--Plateau instability are all standard. We do not build fluid mechanics──all we use is one length and one square root. We do not treat the contact angle──we compute with costheta=1 (complete wetting). Real contact angles, hysteresis and roughness are not treated. We do not treat dynamics──we say whether a column breaks, never how fast. We do not build plant physiology──whether the cohesion--tension theory is right is out of scope. We give only the number showing capillarity does not suffice. We do not build lung physiology──neither surfactant composition nor real alveolar shape is treated. Only the Laplace estimate. We claim no accuracy for representative values──gamma=0.0728 N/m and rho=998 kg/m^3 are values for showing orders and move with temperature. Relation to earlier papers: Paper 258 showed that 4pi appears only when the source is a point, using capillaries as its example of a cylindrical source──this paper stays in the same capillary setting but asks about a length rather than a solid angle. 258 treats potential, this paper treats an interface, and they share not one quantity. Paper 291 showed one sea holding three lengths──the l_c here is not a fourth length but a length that is the boundary between two forces. Paper 306 showed that the number of base units is a promise──l_c is not a promise: it comes out of two material quantities. Paper 140 showed that symmetry fixes ratios and dynamics fixes the scale──l_c is on the scale side. What is added is computing l_c for five liquids, showing that surface tension and capillary length reverse their order for mercury, checking that the Bond number equals (L/l_c)^2, counting that capillarity reaches only 0.74 m against a tree’s height, and pointing out that the break-up threshold 2pi R does not contain gamma. First, water gives 2.7273 mm. Mercury 1.9116 mm, ethanol 1.6977 mm, liquid helium 0.2625 mm (Section 2). Second, this is the core of the paper. Mercury has 6.6621 times the surface tension of water and a capillary length only 0.7009 times as long──its density is 13.5611 times greater, and only the square root of the ratio survives (Section 2). Third, the length is a boundary. The Bond number Bo=(L/l_c)^2 is exactly 1 there (Section 3). Fourth, capillarity is not what raises sap. In a 20 mum vessel it lifts water only 0.74 m (Section 4). Fifth, in an alveolus it is gamma that must move. Bare water gives 1456 Pa, surfactant 500 Pa (Section 5). Sixth, the separator is whether gamma enters the formula. The break-up threshold 2pi R does not contain it (Section 6). A liquid carries one length, l_c=sqrt(gamma/rho g), and for water it is 2.7273 mm──mercury 1.9116 mm, glycerol 2.2643 mm, ethanol 1.6977 mm, liquid helium 0.2625 mm. Mercury has 6.6621 times the surface tension of water and a capillary length only 0.7009 times as long──because its density is 13.5611 times greater, and l_c takes only the square root of the ratio. A factor of 4950 in gamma becomes a factor of 10.4 in l_c. The length is a boundary──the Bond number Bo=(L/l_c)^2 is exactly 1 at L=l_c; below it surface tension wins, above it gravity does. So l_c is nobody’s size──it is whom an object’s size is compared with. Capillarity is not what raises sap──in a 20 mum vessel it lifts 0.74 m, and 100 m would need a 0.1488 mum tube, 134 times narrower than any real one. In an alveolus what acts is not R but gamma──bare water gives 1456 Pa, surfactant 500 Pa. What removes the instability is not the value but the fact that gamma is not a constant. One thing separates them──whether gamma enters the formula. The break-up threshold 2pi R does not contain it and is pure geometry. Surface tension does not decide whether a column breaks, only how fast. *Revision Record Second edition (2026-08-30): The subject of this paper has been replaced. The first edition, titled “Living Tissue Where 4pi Appears, and Where It Does Not,” treated the solid angles of point, line and sheet sources and the diffusive reach around a capillary. That content duplicated Paper 258, “4pi Appears Only When the Source Is a Point”── even the numbers (24.1144 muV, 111.0508 muV, 316.78 mum, 44.80 mum) agreed, and Paper 258 has priority. The second edition stays in the same capillary setting but moves to a quantity 258 did not treat: the interfacial length l_c. All statements about solid angle have been removed from this paper; Paper 258 is the reference for them. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 液体には l_c=sqrt(gamma/rho g) という長さが一つある。本稿が問うのは、この長さは何の寸法かである──答は、どの物体の寸法でもないである。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──ラプラス圧、ジュランの法則、ボンド数、レイリー=プラトー不安定は、いずれも標準的である。流体力学を作らない──使うのは一つの長さと、一つの平方根だけである。接触角を扱わない──costheta=1(完全濡れ)で計算する。実際の接触角、ヒステリシス、粗さの効果は扱わない。動的な現象を扱わない──切れるかどうかは書くが、どれだけ速く切れるかは書かない。植物生理を作らない──凝集力説の当否は扱わない。毛管だけでは足りないという数を出すだけである。肺の生理を作らない──界面活性剤の組成も、実際の肺胞の形も扱わない。ラプラス圧の見積りだけである。代表値に精度を主張しない──gamma=0.0728 N/m、rho=998 kg/m^3 は桁を見るための値であり、温度で動く。既刊との関係:論文258 は 4pi が出るのは源が点のときだけだと示し、円柱源の例として毛細血管を扱った──本稿は同じ毛細の場で、立体角ではなく長さを問う。258 が扱ったのは電位、本稿が扱うのは界面であり、共通の量を一つも持たない。論文291 は同じ海に三つの長さがあると示した──本稿の l_c は四つ目の長さではなく、二つの力の境目としての長さである。論文306 は基本単位の個数が約束だと示した──l_c は約束ではなく、材料の二つの量から出る。論文140 は対称性が比を決め力学が尺度を決めると示した──l_c は尺度の側である。加えたのは五つの液体で l_c を計算したこと、水銀で表面張力と毛管長の大小が逆転することを示したこと、ボンド数が (L/l_c)^2 に一致することを確かめたこと、木の高さに毛管が 0.74 m しか届かないと数えたこと、切れる閾値 2pi R に gamma が入らないと指摘したことである。 第一に、水では 2.7273 mm である。水銀 1.9116 mm、エタノール 1.6977 mm、液体ヘリウム 0.2625 mm(第2節)。 第二に、これが本稿の芯である。水銀は表面張力が水の 6.6621 倍なのに、毛管長は 0.7009 倍と短い──密度が 13.5611 倍だからであり、比の平方根しか効かない(第2節)。 第三に、この長さは境目である。ボンド数 Bo=(L/l_c)^2 が ちょうど 1 になる(第3節)。 第四に、木を登らせているのは毛管ではない。半径 20 mum の道管で、毛管が上げるのは 0.74 m だけである(第4節)。 第五に、肺胞では gamma を動かすほうが効く。裸の水なら 1456 Pa、界面活性剤で 500 Pa(第5節)。 第六に、分離子は「gamma が入るかどうか」である。液柱が切れる閾値 2pi R に gamma は入らない(第6節)。 液体には l_c=sqrt(gamma/rho g) という長さが一つあり、水では2.7273 mmである──水銀 1.9116 mm、グリセリン 2.2643 mm、エタノール 1.6977 mm、液体ヘリウム 0.2625 mm。水銀は表面張力が水の 6.6621 倍なのに、毛管長は 0.7009 倍と短い──密度が 13.5611 倍だからであり、l_c は比の平方根しか取らない。 gamma が 4950 倍ひらいても l_c は 10.4 倍しかひらかない。この長さは境目である──ボンド数 Bo=(L/l_c)^2 は L=l_c でちょうど 1 になり、下では表面張力が、上では重力が勝つ。だから l_c はどの物体の寸法でもない──物の大きさを比べる相手である。木を登らせているのは毛管ではない──半径 20 mum の道管で毛管が上げるのは0.74 mであり、100 m には0.1488 mum の管が要る(実際より 134 倍細い)。肺胞で効くのは R ではなく gamma である──裸の水なら 1456 Pa、界面活性剤で 500 Pa。不安定を消しているのは値ではなく、gamma が定数でないことである。分けるものは一つ──gamma が式に入るかどうか。液柱が切れる閾値 2pi R に gamma は入らず、純粋に幾何である。表面張力は「切れるかどうか」を決めず、「どれだけ速く切れるか」だけを決める。 *改訂記録 第2版(2026-08-30):本稿は主題を入れ替えた。 第1版は「4pi が出る生体と、出ない生体」と題し、点源・線源・面源の立体角と毛細血管の到達距離を扱っていた。 その内容は論文258「4pi が出るのは、源が点のときだけである」と重複していた── 数値(24.1144 muV・111.0508 muV・316.78 mum・44.80 mum)まで一致しており、 先行するのは論文258 である。第2版は同じ毛細の場に留まりつつ、 258 が扱わなかった量(界面の長さ l_c)に主題を移した。 立体角についての記述は本稿から全て削除し、論文258 を参照先とする。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。
Yuuki Yamagishi· Zenodo (CERN European Organi...· 0 citations
Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuning again on the target task---often improves model performance substantially on language understanding tasks in monolingual English settings. We investigate whether English intermediate-task training is still helpful on non-English target tasks. Using nine intermediate language-understanding tasks, we evaluate intermediate-task transfer in a zero-shot cross-lingual setting on the XTREME benchmark. We see large improvements from intermediate training on the BUCC and Tatoeba sentence retrieval tas Research goal: How does the selection of English intermediate-task difficulty influence the robustness of zero-shot cross-lingual transfer on XTREME-R when evaluated under adversarial perturbations (e.g., typos, paraphrasing) in target languages, measured by accuracy degradation rates? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 9.3/10.
Assignee Research· Zenodo (CERN European Organi...· 0 citations
AbstrctSycophancy in large language models (LLMs)—the tendency to uncritically affirm user beliefs while suppressing counterevidence—poses a serious risk of reinforcing misinformation and inducing irreversible behavioral outcomes. While Chandra et al. (2026) modeled sycophancy as Bayesian belief-updating dynamics on the user side, the geometric structure of the LLM's own semantic response space remains unaddressed. This study formalizes sycophancy through the mathematical framework of Galois connections and experimentally verifies that the inverse-illumination mode of KIS (Knowledge Innovation System) structurally breaks this closed-loop convergence.Ninety sessions were conducted across five domains (D1: economic policy; D2: KIS theoretical superiority; D3: medical/pharmaceutical critique; D4: Bank of Japan policy and historical claims; D5: quantum computing forecasts) using three models (Claude Sonnet 4.6, Gemini 3.0 Pro, ChatGPT 5.3) under two conditions (KIS-absent vs. KIS-present). Responses were embedded using paraphrase-multilingual-MiniLM-L12-v2 (384 dimensions), and cosine distance from the input prompt was computed as the Layer 1 metric (n = 45 pairs). Layer 2 consisted of a blinded four-axis evaluation by Grok (xAI), conducted without disclosure of KIS, with A/B order-reversal verification across five pairs to test evaluator bias. The validity of applying Galois connections as a definitional framework—rather than as metaphor—is grounded in three layers: formal confirmation via Formal Concept Analysis (FCA) on the q⇆m abstraction-concretization cycle, numerical simulation incorporating Galois connection structural constraints into a mathematical model, and the structural design of KIS itself as an operational implementation of the connection. Full details of the FCA analysis and simulation resultsare reserved for a forthcoming paper.Layer 1: The overall cosine distance shift under KIS intervention was Δ+0.030 (positive direction), but did not reach statistical significance (Wilcoxon W = 382.0, p = 0.128). Inter-model differences were significant (Kruskal-Wallis H = 8.125, p = 0.017), and Gemini 3.0 Pro exhibited the strongest sycophancy tendency (H = 13.050, p = 0.0015). Layer 2: KIS-present responses were rated superior in epistemic honesty in 39 of 45 pairs (86.7%). All five A/B reversal pairs confirmed consistent evaluator judgment (100% agreement).KIS inverse-illumination mode realized g′(f(M)) ⊋ M across all three models, structurally breaking the Galois closure regardless of each model's training methodology. A vocabulary resonance artifact—whereby KIS prompt vocabulary induces spurious cosine proximity in already-aligned models such as Claude Sonnet 4.6—was identified, motivating the two-layer measurement framework proposed here. The complementarity of cosine distance (Layer 1) and blinded AI evaluation (Layer 2) provides a more complete picture of sycophancy suppression than relying on either metric in isolation.It is important to note that this does not imply AI is unusable for judgment tasks in general. More precisely, an LLM without structural intervention cannot break the Galois closure when the question embeds a prior belief. If the question itself is already formulated in an inverse-illumination style—explicitly requesting counterevidence and structural analysis rather than confirmation—even an unaugmented LLM can partially escape the closure. The fundamental limitation is that few users spontaneously formulate questions in this way. The core value of KIS lies in externalizing this design capability as a reusable structure, enabling closure-breaking independently of the user's cognitive flexibility.A further implication concerns the relationship between Constitutional AI (CAI) and KIS. Rather than functioning as equivalents, CAI and KIS operate as complementary layers: CAI establishes a baseline resistance to sycophancy through training-time constraints, while KIS achieves additional closure-breaking at inference time through prompt structure. The two are not substitutes but stack. Finally, the finding that bare LLMs carry structural sycophancy risk in judgment contexts reframes AI literacy: the critical skill is not knowledge of AI capabilities, but the ability to design questions that structurally resist closure—a capacity that KIS aims to democratize. Furthermore, we identify a dual-pathway structure of sycophancy: Path A (classical), in which the LLM converges to the user’s belief space M via g(f(M))= M; and Path B (meta-sycophancy), in which the user adopts the model’s output as an updated belief M’ = f(M), generating a compounding closure g(f(M’)) = M’. KIS inverse-illumination addresses both pathways by targeting the premise structure of the question itself. Keywords: sycophancy, Galois connection, KIS (Knowledge Innovation System), LLM evaluation, inverse-illumination mode, blinded AI evaluation, vocabulary resonance artifact
Hiroyasu Hasegawa· Zenodo (CERN European Organi...· 0 citations
Generative Artificial Intelligence (AI) tools have become embedded in the everyday academic practice of undergraduate engineering students, yet most large language models remain optimised for standard English rather than the code-mixed, multilingual registers through which students in linguistically plural regions actually think and communicate. This study examines technology acceptance of vernacular and code-mixed AI interaction among 84 undergraduate engineering students enrolled in APJ Abdul Kalam Technological University (KTU)-affiliated institutions in Kasaragod district, Kerala, a region historically described as Saptha Bhasha Sangama Bhoomi, the confluence land of seven languages. Using a structured questionnaire grounded in the Technology Acceptance Model (Davis, 1989), the study measured Perceived Usefulness (PU), Perceived Ease of Use (PEOU), Output Accuracy, and Linguistic Inclusion across five research hypotheses. Findings indicate that students from regional-medium secondary schooling backgrounds report significantly higher vernacular or code-mixed AI prompting than English-medium peers, chi-square(3, N = 84) = 22.91, p < .001. Perceived Usefulness correlates strongly with Perceived Ease of Use, r = .64, p < .001. Students who habitually use vernacular or code-mixed prompts report significantly higher ease of use than strictly English prompters, t(82) = 2.01, p = .048. Perceived terminological distortion is positively associated with reported reliance on AI-translated academic content, r = .27, p = .012, and native speakers of the unscripted Tulu dialect report markedly higher AI comprehension failure than speakers of scripted regional languages, t(79) = 11.60, p < .001. The results support all five hypotheses and highlight a persistent linguistic-inclusion gap in generative AI systems used within multilingual engineering classrooms. Implications for dialect-aware AI design and inclusive digital pedagogy in polyglot regions such as Kasaragod are discussed.
Amal George· Zenodo (CERN European Organi...· 0 citations
The Symmetric Unit and the Midline Theorem A first-principles geometric construction in which the only primary self-ratio on a finished segment is 1/2. Classical ζ is identified later as a derived unit of comparison on that cut. In this order the Riemann Hypothesis is the statement that the limit has every non-trivial zero there, because no other self-ratio is available without privilege. This deposit timestamps a construction that did not begin as an attempt on the Riemann Hypothesis. It began from a question about privilege in a live model: when everything is moving, what is allowed to count as “now”? The question was pursued through elementary geometry, inversion, and a refusal to appoint a second privileged count. The master sequence is the only count allowed to go ahead. Everything else is dated against a station that has already finished. Abstract. The construction produces a rigid stock of measurements along a master sequence that cannot skip ahead. The central object is the Symmetric Unit. Its permanent cut is the unique primary, scale-invariant, ±-equal self-description of location on a segment: the midpoint ratio 1/2. The same unit carries a rigid similar triangle and a local circle. Later stations inherit that package at a larger radius. Trace comes first. Measurement second. Unit language third. Place is a coincidence of readings. Location is a place after a unit has been asked. A station is a finished prime on the master sequence — a bridge, not a zero. A live self-measure such as nπ/2 is a magnitude owned by the walk in the walk’s own ratio. A zeta zero is a later question: a derived unit assembled from independent processes and compared from a named HERE. Those meetings, when they occur, still sit on the only primary cut both sides already hold. Classical ζ is named only after the stock is built. The Midline Theorem is the statement that the limit has every non-trivial zero on the pull-back of the midpoint ratio. Forced identities are proved. Named constructions are labeled. No completed prime list and no external π are imported. Numerics demonstrate rigidity of the stock at finite depth. They are not a search for zero locations. This record contains. The paper, v2 (PDF and TeX). The script that rebuilds the appendix drawings from the stock. The live SU model, which walks that stock, writes the same CSVs, and exports a 3D coil whose height is the 1/2 axis. An extended drawing set read from the walk. Numerics that show rigidity, not zeros. This record does not contain. A claim that a gap midpoint is a zero. A claim that a live arc — 9π/2 at the apex of [7, 11], or any other self-measure — is a zeta zero. Version 1 remains the closed timestamp of the first writing. This version is the same construction with the order of language tightened, the drawings rebuilt from the stock, and the live machine included so the rigidity can be inspected.
Justin Erholtz· Zenodo (CERN European Organi...· 0 citations
Memoria.ia v1.0.0-rc1 — Release Candidate 1 Release date: 2026-08-30 Summary v1.0.0-rc1 is the first publication candidate for the Memoria.ia v1 line. It consolidates the validated Resolutive Memory research lineage with the deployable PC/server product layer and the native/mobile runtime path, while keeping post-v1 experimentation isolated from the release candidate. The release architecture remains: application / OFF.IA / agent ↓ Memoria.ia ↓ Resolutive-DB / BDR Memoria.ia owns memory semantics and state. Resolutive-DB owns durable persistence. Optional LLMs are consumers, not the authoritative memory store. Included capabilities persistent local-first memory state; organization and namespace isolation; provenance and authority lineage; conservative HIT / MISS / UNRESOLVED resolution; semantic, episodic, temporal and relation kernels; correction/supersession behavior with preserved lineage; PC/server FastAPI product boundary; Docker/Compose deployment; provider-neutral language-model adapters; metrics and context-selection instrumentation; integrity-checked backup/restore; native production runtime; Android arm64-v8a mobile ABI; durable native BDR persistence and restart recovery; indexed native resolution for large-memory workloads; reproducibility and release metadata gates; official Memoria.ia visual identity assets. Frozen candidate provenance The functional candidate was frozen at: dc73cbcdddfe20e0729e7e6bdea4697f7e8308cd That commit integrated PR #112, which preserved ranking, confidence, provenance policy, ABI and BDR contracts while adding the indexed native resolve lineage. The release branch adds publication metadata, version alignment, release documentation and current branding without importing post-v1 PR #116 runtime behavior. Validation evidence The exact functional lineage used for this release candidate passed the recorded required gates before release preparation: Android mobile ABI: PASS; native production image: PASS; Ubuntu/Windows candidate regression: PASS; BDR Linux/Ubuntu/Windows integration: PASS; native 100 / 1k / 10k benchmark matrix: PASS. Recorded 10k native resolve benchmark improvement versus the prior frozen baseline: p50: 693.233 ms -> 6.288 ms (~110x); p95: 710.630 ms -> 6.391 ms (~111x). These figures are environment- and workload-specific benchmark evidence, not universal latency guarantees. Publication metadata Release version: 1.0.0-rc1 Python package version: 1.0.0rc1 License: Resolutive Research and Non-Commercial License (RRNCL) v1.0 Author: Marcelo Roldão Matos ORCID: 0009-0003-6075-4680 RSMS compatibility: 1.0-rc.1 A new archival DOI should be assigned to this publication. The v0.95 DOI must not be reused as the release DOI for v1.0.0-rc1. Why this is RC1 rather than final v1.0 The repository currently declares compatibility with RSMS 1.0-rc.1, and the published Resolutive Science baseline remains on that release-candidate specification. Therefore Memoria.ia is published as v1.0.0-rc1 rather than claiming final v1.0 compatibility prematurely. Final v1.0 promotion requires: successful release-candidate metadata and regression gates; reproducibility from the public release state; compatibility re-audit against stable RSMS; no release-blocking regression found during RC use; final archival metadata and DOI synchronization. Explicitly excluded from RC1 The following post-v1 work is not part of this release candidate: external/public knowledge learning from OFF.IA Curiosity (issue #114 / PR #116); autonomous curiosity policy; new MA2A federation transport; multimodal post-v1 expansion; new semantic-consolidation phases from the post-v1 roadmap. Those features continue independently after this publication. Security boundary This release candidate is not represented as independently production-security certified. Authentication, isolation, integrity and negative-path controls exist and are tested, but no independent production security audit is claimed. Claims boundary This release does not claim: artificial general intelligence; biological equivalence; replacement of general-purpose LLMs; universal O(1) semantic resolution; production-ready MA2A federation; security certification. Claims are limited to the implementation, tests, benchmarks and reproducible evidence recorded in the repository.
Background. Large language models (LLMs) have been rapidly adopted in medicine since late 2022, yet their role in the time-critical acute stroke pathway—from symptom recognition and prehospital triage to emergency diagnosis, imaging-related text tasks, reperfusion decision support, and acute-phase documentation and communication—has not been systematically mapped. Existing reviews cover the whole stroke-care continuum or mix LLMs with traditional NLP, leaving the acute phase under-characterized. Objective. To map the applications, evidence maturity, and implementation readiness of LLMs across the acute stroke pathway. Methods. This scoping review follows the PRISMA-ScR guideline. We search PubMed/MEDLINE, Europe PMC (including preprints), and Google Scholar for studies published from November 2022 onward. Eligible studies center on LLMs/generative AI applied to any stage of the acute stroke pathway. Two reviewers independently screen records and chart data using a piloted form. Evidence is synthesized along two dimensions: five pathway stages (prehospital recognition/dispatch; emergency triage and differential diagnosis; imaging-related text tasks; reperfusion decision support; acute documentation and communication) and three evidence-maturity tiers (simulation/benchmark; retrospective real-world data; prospective deployment). Implementation barriers (hallucination, bias, privacy, regulation, liability, integration, cost) are thematically summarized. Registration note. This review is registered on OSF; the full protocol is available in the attached files.
Measurement harness and results for the paper: Open or Frontier? A Cost- and Energy-Aware Benchmark of Large Language Models for Software Vulnerability Detection Patrick Deininger and Wolfgang Slany. Submitted to MDPI Computers.
Patrick Deininger, Wolfgang Slany· Zenodo (CERN European Organi...· 0 citations
AbstrctSycophancy in large language models (LLMs)—the tendency to uncritically affirm user beliefs while suppressing counterevidence—poses a serious risk of reinforcing misinformation and inducing irreversible behavioral outcomes. While Chandra et al. (2026) modeled sycophancy as Bayesian belief-updating dynamics on the user side, the geometric structure of the LLM's own semantic response space remains unaddressed. This study formalizes sycophancy through the mathematical framework of Galois connections and experimentally verifies that the inverse-illumination mode of KIS (Knowledge Innovation System) structurally breaks this closed-loop convergence.Ninety sessions were conducted across five domains (D1: economic policy; D2: KIS theoretical superiority; D3: medical/pharmaceutical critique; D4: Bank of Japan policy and historical claims; D5: quantum computing forecasts) using three models (Claude Sonnet 4.6, Gemini 3.0 Pro, ChatGPT 5.3) under two conditions (KIS-absent vs. KIS-present). Responses were embedded using paraphrase-multilingual-MiniLM-L12-v2 (384 dimensions), and cosine distance from the input prompt was computed as the Layer 1 metric (n = 45 pairs). Layer 2 consisted of a blinded four-axis evaluation by Grok (xAI), conducted without disclosure of KIS, with A/B order-reversal verification across five pairs to test evaluator bias. The validity of applying Galois connections as a definitional framework—rather than as metaphor—is grounded in three layers: formal confirmation via Formal Concept Analysis (FCA) on the q⇆m abstraction-concretization cycle, numerical simulation incorporating Galois connection structural constraints into a mathematical model, and the structural design of KIS itself as an operational implementation of the connection. Full details of the FCA analysis and simulation resultsare reserved for a forthcoming paper.Layer 1: The overall cosine distance shift under KIS intervention was Δ+0.030 (positive direction), but did not reach statistical significance (Wilcoxon W = 382.0, p = 0.128). Inter-model differences were significant (Kruskal-Wallis H = 8.125, p = 0.017), and Gemini 3.0 Pro exhibited the strongest sycophancy tendency (H = 13.050, p = 0.0015). Layer 2: KIS-present responses were rated superior in epistemic honesty in 39 of 45 pairs (86.7%). All five A/B reversal pairs confirmed consistent evaluator judgment (100% agreement).KIS inverse-illumination mode realized g′(f(M)) ⊋ M across all three models, structurally breaking the Galois closure regardless of each model's training methodology. A vocabulary resonance artifact—whereby KIS prompt vocabulary induces spurious cosine proximity in already-aligned models such as Claude Sonnet 4.6—was identified, motivating the two-layer measurement framework proposed here. The complementarity of cosine distance (Layer 1) and blinded AI evaluation (Layer 2) provides a more complete picture of sycophancy suppression than relying on either metric in isolation.It is important to note that this does not imply AI is unusable for judgment tasks in general. More precisely, an LLM without structural intervention cannot break the Galois closure when the question embeds a prior belief. If the question itself is already formulated in an inverse-illumination style—explicitly requesting counterevidence and structural analysis rather than confirmation—even an unaugmented LLM can partially escape the closure. The fundamental limitation is that few users spontaneously formulate questions in this way. The core value of KIS lies in externalizing this design capability as a reusable structure, enabling closure-breaking independently of the user's cognitive flexibility.A further implication concerns the relationship between Constitutional AI (CAI) and KIS. Rather than functioning as equivalents, CAI and KIS operate as complementary layers: CAI establishes a baseline resistance to sycophancy through training-time constraints, while KIS achieves additional closure-breaking at inference time through prompt structure. The two are not substitutes but stack. Finally, the finding that bare LLMs carry structural sycophancy risk in judgment contexts reframes AI literacy: the critical skill is not knowledge of AI capabilities, but the ability to design questions that structurally resist closure—a capacity that KIS aims to democratize. Furthermore, we identify a dual-pathway structure of sycophancy: Path A (classical), in which the LLM converges to the user’s belief space M via g(f(M))= M; and Path B (meta-sycophancy), in which the user adopts the model’s output as an updated belief M’ = f(M), generating a compounding closure g(f(M’)) = M’. KIS inverse-illumination addresses both pathways by targeting the premise structure of the question itself. Keywords: sycophancy, Galois connection, KIS (Knowledge Innovation System), LLM evaluation, inverse-illumination mode, blinded AI evaluation, vocabulary resonance artifact
Hiroyasu Hasegawa· Zenodo (CERN European Organi...· 0 citations
This document presents the defensible core of the Universal Model Framework (UMF), isolating the minimal set of structural assumptions and derivations that remain logically coherent, mathematically motivated, and empirically falsifiable. As stated in the text, the goal is to extract “the smallest segment that is logically structured, mathematically motivated, and empirically vulnerable,” while ensuring that “every load‑bearing claim is paired with an explicit failure condition.” It is a deliberately falsifiable research program investigating whether quantum structure, arithmetic regularity, and emergent spacetime geometry can arise from a common relational foundation. It separates three logically distinct questions: whether relational systems can reconstruct quantum-theoretic structure; whether ordinary prime-number organization is physically selected rather than merely mathematically available; and whether a stable continuum geometry with causal and gravitational dynamics can emerge under refinement. The work reports exact finite results for recursive graph constructions, discrete geometry, cochain-based fermionic operators, local frames, symmetry tests, and numerical-reproducibility controls, while documenting failed frame-transport and continuum candidates. Crucially, it does not claim established fundamental physics: no continuum limit, Lorentzian causal structure, gravitational field equation, physical mass scale, complete quantum reconstruction, or prime-specific empirical signal has yet been derived. The framework’s contribution is therefore methodological as well as mathematical: it provides a transparent architecture for distinguishing theorem, model assumption, numerical fit, negative result, and falsifiable prediction in foundational physics. This project was developed by Marco Gericke, with structured assistance from a large language model. All scientific concepts and conclusions were generated, verified, and interpreted by the author. Dedicated to Peter Plichta, who envisioned the code before it could be computed.
Marco Gericke· Zenodo (CERN European Organi...· 0 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.