Skip to content

Category

small language model

678 papers

#small language model Review Open access Aug 2026

Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study

Abstract Background Systematic literature reviews (SLRs) are essential for evidence synthesis in health research but remain labor-intensive, especially at the screening stage. Manual review of titles and abstracts requires substantial human effort, while existing automation tools still have limited adoption in health technology assessment. The EQ-5D questionnaire, a widely used patient-reported outcome measure for health-related quality of life, provides data that frequently underpin reimbursement and policy decisions. Objective This pilot study evaluated whether recent large language models (LLMs) can support the identification of publications reporting EQ-5D data in PubMed records, using only publicly available metadata (title, abstract, and keywords). Methods A total of 200 publications retrieved through the EuroQol PubMed filter were manually labeled by experts as reporting or not reporting EQ-5D data. The dataset was split into stratified training, validation, and test subsets. Several machine learning approaches were compared, including a Naïve Bayes baseline using bag-of-words features, a decision-tree model based on full-text keyword occurrence, and transformer-based LLMs (Bidirectional Encoder Representations from Transformers [BERT], Biomedical BERT [BioBERT], Scientific BERT [SciBERT], and Biomedical Language Understanding Evaluation BERT [BlueBERT]). Both classifier-only and fine-tuned configurations were tested across multiple learning rates. Model performance was assessed using accuracy, precision, recall, and F1-score. Results Baseline approaches achieved near-random test performance (accuracy around 0.53). Classifier-only LLMs modestly improved results (accuracy up to 0.64 with SciBERT). Fine-tuned models substantially outperformed these baselines, with BERT and BioBERT achieving the best performance (accuracy=0.70; F1-score=0.68). In screening-oriented evaluation, this configuration achieved 90.0% sensitivity, 40.0% specificity, and 6 false negatives on the held-out test set. The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection. Conclusions This study provides the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature. The findings support technical feasibility but do not establish a reliable stand-alone automated screening tool. Although limited by dataset size, the proposed workflow is reproducible and adaptable to other patient-reported outcome measures. Because validation was based on a single small train-validation-test split, the results should be interpreted as preliminary; future work will scale data collection, include statistical testing, and explore semisupervised learning to further reduce manual screening workload.

Gábor Kertész, J. Czere, Zsombor Zrubka et al. · 0 citations
#small language model Open access Aug 2026

The Length That Sets How Thick an Atmosphere Is Also Sets Whether It Stays ── Earth's Scale Height Is 7.319 km, 0.1149% of the Radius ── The Number Deciding Escape Is the Reciprocal of That Ratio ── [Paper 321]

An atmosphere has one length in it, H=kT/(mg). This paper asks what that length decides──the answer is both the thickness and whether the atmosphere stays. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──scale height, the isothermal barometric law, and the Jeans parameter as an escape criterion are all standard. We do not build atmospheric science──all we use is one length and one ratio. We assume isothermality──a real atmosphere changes temperature with height, and H with it. We compute at one representative temperature. We do not treat escape mechanisms──hydrodynamic escape, non-thermal escape and solar-wind stripping are not entered. lambda is a guide, not a verdict. We do not follow compositional evolution──which gas leaves first, and when, is not treated. We claim no accuracy for the representative values──T and mean molecular weight are numbers for seeing orders of magnitude. Relation to earlier papers: Paper 309 showed that l_c=sqrt(gamma/rho g) is the length dividing gravity from surface tension──this is a third of the same kind, dividing gravity from heat this time. Paper 291 showed three lengths inside one ocean──this paper is the reverse: one length answering two questions. Paper 268 showed that the triple point’s being a “point” is settled by the phase rule before measurement──the same territory, pressure, and this paper looks at the distribution inside a single phase. Paper 276 showed the decibel is a ratio and not a quantity──R/H is a ratio too, and carries no dimension. Paper 316 showed the same integer returning from three counts──this paper is its mirror: one ratio answering two questions. What is added is computing the scale height for five bodies, tabulating H/R as a percentage, giving the pressure profile and the cumulative mass in numbers, and showing that the Jeans parameter equals R/H. First, on Earth it is 7.319 km.0.1149% of the 6371 km radius (Section 2). Second, pressure falls by 1/e every H. At Everest’s height, 0.2985 (Section 3). Third, 63.2121% of the mass sits below one H (Section 3). Fourth, this is the core of the paper. The number deciding escape is R/H itself (Section 4). Fifth, the Moon is at 25.20 and leaks. Earth 871.65, Titan 124.90 (Section 4). Sixth, the separator is which of gravity and heat wins, and it can only be written as a ratio (Section 5). An isothermal atmosphere has one length in it, H=kT/(mg), a ratio of heat to gravity that has come out as a length. On Earth it is 7.319 km──only 0.1149% of the 6371 km radius, 0.17 mm on a 30 cm globe, thinner than a sheet of paper. Pressure falls by 1/e every H, reaching 0.2985 at the summit of Everest──and 63.2121% of the mass sits below a single H. That 1-1/e depends on neither T nor g nor mu, and is the same on Venus, Jupiter and Titan. And the number deciding whether the atmosphere escapes comes out of that same H──putting g=GM/R^2 into the Jeans parameter lambda=GMm/(kTR) makes it exactly lambda=R/H. Earth stays at 871.65, Mars at 313.60, Titan at 124.90, and the Moon leaks at 25.20──Titan is smaller than the Moon and keeps its atmosphere because 94 K of cold compensates for the smallness. One thing separates them──which is larger, kT or mgR. And that can only be written as a ratio. H alone settles nothing──Jupiter’s H exceeds Earth’s, and at R/H=2804.51 it is far safer, so the ordering by length runs opposite to the ordering by safety. This is a third of the kind Paper 309 began with “the length dividing gravity from surface tension is not the size of the material”──H is not the size of the body but what the body is compared with. One last thing──one ratio is answering two questions. “How thin is it” and “does it stay” look different, and the answers are the two faces of one number. A thin atmosphere is an atmosphere in the act of leaving. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 大気には H=kT/(mg) という長さが一つある。本稿が問うのは、この長さは何を決めているのかである──答は、厚さと、残るかどうかの両方である。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──スケールハイト、等温大気の気圧分布、ジーンズ係数による散逸の目安は、いずれも標準的である。大気科学を作らない──使うのは一つの長さと、一つの比だけである。等温を仮定する──実際の大気は高さで温度が変わり、H も高さで変わる。本稿は一つの代表温度で計算する。散逸の機構を扱わない──ジーンズ散逸以外(流体力学的散逸・非熱的散逸・太陽風の剥ぎ取り)には立ち入らない。 lambda は目安であって判定ではない。組成の進化を追わない──どの気体が先に逃げるかの時間発展は扱わない。代表値に精度を主張しない──T と平均分子量は桁を見るための値である。既刊との関係:論文309 は l_c=sqrt(gamma/rho g) が重力と表面張力を分ける長さだと示した──本稿は同じ型の三本目であり、今度は重力と熱を分ける。論文291 は同じ海に三つの長さがあると示した──本稿は一つの長さが二つの問いに答えるという、逆向きの話である。論文268 は三重点が「点」であるのは相律が測る前に決めていると示した──気圧の話をする場所が同じであり、本稿は一相の中の分布を見る。論文276 はデシベルが比であって量ではないと示した──R/H も比であり、次元を持たない。論文316 は同じ整数を三つの数え方が返すと示した──本稿は一つの比が二つの問いに答えるという、鏡像である。加えたのは五つの天体でスケールハイトを計算したこと、H/R を百分率で並べたこと、気圧の高度依存と質量の累積を数で出したこと、ジーンズ係数が R/H に一致することを示したことである。 第一に、地球では 7.319 km である。半径 6371 km の0.1149%(第2節)。 第二に、高度 H ごとに 1/e になる。エベレストの高さで 0.2985 倍(第3節)。 第三に、H 一つぶんの下に 63.2121% が入っている(第3節)。 第四に、これが本稿の芯である。逃げるか残るかを決める数は R/H そのものである(第4節)。 第五に、月は 25.20 で漏れる。地球は 871.65、タイタンは 124.90(第4節)。 第六に、分離子は「重力と熱のどちらが勝つか」であり、それは比でしか書けない(第5節)。 等温の大気には H=kT/(mg) という長さが一つあり、熱と重力の比が長さになったものである。地球では 7.319 km──半径 6371 km の0.1149%にすぎず、直径 30 cm の地球儀なら 0.17 mm、紙一枚より薄い。高度 H ごとに気圧は 1/e になり、エベレストの頂上で 0.2985 倍──そしてH 一つぶんの下に 63.2121% の質量が入っている。この 1-1/e は T にも g にも mu にも依らず、金星でも木星でもタイタンでも同じである。そして大気が逃げるか残るかを決める数は、この H そのものから出る──ジーンズ係数 lambda=GMm/(kTR) に g=GM/R^2 を入れるとlambda=R/H に厳密になる。地球は 871.65、火星は 313.60、タイタンは 124.90 で残り、月は 25.20 で漏れる──タイタンが月より小さいのに大気を持つのは、94 K という冷たさが小ささを補っているからである。分けるものは一つ──kT と mgR のどちらが大きいか。そしてそれは比でしか書けない。 H 単独では何も決まらない──木星の H は地球より大きいのに R/H=2804.51 で地球よりずっと安泰であり、長さの大小と安泰さの大小が逆を向いている。論文309 が「重力と表面張力を分ける長さは材料の寸法ではない」と書いたのと同じ型の三本目である──H は天体の寸法ではなく、比べる相手である。最後に一つ──同じ一つの比が、二つの問いに答えている。「どれだけ薄いか」と「残るか」は別の問いに見えて、答は同じ数の表と裏であった。薄い大気とは、逃げかけている大気のことである。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#small language model Open access Aug 2026

Stretching a Concept to Breaking Point: An Empirical Audit of a Taxonomy of Conceptual Tensions

A glass-box creativity engine, the Concept Collider, rests on a single primitive: a concept is not merely a point but a structure held under tension, and pressing it along one of seven fracture types breaks it in a characteristic way. That taxonomy is load-bearing—every profile, every homology score, every negative-space detector is computed from it—yet it had never been audited. We give it the first empirical audit and find it imperfect: redundant, non-orthogonal, incomplete, and in one case a consequence miscast as a mechanism. From the audit we reconstruct rather than discard: we demote that consequence, promote a recurring self-defeating tension, and propose an empirically grounded tree of fractures. But the deeper finding is that no taxonomy is canonical: standard decomposition criteria each return a different basis for the same data. Computational creativity is a relativistic discipline—novelty is relative to a base, value to an observer, and the tension taxonomy to a method—so the reportable object is not a basis but the invariants: a small core of fractures (circularity, contradiction, self-defeating means) that survives every change of decomposition method and, across five model families spanning both Western and Chinese training ecosystems (Anthropic, Meta, OpenAI, Alibaba, Moonshot), every change of vendor—though the agreement weakens once the tensions are judged in Chinese, so the invariance holds within a language more than across it. That partial cross-ecosystem agreement weakens—without eliminating—the worry that such consensus merely records shared training text; we therefore report the invariants as robust descriptors under our measurement, not as proof that concepts possess a mind-independent structure. The contribution is a transferable criterion—an invariant is what stays put when both method and model family vary. --- Changes in this version (v2): This second version incorporates the mid-2026 literature on inter-model agreement: Ding (2026) audits agreement as a confidence signal and finds it a positive but weak predictor of correctness, while Liu (2026) supplies the mechanism — error decorrelation across independently trained models — and names its ceiling as a shared-error floor. Both are used to state the size of the observer-independence problem rather than only its direction. The approach is also situated in the psychometric lineage of van der Wal et al. (JAIR 79, 2024), who bring construct validity and reliability to bear on bias measures in NLP. Reproducibility artefacts are deposited with this version: the anonymised 380×7 fracture-profile matrix with its provenance metadata, the per-judge classifications from five model families spanning Western and Chinese training ecosystems, and the analysis and figure scripts. Concept names, tension texts and prompts are withheld; the matrix carries opaque identifiers, which changes no published value — verified by recomputation — while keeping the knowledge base of the audited system out of the release. One correction: the adjusted Rand index of the emergence test has been recomputed from the source corpus and revised from 0.02 to 0.065, and the silhouette is reported as flat across every number of clusters rather than monotonically rising. The conclusion is unchanged — the taxonomy does not emerge from the tension descriptions — and a language control (ARI = −0.001) is now reported alongside it.

Sebastian Wahl · 0 citations
#small language model Open access Aug 2026

That the Same Ellipse Was Three Months Out Was the Evidence of a Different Cause ── Aberration Is 20.49552 Arcseconds Whatever the Distance, While Parallax Goes as Its Reciprocal ── A Factor of 26.6442 at Proxima and 20495 at a Kiloparsec ── [Paper 328]

Over a year a star traces a small ellipse on the sky. This paper asks whether that ellipse is one thing──the answer is that there are two, and they are told apart by phase. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──aberration, annual parallax, Bradley’s observation of 1728, and the definition of the constant of aberration are all standard. We do not build celestial mechanics──all we use is two angles and the phase between them. We do not treat this relativistically──the relativistic formula for aberration is not entered. We look only where first order in v/c suffices. We do not treat proper motion──a star’s own motion is not periodic and so separates from both of these. It is mentioned and no more. We do not adjudicate the history──who first saw what is not entered. We do not treat the distance ladder──the rungs beyond parallax belong to Paper 297. Relation to earlier papers: Paper 297 showed the cosmic distance scale to be a ladder of seven rungs with only the first measured directly──that first rung is parallax, and this paper looks at the other ellipse standing beside it. Paper 325 showed the radius not entering the angle of a rainbow──the same form, with distance not entering aberration here. Paper 300 showed whether two things share a root is decidable──this is an instance of “same shape, different root”, and what decided it was not a value but a phase. Paper 327 showed two curves agreeing to within 0.083 per cent──there the appearance is the same and the things differ; here the size can be the same and the things differ. Paper 313 showed the corpus using “synchrony” in several senses──phase is again the tool that decides. What is added is giving the constant of aberration from both a simple calculation and the literature and stating the difference, tabulating the ratio across three decades of distance, solving for the distance at which the two would be equal, and putting the separator on position against velocity. First, aberration is 20.49552 arcseconds (Section 2). Second, this is the core of the paper. Aberration does not depend on distance; parallax goes as its reciprocal (Section 3). Third, even the nearest star opens by 26.6442. At a kiloparsec, 20495 (Section 3). Fourth, amplitude cannot always separate them. At some distance the two would be equal (Section 3). Fifth, their phases differ by a quarter. That is how Bradley told them apart (Section 4). Sixth, the separator is whether the effect is set by position or by velocity (Section 5). Over a year a star traces a small ellipse on the sky, and two causes are possible──annual parallax, because the Earth’s position changes, and aberration, because its velocity does. Both have a period of one year, both are ellipses, and the shape does not distinguish them. Aberration follows from tanalpha=v/c and is 20.49552 arcseconds, about 1/88 of the apparent diameter of the full moon──(the simple calculation here, 20.492697, differs by 0.0138 per cent, because kappa is defined to include the eccentricity, so this paper uses the literature value). And the distance to the star is nowhere in that formula──it contains only the observer’s speed and the speed of light, not one piece of information about the thing being looked at. Even Proxima is out by 26.6442, and at a kiloparsec by 20495.5200. The distance at which the two would be equal is 0.159135 light years, and no such star exists──so on amplitude alone, aberration always wins. What separated them was not amplitude but phase──in circular motion the velocity points a quarter period ahead of the position, so the two ellipses are exactly three months apart. In 1728 Bradley, hunting for parallax, found an ellipse in the wrong season and knew it was not his quarry. A mismatch in amplitude would have been weak evidence──without the distance there is no predicted parallax to compare against, whereas the phase can be predicted without knowing the distance at all. One thing separates them──whether the effect is set by the observer’s position or by the observer’s velocity. Settle that and the distance dependence and the phase follow automatically, so three apparently independent differences are one. Paper 297 wrote that the first rung of the distance ladder is parallax──this paper says that immediately beside it sits an ellipse of the same shape containing no distance at all. The first rung does not stand by itself. One last thing──Bradley failed to find what he was looking for and found what he was not. And it became a measurement of the speed of light, because v/c can be read off from an angle of 20 arcseconds. A failure to find became the ruler for something else. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 恒星は一年かけて空に小さな楕円を描く。本稿が問うのは、その楕円は一種類かである──答は、二種類あり、位相で見分けられるである。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──光行差、年周視差、ブラッドリーの 1728 年の観測、光行差定数の定義はいずれも標準的である。天体力学を作らない──使うのは二つの角度と、その位相差だけである。相対論的な扱いをしない──光行差の相対論的公式には立ち入らない。 v/c の一次までで足りる場面だけを見る。固有運動を扱わない──恒星自身の運動は周期的でないので、本稿の二つとは分けられる。触れるにとどめる。測定の歴史を裁定しない──誰が最初に何を見たかの帰属には立ち入らない。距離はしごを扱わない──視差より遠い段は論文297 の射程である。既刊との関係:論文297 は宇宙の距離が七段のはしごで、直接測られているのは一段目だけだと示した──その一段目が視差であり、本稿はその隣に立つもう一つの楕円を見る。論文325 は虹の角度に半径が入らないと示した──同じ型であり、ここでは光行差に距離が入らない。論文300 は同根か別根かは判定できると示した──本稿は「同じ形をしていても別根」の実例であり、判定に使ったのは値ではなく位相である。論文327 は二つの曲線が 0.083 パーセントしか違わないと示した──そこでは見た目が同じで別物、ここでは大きさが同じでも別物である。論文313 は体系が「同期」を複数の意味で使ってきたと示した──位相という語がここでも判定の道具になっている。加えたのは光行差定数を単純計算と文献値の両方で示し、その差を明示したこと、距離ごとの比を三桁にわたって出したこと、二つが同じ振幅になる距離を求めたこと、分離子を位置と速度の別に置いたことである。 第一に、光行差は 20.49552 秒角である(第2節)。 第二に、これが本稿の芯である。光行差は距離に依らず、視差は距離に反比例する(第3節)。 第三に、最も近い恒星でも 26.6442 倍ひらく。1 キロパーセクなら 20495 倍(第3節)。 第四に、振幅では分けられない場合がある。ある距離では二つが同じ大きさになる(第3節)。 第五に、位相が四分の一ずれている。ブラッドリーはそれで見分けた(第4節)。 第六に、分離子は「位置で決まるか、速度で決まるか」である(第5節)。 恒星は一年かけて空に小さな楕円を描き、その原因は二つありうる──地球の位置が変わるからの年周視差と、地球の速度が変わるからの光行差である。どちらも一年周期の楕円で、形では区別がつかない。光行差は tanalpha=v/c から出て 20.49552 秒角、満月の見かけの直径の約 1/88 である──(本稿の単純計算 20.492697 は文献値と 0.0138 パーセント違い、kappa は離心率を含む定義なので、本稿が使うのは文献値である)。そしてこの式には、恒星までの距離が入っていない──現れるのは観測者の速度と光速だけで、見ている相手の情報が一つも入らない。最も近いプロキシマですら視差の 26.6442 倍、1 キロパーセクでは 20495.5200 倍ひらく。二つが同じ大きさになる距離は 0.159135 光年で、そんな恒星は存在しない──振幅の大小では、常に光行差が勝つ。分けたのは振幅ではなく、位相である──円運動では速度は位置の四分の一周期先を向くので、二つの楕円はちょうど三ヶ月ずれる。ブラッドリーは 1728 年、視差を探していて楕円を見つけ、季節がずれていることで探し物ではないと知った。振幅が合わないことは証拠として弱い──距離を知らなければ視差の予想値が立たないからで、位相は距離を知らなくても予想できる。分けるものは一つ──その効果が観測者の位置で決まるか、速度で決まるか。それが決まれば距離依存も位相も自動的に決まり、独立に見えた三つの違いが一つの違いになる。論文297 は距離のはしごの一段目が視差だと書いた──本稿はその一段目のすぐ隣に、距離を一切含まない同じ形の楕円が乗っていることを言う。はしごの一段目は、単独では立っていない。最後に一つ──ブラッドリーは探していたものを見つけられず、探していなかったものを見つけた。そしてそれは光速の測定になった。20 秒角という角度から v/c が読めるからである。見つからなかったことが、別のものの物差しになった。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#small language model Open access Aug 2026

Two Curves Agree to Within 0.083 Per Cent and Are Not the Same Curve ── What Divides the Catenary from the Parabola Is Not Shape but Which Length the Load Is Counted Along ── [Paper 327]

A hanging chain and a suspension-bridge cable cannot be told apart by eye. This paper asks what divides them──the answer is not the shape, but which length the load is counted along. No new mathematical theorem and no new law is claimed. Scope of this paper (scope note): No new mathematical theorem and no new law is claimed──the catenary equation, the parabolic-cable theory of suspension bridges, and the Taylor expansions of both are all standard. We do not do structural design──all we use is two curves and one difference of measure. We do not treat deflection──cable stretch, thermal movement and live-load deflection are not entered. This paper looks only at the idealised static shape. We do not adjudicate which is correct──a real bridge lies between the two, and this paper does not locate it. We claim no engineering accuracy──the sag ratios are values for comparing shapes, not those of any particular bridge. We do not discuss optimal shape──which is structurally better is not treated. Relation to earlier papers: Paper 198 showed that “straight” does not mean “shortest”──198 says one curve has two characters; this says two curves look alike. The direction is reversed. Paper 326 showed the square-cube relation is an identity──both are of the “same appearance, different content” kind, but 326 is about a ratio and this is about functions. Paper 300 showed whether two things share a root is decidable──this is a worked instance of “numerical closeness is not sameness of root”. Paper 245 showed two limits have no answer until their order is written──here too, the curve is undetermined until the measure of integration is written. Paper 183 showed one table having two correct answers──one shape having two correct equations is its counterpart. What is added is computing the difference at six points, stating that it begins at fourth order, giving the maximum difference for each sag ratio, and putting the separator on the measure. First, they agree exactly through second order. At x=0.1 the difference is 0.083292 per cent (Section 2). Second, far out they are different things. At x=5, 82.925818 per cent (Section 2). Third, the difference begins at fourth order. x^4/24 against 0 (Section 2). Fourth, this is the core of the paper. What divides them is counting along arc length or along horizontal length (Section 3). Fifth, on a real bridge the difference is only 0.082674 per cent (Section 4). Sixth, the separator is whether the deck or the cable is the heavier (Section 5). A hanging chain and a suspension-bridge cable cannot be told apart by eye. Expanding cosh shows the leading term to be the parabola itself, so they agree exactly through second order──with a=1 the difference is 0.083292 per cent at x=0.1 and 82.925818 per cent at x=5, and both figures are correct about the same two curves. “How different are they” is undetermined until you say where you are looking. The difference begins at fourth order──first, second and third all agree, and the parabola is the second-order Taylor polynomial of the catenary, not a curve fitted to it. And what divides them is not shape but measure──both come from the same balance H y''=w, and all that differs is whether w is constant against arc length ds or horizontal length dx. Constant along the arc makes the steep ends weigh more per horizontal foot, and so the catenary rises faster. So “a parabola close to a catenary” has it backwards──each is the exact answer to a different physics, and the approximation relation is something mathematics noticed afterwards. At the 1/10 sag of a real suspension bridge the maximum difference is 0.082674 per cent of the sag──8.3 cm on a 1000 m span, sometimes smaller than the construction tolerance. And they are still two different curves. One thing separates them──whether the suspended deck or the cable itself is heavier. A real bridge is neither ideal, and this paper does not locate it. Against Paper 198 the direction is exactly reversed──198 says one curve has two characters, this says two curves have one appearance, and both obey the same discipline: appearance is not evidence. One last thing──Galileo wrote that a hanging chain is a parabola. He was wrong, by 0.083 per cent. Not an error the eye could correct, and with no cosh yet in existence there was nothing to correct it to. Writing the right shape needs the words for it to exist first. On the making of this work: The ideas and content of this work stem from the author's own considerations. Assistance from an AI (a large language model) was used for structuring, English translation, and checking the algebra. Any remaining errors or misinterpretations are solely the author's. Feedback and corrections are sincerely appreciated. ----- 垂れた鎖の形と、吊り橋のケーブルの形は、目では区別がつかない。本稿が問うのは、では何が二つを分けているのかである──答は、形ではなく、荷重をどちらの長さで数えるかである。新しい数学定理も新しい法則も主張しない。 本稿の射程(射程注記):新しい数学定理も新しい法則も主張しない──懸垂線の方程式、放物線ケーブルの吊り橋理論、両者のテイラー展開はいずれも標準的である。構造設計をしない──使うのは二つの曲線と、一つの測度の違いだけである。たわみを扱わない──ケーブルの伸び、温度変化、活荷重によるたわみには立ち入らない。本稿は理想化された静止形状だけを見る。どちらが正しいかを判定しない──実際の橋は両者のあいだにあり、本稿はその位置を決めない。数値に工学上の精度を主張しない──たるみ比は形を比べるための値であって、特定の橋のものではない。最適形状を論じない──どちらが構造として有利かは扱わない。既刊との関係:論文198 は「まっすぐ」は「最短」を意味しないと示した──198 は同じ曲線に二つの性格があると言い、本稿は二つの曲線が同じに見えると言う。向きが逆である。論文326 は二乗三乗が恒等式だと示した──どちらも「見かけが同じでも中身が違う」型だが、326 は比の話、本稿は関数の話である。論文300 は同根か別根かは判定できると示した──本稿は「数値が近いだけでは同根でない」の実例になっている。論文245 は二つの極限が順序を書くまで答を持たないと示した──ここでも、どの測度で積分するかを書くまで曲線が決まらない。論文183 は同じ表が二つの正しい答を持つと示した──同じ形が二つの正しい方程式を持つという、対になる例である。加えたのは両曲線の差を六点で計算したこと、差が四次から出ることを明示したこと、たるみ比ごとの最大差を出したこと、分離子を測度の違いに置いたことである。 第一に、原点では二次まで完全に一致する。 x=0.1 で差は 0.083292 パーセント(第2節)。 第二に、遠くでは別物になる。 x=5 で 82.925818 パーセント(第2節)。 第三に、差は四次の項から出る。 x^4/24 対 0(第2節)。 第四に、これが本稿の芯である。分けているのは、弧長で数えるか水平長で数えるかである(第3節)。 第五に、実際の吊り橋では差は 0.082674 パーセントしかない(第4節)。 第六に、分離子は「桁とケーブル、どちらが重いか」である(第5節)。 垂れた鎖と吊り橋のケーブルは、目では区別がつかない。 cosh を展開すると最初の項が放物線そのもので、二次まで完全に一致する──a=1 とすると x=0.1 での差は 0.083292 パーセント、x=5 では 82.925818 パーセントであり、同じ二本について両方とも正しい。「どれだけ違うか」は、どこを見るかを言わないと決まらない。差は四次の項から出る──一次も二次も三次も一致しているので近くでは見分けようがなく、放物線は懸垂線の二次のテイラー多項式そのものであって、当てはめた曲線ではない。そして二つを分けているのは、形ではなく測度である──どちらも H y''=w という同じ釣り合いから出て、違うのは w が弧長 ds について一定か、水平長 dx について一定かだけである。弧長で一定なら傾いた所ほど水平方向に重くなり、だから懸垂線のほうが速く立ち上がる。だから「懸垂線に近い放物線」という言い方は順序が逆である──どちらも別々の物理から出た正確な答であって、近似関係は後から数学が見つけたものである。実際の吊り橋のたるみ比 1/10 では、最大差はたるみの 0.082674 パーセント──スパン 1000 m・たるみ 100 m の橋で 8.3 cm であり、施工の誤差より小さいこともある。それでも二つは別の曲線である。分けるものは一つ──吊られた桁とケーブル自身の、どちらが重いか。実際の橋はそのあいだにあり、本稿はその位置を決めない。論文198 とは向きがちょうど逆である──198 は一つの曲線が二つの性格を持つと言い、本稿は二つの曲線が一つの見かけを持つと言う。どちらも「見た目は根拠にならない」という同じ規律に服している。最後に一つ──ガリレオは垂れた鎖を放物線だと書いた。間違いだが、0.083 パーセントの間違いである。目で見て直せる誤りではなく、cosh という関数がまだ無かったのだから直しようもなかった。正しい形を書くには、書くための言葉が先に要る。 作成にあたって:本稿の着想と内容は、著者自身の考察に基づくものです。文章の構成整理や英訳、数式の確認には AI(大規模言語モデル)の助力を得ました。最終的な内容の解釈や誤りがあれば、それらはすべて著者の責に帰します。お気づきの点があれば、ご教示いただければ幸いです。

Yuuki Yamagishi · 0 citations
#small language model Open access Aug 2026

The geopolitics of Artificial Intelligence: Technological Dependence and Digital Sovereignty in Developing States

Artificial intelligence is becoming part of the way countries develop their economies, educate their populations and provide digital services. At the same time, the resources needed to build advanced AI systems are concentrated in a relatively small number of countries and companies. This creates an important question for developing states: how can they benefit from technologies developed abroad without becoming too dependent on decisions made outside their own borders? This paper examines this question through the case of Uzbekistan. It focuses on technological dependence and digital sovereignty and uses linguistic inequality in AI as one example of how this dependence may appear in practice. Particular attention is given to tokenization, the process through which text is divided into smaller units before it is processed by a language model. Research has shown that different languages can require very different numbers of tokens to communicate similar information, and that these differences can affect the cost and efficiency of commercial AI systems (Ahia et al., 2023). The paper argues that the Uzbek language should not be viewed only as a technical challenge for AI developers. Its relatively limited computational language resources also illustrate a wider issue: countries with smaller technological ecosystems may have less influence over the systems they increasingly use. The paper therefore considers digital sovereignty not as complete technological independence, but as the ability of a country to develop sufficient knowledge, infrastructure and partnerships to make informed decisions about technologies on which it relies. It also examines the role of UNESCO, international cooperation and regional collaboration in creating a more inclusive system of AI governance.

Gulira'no Abdullayeva · 0 citations
#small language model Open access Aug 2026

DPS (Dynamic Prompt Specialization via Semantic Micro-Model Routing)

WHITE PAPER Dynamic Prompt Specialization via Semantic Micro-Model Routing STRATEGIC TECHNICAL MEMORANDUM: DPS Author: Valerii Khalif (VALEO), AVA LIVE Status: Draft for DOI / Zenodo submission Date: August 31, 2026 License: Open / free to use (attribution requested; voluntary donations welcome, not required) Abstract Modern development of large language models (LLMs) based on the Transformer architecture relies heavily on scaling: increasing parameter counts, training data volumes, and computational resources, alongside regular retraining and model updates. This document proposes an alternative, additive adaptation layer that does not require modifying the weights of the core LLM. The proposed approach is designated as Dynamic Prompt Specialization (DPS). DPS employs a small, high-speed auxiliary model—a semantic micro-model—that continuously analyzes the semantic and communicative structure of the current user input, extracts concise semantic fragments, and maps them to a predefined code table. Each code corresponds to a pre-engineered and validated prompt fragment, optionally accompanied by a confidence level or signal intensity score. Based on the active set of signals, DPS constructs a dynamic specialized context transmitted to the unmodified primary LLM. The primary goal of DPS is user interaction personalization: the system can adapt to stable individual communication patterns, ongoing conversational context, professional expertise level, task domain, feedback to previous answers, and other observable signals. Furthermore, DPS is not restricted to user communication. The same architecture can be applied to domain-specific pre-specialization of queries prior to reaching the primary LLM. This pre-structures multi-faceted queries and potentially reduces the necessity for dedicated specialist agents and complex orchestration pipelines. DPS operates without altering main model weights or requiring retraining. In the absence of a sufficiently confident signal, the system defaults to the baseline behavior of the main LLM. Consequently, DPS represents an orthogonal adaptation layer that can be integrated into existing LLM infrastructure without altering the underlying foundation model. 1. Problem Statement The proposed approach addresses two interconnected constraints in current LLM-based systems. 1.1. Scaling Does Not Solve Individual Specialization The main vector of foundation model advancement focuses on increasing: Parameter counts; Training data volume; Compute resources; Training duration and complexity; Model update frequency. However, expanding overall model capabilities does not yield a proportional enhancement in interaction quality for a specific individual user. In many practical scenarios, a user operates within a relatively bounded domain of knowledge and exhibits consistent individual patterns of task formulation, communication, verification, clarification, and correction. Thus, incremental general model knowledge often carries less practical utility than the system's ability to precisely align with a specific user's traits and the current interaction state. DPS shifts part of the specialization process from model parameter modification to dynamic context adaptation. 1.2. Static System Prompt Most LLM systems rely on a single primary system prompt or a narrow set of predefined instructions. As a result, interactions are typically treated as static: $$\text{System Instructions} \longrightarrow \text{User Request} \longrightarrow \text{Response}$$ DPS frames interaction as a dynamic process where conversational state evolves continuously: $$\text{Current User Signal} \longrightarrow \text{Analysis} \longrightarrow \text{Active Context Update} \longrightarrow \text{Response} \longrightarrow \text{New User Signal} \longrightarrow \text{State Update}$$ Thus, specialization occurs continuously rather than as a single static setup. 2. Core DPS Architecture DPS comprises three functionally separated layers, each holding distinct responsibilities. 2.1. Semantic Micro-Model – Fast Semantic Layer The first layer consists of a lightweight, ultra-fast auxiliary model. It is not designed for full reasoning or final output generation. Its scope is narrowly constrained to: Analyzing input utterances; Extracting short semantic fragments; Identifying corresponding semantic signals; Mapping signals to designated codes; Computing confidence or intensity levels for each signal. A typical semantic fragment may consist of a few words, though the architecture does not impose fixed-length constraints. Fragment size depends on input signal characteristics and implementation design. A single user turn can simultaneously trigger multiple independent codes. For instance, a query might concurrently express: Professional technical context; Financial domain context; Request for conciseness; Requirement for high precision; Emotional tone/signal; Affirmation of a previously established rule. Consequently, DPS is not forced to select a single category or agent, but instead builds a multidimensional set of co-active signals. 2.2. Code-Prompt Correspondence Table The second layer operates as a mapping repository: Code (e.g., PFA-0281) $\rightarrow$ Pre-engineered Prompt Fragment. A key principle of DPS is that prompt fragments are not automatically generated by the system on the fly. They are authored by engineers and undergo a controlled lifecycle: $$\text{Design} \longrightarrow \text{Testing} \longrightarrow \text{Verification} \longrightarrow \text{Approval} \longrightarrow \text{Deployment}$$ This establishes prompt engineering as a structured software engineering discipline. Codes map to behavioral, domain, or contextual specializations, such as: Technical engineering; Software programming; Financial analysis; Scientific discussion; Conversational interaction; Concise messaging; In-depth explanation; Critical evaluation; Re-affirmation of established context. 2.3. Dynamic Composite Prompt At the third layer, active codes are compiled into a unified composite context. Each signal maintains an intensity score, denoted internally as a temperature signal—a metric reflecting the activation strength of the given signal. Note: This term is distinct from the generation temperature parameter used in LLM sampling. DPS temperature measures signal activation strength rather than generative randomness. Based on these active signals, the system constructs a Composite Prompt passed to the primary LLM. Crucially, the primary LLM: Undergoes no retraining; Retains its original weights; Requires no specialized model variant; Executes a standard single generation pass. 3. Dynamic Temperature Profile – Breadcrumb State A core feature of DPS is that user profiling is not a one-off classification. It exists as a dynamic interaction state termed the Breadcrumb State. Profiling begins during the first exchanges of a conversation, eliminating the need to aggregate multiple historical sessions before initial activation. As dialogue progresses, individual signals can strengthen ($\uparrow$), weaken ($\downarrow$), disappear, be superseded, or interact with other active signals. Signal Dimension Trend Professional context $\uparrow$ Conciseness $\uparrow$ Need for detailed explanation $\downarrow$ Resistance to suggestion $\uparrow$ Confirmation of understanding $\uparrow$ This enables tracking not only topic content, but the evolving communicative dynamics of the session. 4. User Personalization The main purpose of DPS is adapting system behavior to an individual user. Across recurring sessions, a profile can persist (stored on the user side as a table of code mappings and signal temperatures) and evolve. Over time, the system continuously adapts to: Consistent professional contexts; Preferred task formulation styles; Desired response depth; Typical modes of confirmation or disagreement; Receptivity to suggestions; Preferred level of directness; Recurring communication patterns. This does not alter core LLM weights; instead, the interaction context shifts, allowing a single foundation model to serve diverse users through distinct dynamic adaptation profiles. 5. Professional Specialization Personalization does not degrade technical rigor. DPS can serve as a domain-specialization pre-layer between user requests and the primary LLM. When a user simultaneously addresses a technical problem, financial constraints, demands rigorous analysis, cites prior rules, and signals skepticism, traditional agentic architectures route the query through multiple stages: $$\text{Router} \longrightarrow \text{Technical Agent} \longrightarrow \text{Financial Agent} \longrightarrow \text{Context Aggregator} \longrightarrow \text{Main LLM}$$ DPS streamlines this pathway: $$\text{User Input} \longrightarrow \text{Semantic Micro-Model} \longrightarrow \text{Active Codes Selection} \longrightarrow \text{Composite Prompt} \longrightarrow \text{Main LLM}$$ Multiple context facets are integrated into a single pre-generation context, avoiding multi-agent pipeline overhead. 6. Reducing Agentic Complexity Modern AI implementations frequently rely on multi-agent orchestration: $$\text{Router} \longrightarrow \text{Specialist Agents (A, B, C)} \longrightarrow \text{Aggregator} \longrightarrow \text{Final LLM}$$ While effective, this paradigm increases inference latency, model call counts, compute costs, context size, and orchestration complexity. DPS offloads specialization that can be expressed via engineered prompt fragments to a lightweight micro-model prior to

Valeo Khalif · 0 citations
#small language model Open access Aug 2026

Literary literacy and mental health in contemporary China: evidence from the China family panel studies

Background Mental health problems impose a substantial social and economic burden in China. Yet population-based mental-health research has paid limited attention to literary literacy as a cultural resource embedded in everyday life. This study examines literary literacy as a multidimensional construct associated with adult mental health. Methods We analysed 114,448 person-wave observations from six China Family Panel Studies waves (2010, 2014, 2016, 2018, 2020, and 2022). The reconstructed literary-literacy score combines two direct domains—reading participation and language-cognitive capacity—using equal weights after within-wave standardization. Respondent education, parental education and family cultural resources were removed from the score and entered as covariates. The analysis used generalized propensity-score inverse probability weighting, covariate adjustment, individual fixed effects, a generalized additive percentile profile, an observed-variable parallel-path structural equation model and product-term moderation models. Results In the weighted covariate-adjusted model, a one-standard-deviation increase in reconstructed literary literacy was associated with a 0.058-SD higher mental-health score (95% CI [0.046, 0.070]). The within-person estimate was smaller ( β = 0.009, 95% CI [−0.001, 0.019]). In component models, the language-cognitive estimate was positive ( β = 0.075, 95% CI [0.066, 0.084]), whereas the average reading-participation estimate was close to zero ( β = −0.003, 95% CI [−0.009, 0.004]). The parallel-path model identified indirect coefficients through language expression, future confidence, interpersonal trust and life satisfaction, all with 95% confidence intervals above zero. The literary-literacy gradient was larger among women, rural residents, lower-income and lower-education groups, and adults aged 60 or older. Conclusion Reconstructing literary literacy from direct reading and language-cognitive indicators produced a positive adjusted association with adult mental health. The results locate the strongest average contribution in language-cognitive capacity and show that the joint profile varies across social groups.

Hang Shi · 0 citations
#small language model Open access Aug 2026

생성형 LLM 활용 비전공자의 실무지식 기반 소프트웨어 구현·보정 과정에 관한 탐색적 사례연구 — EasyHR Office와 자기이해 웹 애플리케이션 개발 사례를 중심으로 —

The spread of Generative Artificial Intelligence and large language models (LLMs) has opened new possibilities for non-specialists to conceptualize and build functioning software on their own. This study exploratorily analyzes how a single user without formal programming training employed generative LLMs to develop, revise, and refine EasyHR Office and a self-understanding web application titled "나는 어떤 일을 어떻게 할 때 가장 잘 작동하는가?". The two programs originated from a combination of the desire to create something directly through generative AI and a practical concern with HR work in small workplaces as well as questions about how people function. In addition, recurring patterns observed across the two cases are tentatively organized under the label of Deliberate Deficit-Consolidation (DDC). Keywords: Generative Artificial Intelligence, Large Language Model, LLM, Non-specialist Development, Self-directed Learning, HR Practice, Software Development, Web Application, Case Study, DDC

Myung‐Jun Lee · 0 citations
#small language model Book Open access Aug 2026

三阶意识与 AF-by-discrete 的结构相容性

This bilingual Chinese–English research monograph develops a controlled structural correspondence between a psychological–phenomenological account of third-order consciousness and the mathematics of AF-by-discrete groupoids. On the psychological side, the work asks how experience becomes organized as “my experience,” how separated episodes become recognized as “this is happening again,” and how the observer that names, compares, evaluates, and narrates experience can itself become part of what is observed. On the mathematical side, it independently develops the required language from finite graphs, Cantor path spaces, groupoids and germs, AF groupoids, graphs with boundary, tile inflation, traverses, returns, and incompressibility. The two domains are connected only through explicitly delimited structural correspondences and a stated translation discipline. A central working hypothesis is that subject-organizing complexity need not be uniformly distributed across experience: a large background may be organized through a small number of repeatedly re-identifiable, future-relevant interfaces. These interfaces are not stipulated in advance; they must be forced into view by repeated observation, return structure, failures of local continuation, and stable changes in what becomes possible afterward. The manuscript develops finite-scale formal tools, counterexamples, failure conditions, and candidate routes back to psychological observation and experimental design. The work does not identify consciousness with a groupoid and does not claim to provide a validated neural mechanism, clinical model, or ontological theory. Its aim is to make cross-disciplinary claims more discriminating, more explicit about their evidential burden, and easier to reject when the proposed correspondence fails. This manuscript has not undergone formal peer review. Zenodo is used here as a permanent archival and citation venue for the complete bilingual research text. Access to the deposited files is currently restricted; access may be granted by the author upon request.

Zheng Kuang · 0 citations
#small language model Open access Aug 2026

Prompt Framing and the Coverage Gap in Automated Security Analysis of Code Generated by Small, Locally Run Language Models

Initial Research Release This is the first public release of Security in LLM-Generated Code. This release contains the research materials, experimental code, datasets, analysis scripts, findings, and supporting documentation associated with the study. Contents Research paper and supporting documentation Experimental datasets LLM-generated code security analysis Analysis and auditing scripts Research findings Reproducibility materials Version v1.0.0 This release represents the initial public version of the project and is intended to provide a stable, citable snapshot of the research.

Talha Imran · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.