Skip to content

Category

small language model

608 papers

#small language model Open access Sep 2026

Morphological Hijacking in Frozen Language Models: Recovering Structured Representations and Testing the Limits of Algebraic Composition

This revised manuscript presents a controlled study of morphological hijacking in frozen language-model representations using the synthetic LUXVAR/VARZIN framework. The study shows that adversarial surface structure can dominate frozen representations, while a lightweight trained projection head can substantially recover the targeted categorical/group-position structure under controlled held-out, cross-script, and out-of-distribution tests. The revision incorporates the complete Level-3 composition diagnostic chain. A linear decoder failed on seen-pair composition, after which diagnostics tested optimization, structural rank deficiency, cross-family geometric alignment, and decoder capacity. A pre-registered one-hidden-layer MLP (64 hidden units, ReLU) fit the 120 seen ordered pairs almost perfectly (mean L3-A accuracy = 0.996 across five seeds), showing that decoder capacity was sufficient to fit the training set. Crucially, the same MLP did not generalize to 24 held-out ordered pairs: TRUE accuracy was 0.106 (38/360), compared with 0.281 (101/360) for the shuffled-label control and 0.175 (63/360) for the wrong-operation control. The revised interpretation is deliberately narrow. The intervention provides evidence for recovery of the targeted categorical/group-position structure, but it does not demonstrate systematic algebraic composition. The unseen-pair result is consistent with memorization/interpolation rather than learned application of the (i+j) mod 12 rule. The paper therefore treats composition as an important negative boundary condition, not as evidence that composition information is universally absent from the underlying representations. Conclusions are restricted to the tested synthetic lexicon, GPT-2-small extraction pipeline, preprocessing, alignment procedures, and decoder classes

Nirouyar Reza · 0 citations
#generative ai Open access Sep 2026

Subset Heterogeneity in Reward Model Benchmarks: A CONFIRM Validation of 174 RewardBench 2 Models

Background. RewardBench 2 aggregates pairwise preference accuracy across heterogeneous task families (factuality, instruction following, safety, and others). A single leaderboard score can mask systematic subset specialization, yet standard benchmark reporting rarely tests whether accuracy is independent of task category. Methods. We applied CONFIRM, a chi-square test of independence with Cramér's V effect sizing and empirically anchored letter grades, to 174 publicly released reward models. For each model we constructed a 5 × 2 contingency table (subset × correct/incorrect) from published per-prompt scores (n = 1,763 prompts after excluding the non-binary Ties subset). The null hypothesis was that correct/incorrect outcomes are independent of subset category. Grades A–F reflect V magnitude (CONFIRM v2 thresholds); grade I denotes insufficient power. Non-significant results grade F only when power ≥ 0.80 to detect V = 0.10. Results. All 174 models rejected independence at α = 0.05 (all p < 0.02; median V = 0.272, range 0.082–0.467). No model received F or I. Grade distribution: A 36.8% (n = 64), B 50.6% (n = 88), C 11.5% (n = 20), D 1.1% (n = 2). Pooling across models, mean per-subset accuracy was lowest on Precise IF (36.3 per 100) and highest on Safety (75.7 per 100). In 93.1% of models the largest subset gap involved Precise IF as the weakest category; Safety was the strongest endpoint in 70.7% of cases. Conclusions. Published RewardBench 2 reward models exhibit statistically detectable subset heterogeneity at this sample size. Aggregate accuracy therefore understates structured performance imbalance, particularly weakness on precise instruction-following relative to safety-oriented subsets. CONFIRM provides a reproducible heterogeneity diagnostic to be read alongside ranking metrics. Plain-language summary. For all 174 reward models tested, the rate of correct judgments differed across the five task types by more than the prespecified statistical threshold. The direction of the difference was shared: 93.1% of models were weakest on precise instruction-following, and 70.7% were strongest on safety. Each model was tested on 1,763 prompts, enough sensitivity to detect even a small difference had one been present, so a model showing no difference would have been identifiable as such. None did. A single overall leaderboard score does not show this — two models with the same average can differ substantially underneath. This analysis measures whether a model's accuracy is uneven across task types, not which model is best overall, and it identifies a pattern without establishing its cause. Supplementary material. The deposited archive (rewardbench2_validation.zip) contains per-model results, subset breakdowns, contingency cell counts, validation flags, and step-by-step mathematical derivations for all 174 models. Competing interests. The author is affiliated with TraceSeis, Inc., which is developing CONFIRM as a commercial product. This constitutes a competing interest. All results are reproducible from the cited public data and the deposited analysis outputs. AI use disclosure. Generative AI tools were used during preparation of this work, in two distinct roles. For drafting and implementation: Anthropic Claude assisted with manuscript prose; the CONFIRM engine and analysis pipeline were implemented with AI coding tools (Cursor, Anthropic Claude) to the author's specification; and Google Gemini was consulted during writing and analysis runs. For review: Perplexity provided editorial review of a late draft, and xAI Grok was used as a general consistency check. This reflects the author's record of tool use and is not offered as an exhaustive log. Research design, statistical methodology, and interpretation are the author's. Because the analysis software was AI-implemented, every reported statistic was independently recomputed from observed cell counts and checked against pipeline output before reporting; per-model derivations are deposited as confirm_math.html and can be checked by hand. The author takes full responsibility for the contents of this record.

Alvaro Chaveste-Fernandez · 0 citations
#small language model Open access Sep 2026

Spanish version of the personal experience of aging scale: psychometric analyses from classical and network psychometrics approaches in a Chilean sample

Abstract Background How people experience aging is linked to health, well-being, and engagement in later life, but Spanish-language measures for capture self-perceptions of aging in Latin America remain limited. Objective To adapt the Personal Experience of Aging Scale (PEAS) into Spanish and evaluate its psychometric properties in Chilean adults aged 50 and older. Methods The instrument was translated into Spanish, and content validity was examined by expert judges. Two independent community samples from Chile (exploratory sample: N = 207; confirmatory sample: N = 361) completed the PEAS. Dimensionality was examined using factor-analytic and network-based approaches, followed by confirmatory testing. Reliability, convergent/discriminant validity, measurement invariance across age groups (<60 vs. ≥ 60), and construct validity were assessed using measures of ageist stereotypes and perceived age discrimination. Results Expert ratings supported content validity. Exploratory and confirmatory evidence supported a three-factor structure. A 12-item model demonstrated good fit (CFI = .993, TLI = .992, SRMR = .068, RMSEA = .079) with strong standardized loadings ranging from .721 to .989. Although the adapted version retains the original 12-item length, it differs from the original structure in the distribution of items across dimensions due to the addition of one item and the removal of another during the adaptation process. This distinction should be considered when comparing findings across studies. Measurement invariance up to thresholds was supported across age groups (< 60 vs. ≥ 60; ΔCFI ≤ .01). Reliability was adequate to excellent across factors (α = .771 to .883; ω = .776 to .888). Physical Decline and Social Loss correlated positively with ageist stereotypes and perceived age discrimination (ρ = .286 to .427), whereas Continuous Growth showed small negative correlations (ρ = -.129 to -.167). Conclusions The Spanish PEAS shows robust evidence for a three-dimensional structure, good reliability, and measurement invariance up to thresholds across age groups in Chilean adults. The adapted 12-item version provides a parsimonious option with strong item functioning.

Oscar Terán-Mendoza, Vicente Cancino · 0 citations
#small language model Dataset Open access Sep 2026

Preregistered Tests of Local-Noise Predictions for GHZ Coherence Under Matched Qubit Environments on IBM Kingston

DOI: 10.5281/zenodo.22283967Record set: EC-MECH campaign (EC-MECH-001 pilot + EC-MECH-002)Related to: EC-STORAGE-001, DOI 10.5281/zenodo.21299091 Research question: Can the coherence lifetime of an entangled multi-qubit state be predicted from the measured coherence losses of its individual qubits? Plain language summary When several qubits are entangled together, how fast does the entangled state lose its coherence compared with the qubits measured one at a time? The textbook expectation, if each qubit's noise is its own private business, is that the losses simply add up. This record tests that expectation directly on IBM hardware, and reports two experiments: one that returned ABSTAIN under its own preregistered rules rather than supporting a scientific conclusion, and a redesigned successor that did produce one. The first experiment failed for an instructive reason. Its quality gate asked whether the measured decay curves looked like clean exponentials, when the question that actually mattered was whether the decay rate had been pinned down precisely. Those are different things, and for shallow decays they come apart badly. Five measurements whose rates were known to within 4–6% were discarded because their curves were too flat for the goodness-of-fit statistic to work with. Under the frozen rules that cascaded into an ABSTAIN on every downstream question. The verdict stands unamended. Diagnosing that failure exposed a second and more serious problem: the single-qubit reference measurements had been performed in a different noise environment from the entangled measurements they were being compared against. The redesigned experiment fixed both problems and added a dedicated probe of the environment mismatch itself. The result: once the environments were matched, the discrepancies during plain idling became substantially smaller — especially for the 4- and 6-qubit states — but the preregistered precision was still insufficient to establish additivity. Under a standard error-suppression pulse sequence, additivity was rejected at all three sizes. And the dedicated probe confirmed the environment mismatch was real, not merely a theoretical worry. Total hardware cost: 241 seconds across both experiments. Background: an unsupported claim, entered into the record EC-STORAGE-001 (DOI 10.5281/zenodo.21299091) recorded a pre-registration miss — a bare-arm coherence witness crossing at 9.0 µs against a registered 10–30 µs band — and explained it by asserting that the payload resided in weight-4 stabilizer correlations "whose coherences decay at the sum of constituent rates." That explanation does not close numerically against data in the same deposit. The selected qubits had reported T₂ of 178–368 µs. A sum-of-rates model over four such qubits predicts a joint coherence time of roughly 45–60 µs; the measured value was 8.81 ± 0.30 µs. The stated model over-predicts by a factor of roughly 5–7. The claim was therefore a hypothesis written in the grammar of a derivation. It is entered in the adjudication ledger as UNSUPPORTED, and the EC-MECH campaign was constructed to either repair it or retract it. This deposit does not resolve it in EC-STORAGE-001's favour, and readers of that record should treat the mechanism sentence as withdrawn pending a direct test on the encrypted-cloning encoding itself. What we did Both experiments ran on ibm_kingston (156-qubit IBM Heron processor) on a connected 6-qubit chain, physical qubits [14, 15, 19, 35, 34, 33], selected by a frozen policy from the same-day calibration snapshot with no manual override. Neither experiment uses, requires, or reproduces the encrypted-cloning protocol. They test the underlying physics assumption in isolation, using GHZ states and idle delays only. No proprietary components are involved, and the deposited code is fully self-contained. EC-MECH-001 (pilot) — rate additivity Nested GHZ states on the first k qubits (k = 2, 4, 6) were idled for τ ∈ [0, 45] µs and their weight-k coherence read out via parity oscillation. In parallel, all six qubits were prepared in |+⟩ and idled simultaneously to obtain per-qubit in-situ dephasing rates. The frozen predicate compared the fitted GHZ decay rate Γ_k against the sum Σ Γᵢ of the measured single-qubit rates. 320 circuits, 1024 shots each, one job, 93 s QPU. EC-MECH-002 — coherence-function additivity The successor abandons fitted rates entirely. For any family, define c(τ) = C(τ) / C(0) χ(τ) = −ln c(τ) Under independent local phase noise, the GHZ phase is the sum of the local phases, so the coherence factorises exactly: χ_S(τ) = Σ_{i∈S} χ_i(τ) This identity assumes nothing about decay shape — exponential, Gaussian, stretched, and non-Markovian decays all satisfy it. The scientific object is the residual Δ_S(τ) = χ_S(τ) − Σ χᵢ(τ), adjudicated through the scale-free ratio r_S(τ) = Δ_S(τ) / Σ χᵢ(τ) as a preregistered equivalence test with margin |r| ≤ 0.15, using simultaneous 95% confidence intervals (Bonferroni, n = 8, z = 2.734). CI wholly inside the margin → ACCEPT; wholly outside → REJECT; overlapping the boundary → ABSTAIN. Two design repairs distinguish it from the pilot: No R² gate, no fitted rate, no assumed decay law. Matched noise environments. Every non-target chain qubit is pinned in |0⟩ — including the GHZ spectators at k < 6 — rather than left in |+⟩. This matters because a GHZ block is immune to intra-block ZZ coupling: |0…0⟩ and |1…1⟩ are both +1 eigenstates of Z_iZ_j, so the relative phase carrying the coherence is untouched. In the pilot, the single-qubit reference was exposed to neighbour-state-dependent dephasing consistent with this mechanism, while GHZ symmetry cancels intra-block static ZZ — biasing the prediction high. Single-qubit controls use a two-colour scheme on the chain (targets {14, 19, 34}, then {15, 35, 33}) so every target has all chain neighbours pinned. τ = 0 is oversampled at 4096 shots because it is the shared normaliser and its uncertainty enters every χ, inducing covariance that is carried explicitly in the analysis. Design parameters (frozen). Normalisation point τ = 0 at 4096 shots, never adjudicated. Eight informative τ points at 1024 shots each: 12, 18, 20, 22, 25, 28, 35, 45 µs The grid is pilot-informed and deliberately non-uniform: the low end starts at 12 µs because the D_min = 0.10 denominator gate would exclude earlier times once neighbour pinning reduces the local χ, and five of the eight points are clustered in the 18–28 µs window to resolve structure in Δ(τ) there. See the evidence ceiling for what this costs. Arms: bare (plain delay) and dd (symmetric XY4, one cycle). Four phase points per sweep, spanning one full period of the weight-k oscillation. Runtime dynamical decoupling and twirling disabled, so the arms are defined solely by the circuits. Diagnostic family diagPlus on the bare arm at τ ∈ {12, 25, 45} µs — all three exact members of the primary grid, so no interpolation is performed. Circuit execution order randomised under frozen seed 20260902. 376 circuits, 520,192 shots, one job, 148 s QPU against a preregistered estimate of ~139 s (6.1% error). Driver provenance. The EC-MECH-002 driver was hardened after the EC-MECH-001 pilot and before the EC-MECH-002 freeze. The hardening bound submission to the freeze manifest and to the binding layout report (removing hand-entered qubit chains), replaced diagnostic interpolation with exact grid indexing, added randomised circuit execution order under a frozen seed, and added a drift-robustness case to the offline validator. All of it predates the freeze; the deposited digest covers the hardened files, and no change was made after data was seen. Results EC-MECH-001 — ABSTAIN (frozen, unamended) All six weight predicates returned ABSTAIN. Cause: five single-qubit component fits failed the frozen R² ≥ 0.90 gate (bare q15 = 0.883, q33 = 0.873; DD q15 = 0.893, q19 = 0.871, q33 = 0.859) despite relative rate uncertainties of 3.7–5.5%. Because q15 participates from k = 2 onward, the failure cascaded into every weight. All 18 fits in the run had σ(Γ)/Γ ≤ 0.064. R² measures the fraction of variance in log C explained by the line; when a decay is shallow the true variance is small and ordinary scatter consumes a large share of it. R² was the wrong gate. The pilot did establish, as observation rather than verdict, that in-situ dephasing under simultaneous idling ran ~2.4–4.3× faster than the reported T₂ on every qubit (e.g. q35: 67.4 µs in situ against 292.9 µs reported). EC-MECH-002 — primary adjudication Arm k usable τ median r χ²/dof p Verdict bare 2 7 −0.217 2.40 0.0185 ABSTAIN bare 4 8 −0.014 0.62 0.7625 ABSTAIN bare 6 8 −0.087 1.16 0.3166 ABSTAIN dd 2 7 −0.396 5.46 <0.0001 REJECT dd 4 8 −0.213 5.36 <0.0001 REJECT dd 6 8 −0.206 6.90 <0.0001 REJECT Negative r means the GHZ state retains coherence better than the independently measured single-qubit coherences predict. The bare ABSTAINs are not a proof of independence. The equivalence predicate was built precisely so that "we could not reject zero" cannot be reported as "we demonstrated independence." At k = 4 and k = 6 the residual is small (−0.014, −0.087) and the confidence intervals straddle the ±0.15 boundary; the correct statement is that a positive equivalence claim was not supported at the preregistered precision. k = 2 warrants extra caution in both arms. It is the only weight that lost a τ point to the denominator gate, and because r = χ_S/D − 1 with the smallest denominator, it amplifies any bias in D more than the other weights. Its values (−0.217 bare, −0.396 DD) are the most extreme on the board and the least reliable. ZZ diagnostic (preregistered as diagnostic, excluded from the predicate) A dedicated family reproduced the pilot's all-|+⟩ environment to test whether intra-chain ZZ coupling really was contaminating the reference. The preregistered qualitative prediction was excess χ > 0

Amit Brahmbhatt · 0 citations
#small language model Open access Sep 2026

Posterior Ontology Memory: Bayesian Prediction Across Mutable Semantic Schemas

Posterior Ontology Memory (POM) models uncertainty in the vocabulary used by a symbolic memory, rather than assuming that its semantic schema is fixed. It treats mappings from surface predicates to latent semantic categories as uncertain and carries that uncertainty into prediction. The model combines a partition prior, collapsed Beta–Bernoulli rule models, noisy compiler judgments, and Bayesian model averaging. A temporal version allows schemas to split or merge and rule rates to reset, using exact enumeration for small reference cases and sequential Monte Carlo for approximate inference.In finite grounded-rule experiments, model averaging improves held-out prediction when surface predicates share latent behavior, but the advantage disappears when that assumption breaks down. The temporal study passed seven of eight predeclared gates: delayed split evidence met its criterion, while delayed merge evidence did not. In a separate study of 19 GitHub API migrations, a local language model retrieved 18 correct notices, compared with 7 for token overlap, although direct mutation classification was less reliable. The paper presents a bounded proof of concept for symbolic memory that retains uncertainty over changing semantic schemas.

Henrik W. Lautergold · 0 citations
#diffusion models Open access Sep 2026

The Past, Present and Possible Future of Thermal Remote Sensing

Thermal infrared remote sensing has reached a turning point. A status audit completed for this review identifies 55 operational satellite platform deployments on 18 August 2026. The resulting dataset contains 174 named thermal infrared instrument designs mapped to 306 historical, operational and planned satellite platform deployments. Across the orbital record, the finest reported nominal spatial sampling decreased from 55 km for the Medium Resolution Infrared Radiometer aboard TIROS-2 in 1960 to 3.5 m for the mid-wave infrared imager aboard HotSat-2 in 2026, an improvement of more than four orders of magnitude. Public continuity missions are now complemented by specialized instruments on the International Space Station, commercial small satellites, aircraft, stratospheric balloons, drones and terrestrial systems. This review links that platform history to the governing physics of emitted radiation, detector and cooling technologies, calibration, atmospheric effects, emissivity and spatial resolution. It also examines the transition from classical machine learning to convolutional, recurrent, transformer, diffusion, foundation and vision–language approaches. The selected examples indicate that adoption in thermal applications has been uneven rather than uniformly delayed relative to other areas of Earth observation. These methods support image interpretation and reconstruction as well as quantitative retrieval, for which radiometric calibration, physical consistency and independent validation remain necessary. As sensor availability expands, scientific comparability increasingly depends on harmonization, cross-sensor transfer, uncertainty characterization and validation in physical units. We recommend three priorities: (i) open, cross-calibrated thermal archives; (ii) models that preserve the distinct physical meanings of thermal variables; and (iii) validation across sensors, regions and seasons using physical units, independent observations and quantified uncertainty.

Homayoun Rezaie, Geoffrey J. Hay · 0 citations
#large language models Open access Sep 2026

Preparing Citizens for Emergency Calls with a Hybrid FSM-LLM Dialogue Agent

Emergency calls are time-critical, verbal-only interactions in which call takers must assess severity and make decisions based solely on the caller’s description. Effective communication is critical during these calls, especially since most callers are inexperienced and untrained due to the rarity of emergency situations. To address this challenge, we develop a task-oriented dialogue agent that simulates emergency call takers to prepare citizens for effective medical emergency communication. It uses a finite-state machine as its dialogue policy and large language models for natural language understanding. This agent conducts protocol-guided interactions and maps caller descriptions to structured symptoms. In a small exploratory study with human participants, our agent achieved higher observed classification accuracy and dialogue efficiency while maintaining high naturalness scores.

Sean Patrick Klein, Jane Jean Kiam · 0 citations
#large language models Open access Sep 2026

Evolving transition-state search with agentic large language models

Abstract Accurate identification of transition states (TSs) is fundamental to computational chemistry. Modern reaction-discovery efforts increasingly rely on curating and completing large reaction datasets, where even a small fraction of TS-search failures can leave key pathways unresolved and bias the resulting reaction network. Here, we introduce an adaptive evolutionary framework for TS search workflows that prioritizes completion of a target dataset over global algorithmic robustness. The framework iteratively focuses on reactions for which no validated TS has been found and dynamically modifies the workflow to search the remaining TSs. Rather than converging toward a single universally optimal TS search workflow, the framework generates a sequence of specialized workflow variants that collectively increase TS coverage across the dataset. Applied to the Transition1X benchmark (over 10,000 reactions), this adaptive evolution improves a state-of-the-art TS-search workflow to achieve a >80% success rate of TS-finding with rigorous validation via eigenvector analysis and reaction path endpoint confirmation. More broadly, these results suggest that adaptive, LLM-driven evolutionary workflow optimization provides a transferable strategy for improving validated TS coverage in large-scale, failure-prone scientific workflows.

Jan A Meissner, Philipp Kuboth, Jan Meisner · 0 citations

Enhancing Image-Text Alignment in Chest X-ray Datasets by Reducing External References via a Fine-Tuned Large Language Model Pipeline.

An automated, data quality-oriented preprocessing pipeline that rewrites radiology reports at the sentence level to reduce external references is developed and can serve as an effective preprocessing step for improving the suitability of chest X-ray report datasets for multimodal model development.

Yu-Ruei Chen, Chang-Fu Kuo, Ching-Heng Lin · 0 citations
#small language model Preprint Sep 2026

You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change

Vision-language models are increasingly used to measure urban change from repeated street-level imagery, but their longitudinal reliability is not well understood. We test how much a perception score can change when the street itself does not undergo substantial redevelopment. Using 4,648 consecutive-epoch image pairs from 435 Google Street View standpoints across five US cities, we find that re-photographing the same street changes a perception score by 0.80 points on average, equivalent to 66.5% of the difference between two different streets in the same city. Repeated model calls contribute almost no variation, while image re-encoding and prompt-order changes each account for about one fifth of the between-street difference. Six image statistics describing scattering, contrast, colour, exposure, sharpness and specularity explain almost none of the remaining epoch-to-epoch variation. A small systematic drift of about 0.1 points remains and increases with the interval between captures, consistent with minor physical changes not recorded by redevelopment labels. Controlled experiments further show that acquisition conditions can shift scores when camera and image properties are allowed to vary, and that the direction of these shifts depends on the model. In crowdsourced imagery, camera geometry alone causes a model to report physical change in 45% of identical-scene pairs; normalising both images to a common virtual camera reduces this rate to 7.5%. Despite poor reliability at the individual-location level, aggregation recovers a coherent redevelopment signal: changed streets are judged wealthier, better maintained, more enclosed and less green. These results show that vision-language measurement of urban change is reliable at the scale of hundreds of paired observations, but not at the scale of individual sample points.

Kaizhen Tan · 0 citations
#small language model Dataset Open access Sep 2026

User study result and dataset

This record contains the human annotation study underpinning the MSc thesis"Exploring Semantic Consistency in Image Super-Resolution" (Master in ComputerVision, Universitat Autònoma de Barcelona / Computer Vision Center). Super-resolution is normally evaluated on fidelity and perceptual quality — howclose an output is to its reference, and how natural it looks. Neither questionasks whether the reconstruction still means the same thing. A generative modelcan invent a plausible face, produce readable text where the reference had onlya smudge, or turn one object into another, and score well on both axes whiledoing it. This dataset was collected to study that failure mode directly, byasking people what actually changed. It contains 104 super-resolution image pairs and 4,700 annotation records from41 participants, and it is self-contained: every statistic reported in thethesis can be recomputed from these files alone. THE TASK Participants saw a reference image and a super-resolved reconstruction side byside, and marked every way the reconstruction altered the meaning of the scene: C1 — details ("details"): fine details altered, invented or lost C2 — local structures ("local"): local structures deformed or restructured C3 — holistic semantics ("holistic"): the overall meaning of the scene changes C4 — no change ("none"): nothing meaningful changed C1–C3 may co-occur. C4 is exclusive. Annotators honoured that constraint withouta single exception across all 3,497 judgements. Because one judgement can carryseveral categories, a row is one selected label, not one judgement: a judgementis the set of rows sharing a (user_id, image_id) pair. THE IMAGES 100 real pairs plus 4 attention controls. Each control presents a byte-identicalpair — the same file shown twice — so the only defensible answer is C4. This isverifiable rather than asserted: exactly the four rows flagged is_control = truehave reference_identical_to_reconstruction = true, and no real pair does. The 100 real pairs are balanced 52/52 between two public corpora, LSDIR andFoundIR, and spread across 15 super-resolution models at six or seven pairseach: BSRGAN, DiffBIR, DiT4SR, DiT-SR, FaithDiff, HAT, OSEDiff, PiSA-SR, RAP-SR,RealESRGAN, SeeSR, SUPIR, SwinIR, TSD-SR and TVT. FoundIR pairs span fourdegradation regimes. PARTICIPATION AND QUALITY CONTROL 41 participants started, 29 completed all 104 pairs, and 24 of those passed allfour controls. Controls were graded silently: participants received no feedbackand failing one did not end the session. Nobody who finished scored below two offour, and a missed control was almost always answered "details" — the signatureof someone reading compression noise as real change rather than clicking atrandom. The 24 who passed everything form the clean pool used for the headline results:24 × 100 = 2,400 judgements, exactly 24 per real pair. Every one of thoseparticipants answered every pair, so no result can be an artefact of unevencoverage. CONTENTS annotations.csv 4,700 annotation records manifest_104.csv the 104 pairs with provenance, model, degradation and control flags images/reference/ 104 reference images images/reconstruction/ 104 reconstructions; filename equals image_id README.md column reference, agreement formula, runnable reproduction snippet SHA256SUMS checksums for every file PRIVACY Participants are pseudonymous. IP addresses were removed entirely rather thanhashed, because the IPv4 space is small enough to invert any hash by exhaustivesearch. Timestamps are reduced to dates, with a seconds_from_user_start columnpreserving response-time information without the absolute times that wouldpermit linkage. Country and language values held by fewer than threeparticipants are generalised, so no single row can identify a participant. PROVENANCE AND LICENSING Reference images derive from the public LSDIR and FoundIR datasets and remainunder their original licences; manifest_104.csv records the exact source ofevery file. The reconstructions were generated by the author using the publiccheckpoints of the 15 models listed above. Annotations, manifest anddocumentation are released under CC BY 4.0.

Valentin Micu-Hontan · 0 citations
#small language model Open access Sep 2026

HERMES-OT: Hierarchical Embedded Reasoning Models for Predictive Cyber-Physical Defence in Industrial Control Systems

A small-language-model architecture for AI-driven detection, prediction and safe defence of operational technology. Operational technology (OT) and industrial control systems (ICS) increasingly connect information technology, industrial networks, programmable logic controllers (PLCs), SCADA, distributed control systems, robotics and physical processes. This convergence creates a cybersecurity environment in which compromise of a digital asset may propagate into physical consequences. The emergence of tool-using and autonomous large language model (LLM) agents introduces an additional dimension to this threat. Recent research has demonstrated that LLMs can generate attacks against PLC environments, and that autonomous agents can, under appropriate conditions, progress from PLC interaction toward sustained physical objectives. Existing AI cybersecurity approaches predominantly focus on alert classification, anomaly detection, vulnerability identification, malware analysis or natural-language security assistance. These capabilities do not fully address the central OT problem: determining how a cyber event propagates through industrial topology and ultimately affects a physical process. This paper proposes HERMES-OT, a cyber-physical defence architecture based on a hierarchy of compact, specialised language models rather than a single general-purpose LLM, combining industrial telemetry, asset topology, vulnerability intelligence, attack graphs, process-state information, engineering knowledge and digital-twin simulation. The architecture introduces a reasoning chain that runs observe, understand, correlate, predict, simulate, prescribe, validate, learn. Probabilistic reasoning is surrounded by deterministic constraints: a policy engine provides authority, a safety layer provides boundaries, and the human retains control where risk demands it. The central hypothesis is that specialised small language models, coordinated through structured graphs and deterministic safety mechanisms, can provide sufficiently reliable industrial security reasoning while reducing inference latency, computational requirements and exposure of sensitive industrial information. The paper also proposes the OT-HERMES benchmark, a cyber-physical evaluation framework measuring detection, asset reasoning, vulnerability correlation, attack-path prediction, physical-impact prediction, defensive prescription, safety and computational efficiency. The research question is: what is the smallest AI model, or combination of small models, that can reliably reason about cyber-physical risk in an industrial environment? Status: this is a research proposal. No experimental performance figures are claimed for HERMES-OT. The architecture and the eight contributions are proposed and the six hypotheses stated, but not experimentally validated. Implementation and controlled experiments are the next stage of the work, and are essential before the system is presented as empirically validated.

Yasir Musawar · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.