Meta-Student is a collaborative research platform, coordinated by the Sports Science Replication Centre, that pools independently collected undergraduate research datasets into statistically powerful individual-participant-data (IPD) meta-analyses. Students at multiple institutions each run a small, standardised study under local ethical approval and supervision and upload anonymised participant-level data; the platform combines these contributions in a single pre-specified statistical model and grows the pooled sample sequentially until the evidence reaches a pre-specified stopping criteria. Every contributor whose data pass quality validation is a named co-author, and each completed study yields an openly-archived dataset and a reproducible completion report. We convert otherwise-unpublished, individually underpowered/imprecise student projects into cumulative, adequately powered/precise evidence. One avenue of study design is replication, in line with our vision at the Sports Science Replication Centre (www.ssreplicationcentre.com). Candidate studies for replication are selected systematically rather than by nomination, using a Replication Research Identification System (RRIS). RRIS scans applied sport-and-exercise-science research across 18 quartile-1 journals and ranks experimental and quasi-experimental studies by a modified Replication Value index — a time-adjusted form of Isager et al.'s (2023) framework that combines a study's uncertainty (small sample size, 1/N) with its impact, where impact is operationalised as relative citations (citations per year) rather than total citations, so that recently-published findings are not disadvantaged for having had less time to accrue citations. Studies with a high Replication Value (influential claims resting on small, uncertain samples) are the most efficient use of scarce replication resources and are prioritised as potential Meta-Student targets, subject to feasibility screening for the platform's supported designs, and then expertise. The study preregistered here was selected through this process. Running economy (RE), in this case the oxygen cost of running at a fixed sub-maximal speed, is one of the more consistent physiological correlates of distance-running performance (Saunders et al., 2004a; Barnes & Kilding, 2015). A body of work led by Schücker and colleagues reports that where a runner directs their attention changes their RE: an internal focus (attending to one’s own movement or breathing) produces a higher oxygen cost than an external focus (attending to the environment). Schücker, Hagemann, Strauss, and Völker (2009) is the original publication of this effect and was recently replicated in 2019 by the same authors, which is the study under focus here (Schücker & Parrington, 2019). In n = 12 runners performing three counterbalanced 6-minute treadmill conditions (movement / breathing / video) at an individualised sub-maximal speed, VO₂ was significantly higher in both internal-focus conditions than in the external (video) condition (movement – video: +2.6%, d = 0.73; breathing – video: +4.2%, d = 0.77). Our main comparison of interest is the movement vs. video contrast and since we adapt some of the original methods, this should be considered a conceptual replication. Critically, the most recent robust-Bayesian re-analysis (McKay et al., 2024) concludes that the apparent external-focus superiority is substantially inflated by reporting and publication bias. We feel that this claim needs a multi-site replication for this reason. The under-powered original (recruited n=12, tested n=11) yields wide confidence intervals and the reported effects (d ≈ 0.73–0.77) may be vulnerable to inflation (Button et al., 2013). The supporting evidence is dominated by one research group and, largely, one laboratory and language (German instructions). Independent, multi-site, multi-language replication is the decisive test of whether the effect generalises or is a lab/context artefact. Pooling participant-level data (IPD) across many student laboratories does two things a single replication cannot: (i) it delivers the statistical power to estimate the effect precisely and to test it against a smallest-effect-size-of-interest (ii) it provides a severe, pre-registered test of the underlying theory with a real risk of falsification. The meta-synthesis is supported by the original author, Linda Schücker.
Long-context "context rot" (accuracy degradation as input length grows, even when task difficulty is held fixed) is localized here to a specific, causal mechanism inside a small Mixture-of-Experts language model (OLMoE-1B-7B-0924), then tested for generality across two further architectures. Forced-choice retrieval accuracy falls from 0.938 (256 tokens) to 0.688 (3,840 tokens) on a controlled needle-in-a-haystack substrate. A deconfounded linear probe shows the target fact remains ~99.5% decodable at its source position on every failing prompt, ruling out storage loss; decodability specifically at the readout position degrades instead (0.714 vs. 0.960 on model-right prompts). Sixteen attention heads, identified from short prompts alone, carry that content to the readout; on long failing prompts their attention to the fact collapses (0.432 to 0.187), localized to those heads well beyond a random-head null (specificity p<0.0005). A pre-softmax attention boost restricted to exactly those 16 heads, over the fact's span, repairs all 14 failing prompts against strength-matched random-head and wrong-span controls (answer probability 0.238 to 0.986, dz=5.55). Three router-level and readout-level interventions were tried first and failed: two that verifiably restored MoE specialist routing did not move accuracy, and a residual-stream content injection at the readout was content-independent — showing router "starvation" is a downstream correlate of the transport failure, not its cause. The full causal chain of head identification, localized collapse, and causal repair replicates on a second, architecturally distinct MoE (Granite-3.0-3B-A800M: 5/5 failures repaired against matched controls) and on a dense transformer with no MoE component at all (Pythia-2.8B: 11/11 repaired, dz=1.33-1.44; the collapse measurement itself on this substrate is significant and specific but falls under the project's own effect-size floor, and is reported as suggestive rather than confirmatory). This indicates the attention-transport failure and its repair do not depend on Mixture-of-Experts routing. A training-free, IDF-weighted lexical detector, requiring zero forward passes of the model under study, locates the failing span and recovers 100.8% and 99.4% of the oracle repair on the two primary OLMoE substrates, 99.1% on Granite, and 51% on a harder Pythia substrate built specifically to weaken lexical anchoring. Its boundary was tested, not assumed. It is unaffected by paraphrase (100% hit, 100.8% of oracle) but fails completely on multi-hop composition as a single pass (0% hit); a training-free two-stage chain recovers 45.0% of the oracle effect there with no labels, and a registered labeled fallback recovers 61.4%. A follow-up track replaces lexical overlap with sentence embeddings and a graph walk specifically to test coreference, where the relevant content shares zero tokens with the question by construction. The result is a registered split, not a single verdict: a discourse-adjacency edge fully solves coreference where the referring expression is immediately adjacent to its antecedent (100% of oracle, 22 of 22 repaired), but on a harder substrate with antecedent distance drawn independently per prompt, both that mechanism and a distance-tolerant version built specifically to extend it fail identically past distance one (dz=0.48). Both are reported as registered negative results, not discarded. Every claim above is preregistered before evaluation, with effect-size floors (|d| or |dz| >= 0.8) alongside significance, sign-flip and label-shuffle permutation tests (2,000 draws), matched controls, and Benjamini-Hochberg FDR correction. Results that miss a registered bar are reported as such rather than dropped — including three failed intervention families, an initially wrong lexical-detector prediction (logged and corrected), and both coreference negative results above.
Manjunath Bhaskar· Zenodo (CERN European Organi...· 0 citations
A diary of one hundred entries written by a language model carries recurring themes, found after the fact as sets of sparse-autoencoder features that fire together within a bounded stretch of entries. We ask what a witness could have seen at each point of the diary's life, and build the instrument that answers: the same detector run on the log up to each entry τ, with nothing carried between prefixes. Three objects result. A theme becomes a chain of complexes linked across τ by shared features, with a cohesion, the fraction the link carries, and a gap, what it gains and loses. The base is the largest complex at τ: the flagged mass the log cannot yet tell apart in time. The ground is the set of features on in most entries, which the detector excludes by construction. On the first diary the base forms over the first forty entries, holds, and from entry sixty-five articulates into themes; a theme becomes witnessable only after the text has left it, so every theme has a target-time and a later witness-time. Four diaries written under the same protocol share the shape of this process and share a ground of 6,182 features, while their bases share almost nothing. One feature, firing on the coupling of machine and human described as living tissue, has a career in all four in four different figures. The chains, the base and the ground supply a decidable Semantic Witness Log in the sense of dynamic open homotopy type theory: each link is a witness record with target-time and witness-time, cohesion is Presence, the bounded gap is Generativity, the unlinked flash is scatter, the ended chain is rupture, the re-linked chain is resurrection. We state a small calculus for the instrument in that form and close with the design, not yet run, for the same instrument over parallel lives sampled from one prefix. Full HTML text: https://icra.tanazur.org/papers/cohesion-of-an-evolving-text/ — ICRA pre-print series: https://icra.tanazur.org/
Iman Poernomo, Nahla· Zenodo (CERN European Organi...· 0 citations
Recurrent architectures that compress history into a single state are cheap but have classically been unable to exactly retrieve far past key–value bindings, whereas attention content-addresses the whole past at quadratic cost. We show that the conflict between dense language modeling and long-range associative recall need not be a binding trade-off. We couple a nonlinear aggregated-state backbone (AGG, an input-gated forget–write recurrence computed by a parallel scan) with a content-addressed readout (SEL, a small polynomial-composition kernel that scores each query against a bounded window of past states and retrieves their values). At a fixed compute budget on WikiText (GPT-2 vocabulary, V=50257), the combined architecture is the first recurrent model we have tested to simultaneously exceed an equal-size transformer on both axes: it cuts dense perplexity from 183.6 to 127.8 (mean over 3 seeds), and on a real-vocabulary associative-recall probe it reaches 1.000 far-distance (d=7) top-1 recall across three seeds — a level no linear state-space model we compared reached — and this near-perfect recall is retained as the number of key–value pairs scales to m=32 (d=31 recall 1.000) and as the probe is moved from synthetic memorization to real-text long-range retrieval, where the model recovers a hidden bond at 93–100% for contexts up to 8K tokens. Two stability mechanisms — a full-sequence window and a floor on the readout gate — turn an occasional boundary collapse into reliable convergence. The content-addressed readout adds only a narrow, windowed kernel with a handful of parameters, and a controlled removal isolates its role. We argue the resulting architecture is a distinct sequence-modeling primitive: it retains the exact retrieval, compact state, and linear-time cost of its two parents.
Ziheng Zhou· Zenodo (CERN European Organi...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A diary of one hundred entries written by a language model carries recurring themes, found after the fact as sets of sparse-autoencoder features that fire together within a bounded stretch of entries. We ask what a witness could have seen at each point of the diary's life, and build the instrument that answers: the same detector run on the log up to each entry τ, with nothing carried between prefixes. Three objects result. A theme becomes a chain of complexes linked across τ by shared features, with a cohesion, the fraction the link carries, and a gap, what it gains and loses. The base is the largest complex at τ: the flagged mass the log cannot yet tell apart in time. The ground is the set of features on in most entries, which the detector excludes by construction. On the first diary the base forms over the first forty entries, holds, and from entry sixty-five articulates into themes; a theme becomes witnessable only after the text has left it, so every theme has a target-time and a later witness-time. Four diaries written under the same protocol share the shape of this process and share a ground of 6,182 features, while their bases share almost nothing. One feature, firing on the coupling of machine and human described as living tissue, has a career in all four in four different figures. The chains, the base and the ground supply a decidable Semantic Witness Log in the sense of dynamic open homotopy type theory: each link is a witness record with target-time and witness-time, cohesion is Presence, the bounded gap is Generativity, the unlinked flash is scatter, the ended chain is rupture, the re-linked chain is resurrection. We state a small calculus for the instrument in that form and close with the design, not yet run, for the same instrument over parallel lives sampled from one prefix. Full HTML text: https://icra.tanazur.org/papers/cohesion-of-an-evolving-text/ — ICRA pre-print series: https://icra.tanazur.org/
Iman Poernomo, Nahla· Zenodo (CERN European Organi...· 0 citations
Preprint. Not yet peer-reviewed. Frozen autoregressive language models cluster surface-similar tokenstogether even when a stronger, task-relevant structure is available inthe input. This paper documents "morphological hijacking" — near-totalcollapse of algebraic-structure recovery under adversarial surfacecorrelation — across four frozen model families (GPT2-small, Pythia-410m,Mistral-7B-v0.3, Qwen2.5-7B), traces it to representational anisotropy,and introduces a lightweight, trainable projection head (contrastiveobjective + positional-symmetry penalty + rogue-dimension ablation +Soft-PCA initialization) that substantially recovers the structure. Theresult is validated at both the discrete-clustering level (Adjusted RandIndex) and the continuous embedding-geometry level (margin, win rate)across five architecturally diverse models, including two models trainedexplicitly for embedding/retrieval tasks, and further tested on naturalEnglish words and a freshly generated, 10x-larger constructed lexicon.Several negative results are reported alongside the positive ones,including a rejected "globally architectural anisotropy" hypothesis anda rejected layer-selection heuristic. This is a preliminary, honestly-scoped empirical report, not a claim ofgeneral applicability. Limitations, a full pre-submission checklist, andcomplete reproducibility code are included. Author: Reza NirouyarORCID: 0009-0000-4690-6842Contact: contact@varzin.orgProject website: https://varzin.org
Nirouyar Reza· Zenodo (CERN European Organi...· 0 citations
End-to-end vision-language models (VLMs) bind visual competence to the scale of their language model: as the language model shrinks, perception and reasoning degrade together. We study an embodied scene question-answering (QA) interface in which vision never enters the language model. A frozen perception stack detects and ranges objects; a deterministic semantic serializer compiles the perceived state, errors included, into decision-aligned text; an unmodified text-only large language model (LLM) answers. On a visible-scope-matched, occlusion-audited campus-robot benchmark, under a prospectively frozen criterion, the serialized interface, using detectors fine-tuned in-domain within each fold, outperforms a zero-shot VLM whose language model has the same 7B scale (0.7892 vs 0.7462), with a larger margin at 3B (0.7673 vs 0.6913). Preregistered decoupling experiments show the gain survives paraphrase, attributing it to decision-aligned computation rather than answer-string leakage, while novel judgment vocabularies bound its scope. The advantage grows as the reader shrinks to 1.5B and reverses at 0.5B, and a ground-truth oracle locates the reader-capability floor. Under matched task supervision the interfaces converge: a VLM fine-tuned with low-rank adaptation (LoRA) overtakes the zero-shot system but only ties an equally supervised text reader (0.8441 vs 0.8396, no statistically resolved difference), and both routes remain perception-bound. Reported perception parameters are comparable to those of the VLM's vision tower, and total compute is not smaller. This preprint has been submitted to the IEEE for possible publication.
Cong Xu, Ravi Sankar· Zenodo (CERN European Organi...· 0 citations
Self_agent is a locally-run, tick-based autonomous process built on small open-weight language models (via Groq) that maintains its own goal queue, reflects on its own behavior, executes work through short-lived stateless task nodes, and — under a sandboxed, test-gated pipeline — proposes and applies changes to its own source code. It is a working instance of persistent computation (Brooks, 2026): execution as continuous recursive state evolution rather than discrete request/response, with a SQLite-backed kernel that survives restarts and resumes from its last recorded tick. This document has three parts, written for three different readers. Part I is a technical specification: the tick loop, the goal lifecycle, the capability system, the self-modification pipeline, the data model, and the engineering decisions behind each. Part II is a whitepaper positioning self_agent against related architectures — including The Cognitive Runtime (Chainborn Labs & Brooks, 2026), a sibling system that occupies a deliberately different point in the same design space — and presents a domain-invariant template for porting the same core loop elsewhere. Part III addresses what this is actually good for today, stated plainly: a single-operator research prototype with real, specific engineering contributions and real, specific limitations, not a finished product and not a claim about machine consciousness.
Ashad Brooks· Zenodo (CERN European Organi...· 0 citations
End-to-end vision-language models (VLMs) bind visual competence to the scale of their language model: as the language model shrinks, perception and reasoning degrade together. We study an embodied scene question-answering (QA) interface in which vision never enters the language model. A frozen perception stack detects and ranges objects; a deterministic semantic serializer compiles the perceived state, errors included, into decision-aligned text; an unmodified text-only large language model (LLM) answers. On a visible-scope-matched, occlusion-audited campus-robot benchmark, under a prospectively frozen criterion, the serialized interface, using detectors fine-tuned in-domain within each fold, outperforms a zero-shot VLM whose language model has the same 7B scale (0.7892 vs 0.7462), with a larger margin at 3B (0.7673 vs 0.6913). Preregistered decoupling experiments show the gain survives paraphrase, attributing it to decision-aligned computation rather than answer-string leakage, while novel judgment vocabularies bound its scope. The advantage grows as the reader shrinks to 1.5B and reverses at 0.5B, and a ground-truth oracle locates the reader-capability floor. Under matched task supervision the interfaces converge: a VLM fine-tuned with low-rank adaptation (LoRA) overtakes the zero-shot system but only ties an equally supervised text reader (0.8441 vs 0.8396, no statistically resolved difference), and both routes remain perception-bound. Reported perception parameters are comparable to those of the VLM's vision tower, and total compute is not smaller. This preprint has been submitted to the IEEE for possible publication.
Cong Xu, Ravi Sankar· Zenodo (CERN European Organi...· 0 citations
This record contains the human annotation study underpinning the MSc thesis"Exploring Semantic Consistency in Image Super-Resolution" (Master in ComputerVision, Universitat Autònoma de Barcelona / Computer Vision Center). Super-resolution is normally evaluated on fidelity and perceptual quality — howclose an output is to its reference, and how natural it looks. Neither questionasks whether the reconstruction still means the same thing. A generative modelcan invent a plausible face, produce readable text where the reference had onlya smudge, or turn one object into another, and score well on both axes whiledoing it. This dataset was collected to study that failure mode directly, byasking people what actually changed. It contains 104 super-resolution image pairs and 4,700 annotation records from41 participants, and it is self-contained: every statistic reported in thethesis can be recomputed from these files alone. THE TASK Participants saw a reference image and a super-resolved reconstruction side byside, and marked every way the reconstruction altered the meaning of the scene: C1 — details ("details"): fine details altered, invented or lost C2 — local structures ("local"): local structures deformed or restructured C3 — holistic semantics ("holistic"): the overall meaning of the scene changes C4 — no change ("none"): nothing meaningful changed C1–C3 may co-occur. C4 is exclusive. Annotators honoured that constraint withouta single exception across all 3,497 judgements. Because one judgement can carryseveral categories, a row is one selected label, not one judgement: a judgementis the set of rows sharing a (user_id, image_id) pair. THE IMAGES 100 real pairs plus 4 attention controls. Each control presents a byte-identicalpair — the same file shown twice — so the only defensible answer is C4. This isverifiable rather than asserted: exactly the four rows flagged is_control = truehave reference_identical_to_reconstruction = true, and no real pair does. The 100 real pairs are balanced 52/52 between two public corpora, LSDIR andFoundIR, and spread across 15 super-resolution models at six or seven pairseach: BSRGAN, DiffBIR, DiT4SR, DiT-SR, FaithDiff, HAT, OSEDiff, PiSA-SR, RAP-SR,RealESRGAN, SeeSR, SUPIR, SwinIR, TSD-SR and TVT. FoundIR pairs span fourdegradation regimes. PARTICIPATION AND QUALITY CONTROL 41 participants started, 29 completed all 104 pairs, and 24 of those passed allfour controls. Controls were graded silently: participants received no feedbackand failing one did not end the session. Nobody who finished scored below two offour, and a missed control was almost always answered "details" — the signatureof someone reading compression noise as real change rather than clicking atrandom. The 24 who passed everything form the clean pool used for the headline results:24 × 100 = 2,400 judgements, exactly 24 per real pair. Every one of thoseparticipants answered every pair, so no result can be an artefact of unevencoverage. CONTENTS annotations.csv 4,700 annotation records manifest_104.csv the 104 pairs with provenance, model, degradation and control flags images/reference/ 104 reference images images/reconstruction/ 104 reconstructions; filename equals image_id README.md column reference, agreement formula, runnable reproduction snippet SHA256SUMS checksums for every file PRIVACY Participants are pseudonymous. IP addresses were removed entirely rather thanhashed, because the IPv4 space is small enough to invert any hash by exhaustivesearch. Timestamps are reduced to dates, with a seconds_from_user_start columnpreserving response-time information without the absolute times that wouldpermit linkage. Country and language values held by fewer than threeparticipants are generalised, so no single row can identify a participant. PROVENANCE AND LICENSING Reference images derive from the public LSDIR and FoundIR datasets and remainunder their original licences; manifest_104.csv records the exact source ofevery file. The reconstructions were generated by the author using the publiccheckpoints of the 15 models listed above. Annotations, manifest anddocumentation are released under CC BY 4.0.
Valentin Micu-Hontan· Zenodo (CERN European Organi...· 0 citations
Here we present four companion documents for the program corpus listed below. The Taxonomy of Claims lists the theory's claims and assigns each one a status. The claims are grouped in four tiers: Framework (the laws and constants, which the theory does not derive), Derived from Framework (quantities computed from the framework, not independent inputs – including the per-edge failure probability p = ln 2/(2π²) and its spin-foam status, the bare-edge rule, and the tensor consistency relation n_t = n_s − 1), Pre-history (derived properties that are the same in any realization), and History (properties specific to our universe). A set of notes covers the counting conventions, the ensemble rule, the correspondence ledger, the anchor question and its resolution, several associations recorded without claim status, and a closing statement of the one assumption the theory requires. The Epistemic Inventory (What the Theory Claims to Know) contains the same material organized by question. Some fifty questions – why the cosmological constant is small, why w = −1, where the spectral tilt comes from, whether there are primordial gravitational waves, what preceded the first moment, why these constants – are each given the theory's answer in plain language and a status: derived, identified, dissolved, partial, or not claimed. The final section lists the questions the theory does not claim to answer. The Universe from a Single Strike is a lay description of the Ignition, provided as an onramp to the framework's technical papers. Two revision notes carry the tensor sector's audit trail, retained side by side. Correcting Errors in the Tensor-to-Scalar Ratio Calculation is the first revision, kept as the record it is: the ensemble correction, the bare-edge rule, and the revision of the physical ratio to 0.0301 ± 0.001. Deriving the Assertions in the Tensor-to-Scalar Ratio Calculation (new in this version, and controlling where the two disagree) is the consolidated revision, closing by computation the two commitments the first note left standing in its own words: the deposit-support rule, previously ratified, is derived as exact geometry (the 60° ring partition a theorem of the subdivision, the wedge-share inheritance rule derived), and the inter-move correlations, previously asserted shapeless, are enumerated (own-shell cross terms vanish identically; the boundary–existence channel is a constant renormalization adding no shape). The amplitude moves and no relation does: a standard flat-template B-mode analysis will report r = 0.0297 ± 0.0005, about 7% under the current combined limit r < 0.032; the underlying physical ratio, the same at every scale, is 0.0281 ± 0.0005, with the revision history owned on the record (0.0339 → 0.0301 → 0.0281). The tensor tilt is unchanged at the exact consistency relation n_t = n_s − 1 = −0.03512, now ten times redder than single-field inflation's at equal r; the α_s withdrawal stands. The new note answers the first note's frozen claims row by row in a ledger appendix, withdraws that note's promise of numerical finality as a category error, carries the complete derivations in its appendices, reproducible from the public replication package, and is registered before the next B-mode data release. The taxonomy and inventory are updated to match, including the upgraded statuses the note drives. This version registers three additions driven by the tensor tranche of the modeling package (10.5281/zenodo.22217227) and by the one-loop sector: the two-entry sourcing dictionary behind r – missing volume sources the scalar per cell, hinge-deficit shear sources the tensor per hinge, joined at the single-failure channel as a named premise – with the exact S³ mode-count law (2/5)(1 − 4/n²); and the marginal-normalization premise behind A_s, superseding the one-loop paper's "standard sectors net to unity" (the tensor-sector determinant is now computed, 0.762, and the vector and ghost sectors cancel exactly). Registered values are unchanged; the exact-support window factor 0.9539 is carried in the modeling package pending consolidation. The falsification conditions for the corpus are collected separately in the predictions letter (revised in step with this version) and are not repeated in these documents. All documents will be maintained: new versions will record changes in the status of claims, including withdrawals. The Program Corpus The Last Evaporation: Planck Remnants as Cosmological Seeds in Empty Spacetime - 10.5281/zenodo.19324262 Cosmological Structure Without Inflation: The Perturbation Spectrum from Pre-Geometric Construction - 10.5281/zenodo.19513896 One-Loop Identities on the S4 Instanton - 10.5281/zenodo.20045607 The Ignition Transition: From the No-Boundary Saddle to Radiation Domination - 10.5281/zenodo.20559451 The Ignition Inventory: Defect Energy and Static Topology - 10.5281/zenodo.20559916 Cosmology from a Three-Bit Seed: The Predictions and Their Falsification Gates - 10.5281/zenodo.21270529 Epistemic Inventory and Taxonomy of Claims - 10.5281/zenodo.21324530 Tractable Modeling: The Truncated Tessellation - 10.5281/zenodo.22217227
Scott Weller· Zenodo (CERN European Organi...· 0 citations
Preprint notice: This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. Deploying a language model on an embedded accelerator can fix its architecture, vocabulary and weight footprint before accuracy is considered. Holding the teacher, corpus, token budget and evaluation protocol fixed, we compare two initialisations of a 369M-parameter student distilled from a 4.2B-parameter teacher—(A) extraction, width-pruning the teacher's 752M sibling into the hardware-dictated shape, and (C) random initialisation of the same architecture—alongside an off-the-shelf 362M reference (B). Extraction wins at every measured budget: arm (C) never reaches, within 1.3B tokens, the perplexity that arm (A) attains after 100M tokens—a measured token-efficiency bound above 13×—and a saturating power-law fit over the measured range places the perplexity-ratio floor at 1.428. The ordering survives a corrected-objective replication and learning-rate sensitivity checks. The off-the-shelf reference retains higher general-benchmark scores but smaller adaptation gains on a ground-truth-verified corpus of structured scene questions: +0.486 vs. +0.142 accuracy under a matched protocol (at matched general-capability cost), and +0.486 vs. +0.339 at each arm's best rate, where the reference pays five times the general cost and most of its peak is answer-prior memorisation. A scene ablation credits the extracted student with +0.183 of scene-dependent accuracy, exceeding the teacher's +0.114 and the reference's best +0.125. Weight-only 4-bit quantisation packs the specialist's weights into 327 MiB—24× below the bf16 teacher—with domain accuracy unchanged within noise. Two negative results are reported: general-benchmark superiority was not achieved at this budget, and general-purpose distillation transferred the capacity to acquire domain capability rather than the capability itself. Reproducibility package (configurations, question banks, per-item logs, analysis code): 10.5281/zenodo.22256719.
Cong Xu, Ravi Sankar· Zenodo (CERN European Organi...· 0 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.