Non-extensive statistics of the glioblastoma transcriptome: analysis code, derived results and validation-cohort inputs
Analysis code, derived result tables, figures and validation-cohort inputs for a study of non-extensive (Tsallis) statistics in the glioblastoma transcriptome. Expression deviations of each sample are fitted with q-Gaussian laws by maximum likelihood. In a longitudinal GLASS cohort (230 libraries from 115 IDH-wildtype patients, at diagnosis and at recurrence after temozolomide chemoradiotherapy) the entropic index has a cohort mean of 1.330 with a standard error of 0.007, and reproduces in two further independent cohorts analysed identically: 1.331 in CGGA-693 (n = 109) and 1.372 in TCGA-GBM (391 libraries from 293 participants). What is new in version 4. Estimator validated against data with a known answer. On 350 synthetic samples drawn from q-Gaussians with known parameters at the observed gene counts, the estimator recovers q with a bias of -0.0008, no fit fails to converge, and the nominal 95% profile-likelihood intervals achieve 94.0% empirical coverage. Two further cohorts. Both CGGA mRNA-seq releases, stratified to match the GLASS design (primary, WHO IV, IDH-wildtype). Three cohorts agree to within 0.042, inside the a priori equivalence margin. A discrepant fourth cohort, with three explanations tested and excluded. CGGA-325 gives 1.420. Sequencing depth was tested by binomial thinning of real integer counts (shift -0.004), cohort size by subsampling (+0.003 at n = 74), and clinical composition by a pre-registered test of MGMT status (predicts 0.002 of 0.089, and the gap survives within both strata). Reported as unexplained. Model comparison against Laplace. The q-Gaussian is preferred by AIC in 87% of samples against an exponential-tailed alternative, with median dAIC above 120 in every cohort. A cohort-size effect in the estimator, characterised. Deviations from a cohort-median reference inflate the fitted tails when few samples are present (+0.032 at n = 20, +0.003 at n = 74); a fixed reference profile removes the effect, identifying the pathway. Applies to any index built the same way. Carried over from version 3. Estimation error measured by within-sample gene splitting (1–14% of the observed variance); a repeatability floor of 0.070 from 90 TCGA participants with replicate libraries, shown not to be explained by library composition; the finding that chemoradiotherapy moves the index no more than re-sampling the same tumour (F = 0.86, p = 0.78); equivalence testing with patient-clustered standard errors verified by a pairs-cluster bootstrap; and a circular covariate documented (the interquartile range of the deviation vector predicts q at R² = 0.94 but is a deterministic function of the fitted parameters). Contents. code/ analysis scripts in execution order; data/ open-access TCGA-GBM inputs with the GDC manifest; results/ one CSV per analysis; figures/ main and supplementary; legacy/ superseded artefacts, each with a file explaining why it is retired, including one retracted analysis. Reproducibility. The deposit is built from a git commit, recorded in the README. code/19_verify_manuscript_numbers.py recomputes every number printed in the manuscript from the tables in results/, confirms that each appears in the manuscript source, and exits non-zero if any disagrees. It passes on this build: 107 checks, 0 failures. Data availability. Raw RNA-seq and clinical data for the primary cohort are available from the GLASS consortium (Synapse syn17038081) under its terms of use and are not redistributed here; derived per-sample values remain subject to those terms, including a prohibition on commercialisation. See NOTICE.txt. TCGA-GBM data are open access from the NCI Genomic Data Commons and are included. CGGA data are openly available from cgga.org.cn and are not redistributed. Use of AI-assisted tools. Part of the analysis code was written with the assistance of a large language model (Claude, Anthropic), as described in the manuscript. All code was read, executed and verified by the authors, who take full responsibility for its correctness.