Skip to content

Author

Nirouyar Reza

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Morphological Hijacking in Frozen Language Models: Recovering Structured Representations and Testing the Limits of Algebraic Composition

This revised manuscript presents a controlled study of morphological hijacking in frozen language-model representations using the synthetic LUXVAR/VARZIN framework. The study shows that adversarial surface structure can dominate frozen representations, while a lightweight trained projection head can substantially recover the targeted categorical/group-position structure under controlled held-out, cross-script, and out-of-distribution tests. The revision incorporates the complete Level-3 composition diagnostic chain. A linear decoder failed on seen-pair composition, after which diagnostics tested optimization, structural rank deficiency, cross-family geometric alignment, and decoder capacity. A pre-registered one-hidden-layer MLP (64 hidden units, ReLU) fit the 120 seen ordered pairs almost perfectly (mean L3-A accuracy = 0.996 across five seeds), showing that decoder capacity was sufficient to fit the training set. Crucially, the same MLP did not generalize to 24 held-out ordered pairs: TRUE accuracy was 0.106 (38/360), compared with 0.281 (101/360) for the shuffled-label control and 0.175 (63/360) for the wrong-operation control. The revised interpretation is deliberately narrow. The intervention provides evidence for recovery of the targeted categorical/group-position structure, but it does not demonstrate systematic algebraic composition. The unseen-pair result is consistent with memorization/interpolation rather than learned application of the (i+j) mod 12 rule. The paper therefore treats composition as an important negative boundary condition, not as evidence that composition information is universally absent from the underlying representations. Conclusions are restricted to the tested synthetic lexicon, GPT-2-small extraction pipeline, preprocessing, alignment procedures, and decoder classes

Nirouyar Reza · 0 citations
#small language model Open access Sep 2026

Morphological Hijacking in Frozen Language Models: A Contrastive, Symmetry-Regularized Projection Head for Algebraic Structure Recovery

Preprint. Not yet peer-reviewed. Frozen autoregressive language models cluster surface-similar tokenstogether even when a stronger, task-relevant structure is available inthe input. This paper documents "morphological hijacking" — near-totalcollapse of algebraic-structure recovery under adversarial surfacecorrelation — across four frozen model families (GPT2-small, Pythia-410m,Mistral-7B-v0.3, Qwen2.5-7B), traces it to representational anisotropy,and introduces a lightweight, trainable projection head (contrastiveobjective + positional-symmetry penalty + rogue-dimension ablation +Soft-PCA initialization) that substantially recovers the structure. Theresult is validated at both the discrete-clustering level (Adjusted RandIndex) and the continuous embedding-geometry level (margin, win rate)across five architecturally diverse models, including two models trainedexplicitly for embedding/retrieval tasks, and further tested on naturalEnglish words and a freshly generated, 10x-larger constructed lexicon.Several negative results are reported alongside the positive ones,including a rejected "globally architectural anisotropy" hypothesis anda rejected layer-selection heuristic. This is a preliminary, honestly-scoped empirical report, not a claim ofgeneral applicability. Limitations, a full pre-submission checklist, andcomplete reproducibility code are included. Author: Reza NirouyarORCID: 0009-0000-4690-6842Contact: contact@varzin.orgProject website: https://varzin.org

Nirouyar Reza · 0 citations
#small language model Open access Sep 2026

Morphological Hijacking in Frozen Language Models: A Contrastive, Symmetry-Regularized Projection Head for Algebraic Structure Recovery

Preprint. Not yet peer-reviewed. Abstract: Frozen autoregressive language models cluster surface-similar tokens together even when a stronger, task-relevant structure is available in the input. We construct an adversarial lexicon (LUXVAR Core-30) in which word meaning is defined by orbit membership under the finite algebraic group Aff(Z12), while a long, orthographically salient prefix is deliberately uncorrelated with that meaning. Across four frozen backbones (GPT2-small, Pythia-410m, Mistral-7B-v0.3, Qwen2.5-7B), unsupervised clustering of final-layer hidden states recovers the surface prefix almost perfectly (Prefix Hijacking Ratio = 1.000) while failing to recover the underlying orbit structure (Adjusted Rand Index ≈ −0.148 on all four models) — a phenomenon we term morphological hijacking. We trace this in part to extreme, low-dimensional anisotropy (“rogue dimensions”) and in part to token-length imbalance under mean pooling. A cross-script replication of the rogue-dimension analysis shows this anisotropy is strongly script-invariant on GPT2-small (10/10 dimension overlap between English and Persian backgrounds) but substantially weaker at 7B scale (3–4/10), moderating an initial “globally architectural” hypothesis. We train a lightweight, frozen-backbone projection head with a supervised contrastive objective, a positional-symmetry penalty, rogue-dimension ablation, and a soft PCA-blended initialization, and show it substantially improves orbit recovery on held-out, cross-script, out-of-distribution word forms. On GPT2-small, the full recipe reaches perfect, zero-variance clustering (ARI = 1.000, σ = 0.000 across 5 seeds). On Mistral-7B and Qwen2.5-7B, the same class of intervention, independently re-tuned per model, yields a substantial and bootstrap-significant improvement (ARI = 0.957 ± 0.090 and 0.880 ± 0.115, respectively) though with residual seed variance. Continuous validation directly on embedding geometry confirms that every tested model moves from a hard 0.000 win rate to a hard 1.000 win rate after training, ruling out a discrete-clustering artifact as the source of the reported gains. Critically, since recovering a category label does not guarantee a verified group action, we additionally test the harder, literal claim on a dedicated Aff(Z12) construction with an opaque (non-leaking) position encoding and a shuffled-label control (Appendix F). The retargeted remediation recovers true group position substantially above the shuffled-label control across all three tested architectures (true-label ARI 0.79–1.00 vs. shuffled-control ARI 0.02–0.30), proving the recovery of genuine algebraic structure rather than a correlated category label. We report all of this together with its limitations, an expanded 10x-larger constructed lexicon (Appendix E), natural language generalization tests (Appendix C), a full pre-submission checklist, and complete open-source reproducibility code. Author: Reza Nirouyar ORCID: 0009-0000-4690-6842 Contact: contact@varzin.org Project website: https://varzin.org

Nirouyar Reza · 0 citations
#small language model Open access Sep 2026

Morphological Hijacking in Frozen Language Models: A Contrastive, Symmetry-Regularized Projection Head for Algebraic Structure Recovery

Preprint. Not yet peer-reviewed. Abstract: Frozen autoregressive language models cluster surface-similar tokens together even when a stronger, task-relevant structure is available in the input. We construct an adversarial lexicon (LUXVAR Core-30) in which word meaning is defined by orbit membership under the finite algebraic group Aff(Z12), while a long, orthographically salient prefix is deliberately uncorrelated with that meaning. Across four frozen backbones (GPT2-small, Pythia-410m, Mistral-7B-v0.3, Qwen2.5-7B), unsupervised clustering of final-layer hidden states recovers the surface prefix almost perfectly (Prefix Hijacking Ratio = 1.000) while failing to recover the underlying orbit structure (Adjusted Rand Index ≈ −0.148 on all four models) — a phenomenon we term morphological hijacking. We trace this in part to extreme, low-dimensional anisotropy (“rogue dimensions”) and in part to token-length imbalance under mean pooling. A cross-script replication of the rogue-dimension analysis shows this anisotropy is strongly script-invariant on GPT2-small (10/10 dimension overlap between English and Persian backgrounds) but substantially weaker at 7B scale (3–4/10), moderating an initial “globally architectural” hypothesis. We train a lightweight, frozen-backbone projection head with a supervised contrastive objective, a positional-symmetry penalty, rogue-dimension ablation, and a soft PCA-blended initialization, and show it substantially improves orbit recovery on held-out, cross-script, out-of-distribution word forms. On GPT2-small, the full recipe reaches perfect, zero-variance clustering (ARI = 1.000, σ = 0.000 across 5 seeds). On Mistral-7B and Qwen2.5-7B, the same class of intervention, independently re-tuned per model, yields a substantial and bootstrap-significant improvement (ARI = 0.957 ± 0.090 and 0.880 ± 0.115, respectively) though with residual seed variance. Continuous validation directly on embedding geometry confirms that every tested model moves from a hard 0.000 win rate to a hard 1.000 win rate after training, ruling out a discrete-clustering artifact as the source of the reported gains. Critically, since recovering a category label does not guarantee a verified group action, we additionally test the harder, literal claim on a dedicated Aff(Z12) construction with an opaque (non-leaking) position encoding and a shuffled-label control (Appendix F). The retargeted remediation recovers true group position substantially above the shuffled-label control across all three tested architectures (true-label ARI 0.79–1.00 vs. shuffled-control ARI 0.02–0.30), proving the recovery of genuine algebraic structure rather than a correlated category label. We report all of this together with its limitations, an expanded 10x-larger constructed lexicon (Appendix E), natural language generalization tests (Appendix C), a full pre-submission checklist, and complete open-source reproducibility code. Author: Reza Nirouyar ORCID: 0009-0000-4690-6842 Contact: contact@varzin.org Project website: https://varzin.org

Nirouyar Reza · 0 citations