Skip to content

Author

Sahir Maharaj

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#explainable ai Open access Sep 2026

Machine Intuition

Can an artificial intelligence system select the right action before it can articulate a valid reason for that action? Addressing this question at the intersection of reasoning, interpretability, agent safety, and latent computation, we develop a non-anthropomorphic definition of machine intuition as pre-explanatory competence: an action-relevant internal state that is sufficiently informative and causally involved to support a correct decision before a faithful natural-language explanation is available. Synthesizing evidence from chain-of-thought prompting, faithfulness interventions, hidden-state probing, activation steering, mechanistic interpretability, and latent-reasoning systems through 8 August 2026, we note that while final answers and tool-use choices can sometimes be decoded from activations before explicit reasoning begins, difficult multi-step problems are often solved during the generated reasoning trace itself - making token-level deliberation computationally consequential rather than merely explanatory. To separate these regimes, we introduce the Action-Explanation Timing (AET) framework and MINT-Eval, a causal evaluation protocol that distinguishes the earliest causally validated action state from the earliest sufficient, faithful, and interventionally supported explanation, while categorizing faithful latent competence, opaque success, rationalized error, and genuinely deliberative reasoning. The resulting answer is qualified: AI can sometimes act correctly before explaining why, but correctness alone does not establish human-like intuition or trustworthy reasoning; because the same timing gap can arise from useful latent computation, learned heuristics, shortcut features, or post-hoc rationalization, explanations must be treated as evidence to test rather than automatic proof of process, requiring high-consequence actions to undergo external verification, authority controls, and causal audits even when the model's first move is right.

Sahir Maharaj · 0 citations
#explainable ai Open access Sep 2026

Machine Intuition

Can an artificial intelligence system select the right action before it can articulate a valid reason for that action? Addressing this question at the intersection of reasoning, interpretability, agent safety, and latent computation, we develop a non-anthropomorphic definition of machine intuition as pre-explanatory competence: an action-relevant internal state that is sufficiently informative and causally involved to support a correct decision before a faithful natural-language explanation is available. Synthesizing evidence from chain-of-thought prompting, faithfulness interventions, hidden-state probing, activation steering, mechanistic interpretability, and latent-reasoning systems through 8 August 2026, we note that while final answers and tool-use choices can sometimes be decoded from activations before explicit reasoning begins, difficult multi-step problems are often solved during the generated reasoning trace itself - making token-level deliberation computationally consequential rather than merely explanatory. To separate these regimes, we introduce the Action-Explanation Timing (AET) framework and MINT-Eval, a causal evaluation protocol that distinguishes the earliest causally validated action state from the earliest sufficient, faithful, and interventionally supported explanation, while categorizing faithful latent competence, opaque success, rationalized error, and genuinely deliberative reasoning. The resulting answer is qualified: AI can sometimes act correctly before explaining why, but correctness alone does not establish human-like intuition or trustworthy reasoning; because the same timing gap can arise from useful latent computation, learned heuristics, shortcut features, or post-hoc rationalization, explanations must be treated as evidence to test rather than automatic proof of process, requiring high-consequence actions to undergo external verification, authority controls, and causal audits even when the model's first move is right.

Sahir Maharaj · 0 citations
#large language models Open access Sep 2026

Does Becoming Exceptional at One Domain Reduce Transfer Elsewhere

Foundation models derive their value from broad general capability across domains, yet deployment usually rewards specialization - creating a fundamental question for general intelligence: when performance is pushed upward in one region of capability space, is competence elsewhere conserved, redistributed, or destroyed? Synthesizing evidence from continual learning, transfer learning, multi-task optimization, parameter-efficient adaptation, model merging, vision-language adaptation, alignment, and 2023–2026 large-language-model studies, we find that the evidence rejects a universal specialization tax: domain-adaptive pretraining can produce positive transfer, whereas sequential fine-tuning, narrow supervised adaptation, and conflicting objectives can cause catastrophic forgetting, feature distortion, degraded zero-shot transfer, weakened instruction following, or loss of safety behavior depending on task relatedness, update locality, data mixture, optimization geometry, and model capacity. We propose the Generality-Specialization Frontier (GSF), a deployment-oriented framework that treats specialization gain and transfer retention as a Pareto problem, introducing distance-stratified transfer evaluation, invariant retention tests, worst-case regression reporting, and a normalized transfer-elasticity measure. We further propose G-S Bench, an evaluation protocol comparing full fine-tuning, replay, parameter-efficient updates, modular routing, weight interpolation, and non-parametric alternatives under matched target gains, concluding that becoming exceptional at one domain reduces transfer elsewhere only when specialization overwrites shared representations faster than the system preserves broadly useful structure - meaning the tradeoff is an architectural and optimization choice rather than an inevitable law of intelligence.

Sahir Maharaj · 0 citations
#large language models Open access Sep 2026

Does Becoming Exceptional at One Domain Reduce Transfer Elsewhere

Foundation models derive their value from broad general capability across domains, yet deployment usually rewards specialization - creating a fundamental question for general intelligence: when performance is pushed upward in one region of capability space, is competence elsewhere conserved, redistributed, or destroyed? Synthesizing evidence from continual learning, transfer learning, multi-task optimization, parameter-efficient adaptation, model merging, vision-language adaptation, alignment, and 2023–2026 large-language-model studies, we find that the evidence rejects a universal specialization tax: domain-adaptive pretraining can produce positive transfer, whereas sequential fine-tuning, narrow supervised adaptation, and conflicting objectives can cause catastrophic forgetting, feature distortion, degraded zero-shot transfer, weakened instruction following, or loss of safety behavior depending on task relatedness, update locality, data mixture, optimization geometry, and model capacity. We propose the Generality-Specialization Frontier (GSF), a deployment-oriented framework that treats specialization gain and transfer retention as a Pareto problem, introducing distance-stratified transfer evaluation, invariant retention tests, worst-case regression reporting, and a normalized transfer-elasticity measure. We further propose G-S Bench, an evaluation protocol comparing full fine-tuning, replay, parameter-efficient updates, modular routing, weight interpolation, and non-parametric alternatives under matched target gains, concluding that becoming exceptional at one domain reduces transfer elsewhere only when specialization overwrites shared representations faster than the system preserves broadly useful structure - meaning the tradeoff is an architectural and optimization choice rather than an inevitable law of intelligence.

Sahir Maharaj · 0 citations