The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical discussion. We provide a mathematical view of this interpretation and of its assumed causal structure. Besides treating the J-lens as a heuristic probe, we further regard it as a first-order causal transfer operator from intermediate activations to expected future readouts. We study the Jacobian matrix as the optimal local linear approximation of the downstream mapping, analyze its global approximation behavior and bias, and identify its mathematical meaning as an expectation over anticipated future readouts. Further analysis of the Jacobian energy distribution reveals that its causal geometry is highly sparse. The energy decays with depth, concentrates in an extremely small proportion, and decomposes into diagonal pathways and specific critical positions. This decomposition further resolves the expectation of the J-lens over future outputs into short-horizon and sparse concept predictions, providing a more intuitive attribution and explanation for the ability of the J-lens to visualize concepts during the thinking process. Based on the theory, we propose a simple but effective improvement strategy and decoupling method for the J-lens, which significantly enhances the ability of the J-lens to read out correct intermediate concepts.
Mapping anisotropies in the gravitational-wave background (GWB) requires choosing a basis to represent the sky intensity, such as pixel and spherical harmonic bases. Each introduces a truncation---e.g., a number of pixels or a maximum multipole l_max---often set by a naive counting argument that relate the number of measurable modes to the number of independent cross-correlations, N_pair, in a pulsar timing array (PTA). However, such truncations can lead to reconstruction artefacts, as they do not reflect the true information content of the PTA response. Increasing the value of truncation parameters spans the observable space better but reveals poorly constrained modes, making the inverse problem ill-conditioned and requiring regularization. A natural approach is to restrict the reconstruction to a well-measured subspace via principal maps, defined by the dominant eigenmodes of the detector response (or Fisher) matrix. However, these maps are not a fundamental parameterization of the sky, but rather, are derived from an underlying representation---such as a pixelization or a spherical harmonic expansion. While their explicit form depends on basis choice, they can span the \textit{same subspace} when the underlying representation is sufficiently complete. Here, we show that reconstructed anisotropy maps via different bases are equivalent, provided they retain the same information content, i.e., span the same principal subspace. As illustrative cases, we consider a toy model for PTA configuration and several GWB anisotropy shapes: point source, an extended source with deterministic anisotropy, and a statistical isotropic background, along with its summary statistic---the angular power spectrum. Although the reconstructions are equivalent, their computational costs can differ. We conclude with brief comments on the construction of principal maps for ground-based interferometers.
Hybrid inference systems that pair an optical frontend with a digital backend offer a route to offload computation to the physical layer. Yet what the optics should compute, and when this is beneficial, has remained unclear. Here, we show that, for classification tasks, a well-designed optical frontend reshapes the statistics of the sensor intensity readout to improve class separability, quantified by the Bhattacharyya distance. This training-free metric predicts downstream accuracy and reveals that much of the discriminative information resides in inter-pixel correlations. We then identify the roles of coherence and different forms of nonlocality. Because the sensor measures intensity, a linear frontend produces features that are quadratic in the input field; however, only nonlocal, coherent optical systems can exploit the associated information. Such systems can yield significant performance gains, surpassing the best trained linear preprocessor. These results provide physical insights and new design principles for optimal optical--electronic inference systems.
Nicholas Behrens, Yandong Li, F. Monticone· 0 citations
Strong gravitational lensing is a unique probe of the matter power spectrum on small scales, where the abundance of dark matter subhalos in galaxies could be used to distinguish the predictions of the concordance cold dark matter model from alternatives such as warm dark matter. Extracting this signal poses major computational challenges, since perturbations of the lensing potential induced by substructure can be degenerate with both the macro-model of the lens and highly flexible models of the background source. Here, we introduce a framework to quantify these degeneracies using a differentiable strong-lensing simulator. Substructure is represented in a spectral basis confined to an annular domain surrounding the lensed image, allowing signals from a population of NFW subhalos to be encoded in a finite vector space. We then use the Fisher matrix to determine how much information about substructure is absorbed by nuisance components of the model. We find that macro-model degeneracies are largely confined to low-order perturbations of the lensing potential, while degeneracies with the source model can strongly suppress sensitivity across a broad range of scales as the expressivity of the source model is increased. Finally, we introduce the Fisher Graph Laplacian prior as a diagnostic tool to study how the internal degeneracies of the source model can be used to regulate the sensitivity of the data to an unresolved population of dark matter subhalos.
Mutual coupling between antennas has emerged as the dominant direction-dependent corruption in dense aperture arrays, imprinting pronounced sub-MHz spatial and spectral structure that compromises the time-gating and foreground-separation strategies used to isolate the faint 21-cm signal. In this work, we introduce \textit{Direct Primary Beam Correction}, a domain-agnostic framework for reconstructing the far-field radiation pattern relative to an arbitrary reference via a regularised, direction-weighted linear inversion of stacked Jones matrices, thereby enabling the removal of direction-dependent distortions such as mutual coupling. Using full-wave electromagnetic simulations of SKA-Low, we demonstrate that this framework reconstructs the radiation pattern down to the numerical noise floor within a suitably conditioned field of view, with the reconstruction accuracy governed by the regularised inversion and the fidelity of the underlying beam model. Applying the framework to a simulated 4-hour observation of the EoR0 field in the $122$--$134$~MHz band, we identify two principal implications for 21-cm power-spectrum analysis. First, restricting the correction to the main lobe and near sidelobes is inadequate: chromatic grating lobe contributions, whether left insufficiently or entirely uncorrected, continue to contaminate the EoR window. Second, mutual-coupling-induced contamination is temporally coherent and, being anchored to the fixed array geometry, does not average down across snapshots as the EoR field is tracked. Direct primary beam correction, therefore, provides a computationally efficient means of mitigating mutual coupling; however, robust recovery of the EoR window necessitates either full-sky correction or explicit separation of main-beam and sidelobe contributions prior to power-spectrum estimation.
Oscar S D O’Hara, Quentin Gueuning, E. D. L. Acedo et al.· 0 citations
Reconstructing multidimensional vector fields from path-integrated projection data is a fundamental challenge in high-energy-density physics, particularly when experimental sources exhibit spectral broadening and shot-to-shot jitter. We present a physics-guided deep-learning framework that addresses this ill-posed inverse problem by formulating global reconstruction as an aggregation of local inference tasks. By training a neural network on single-particle trajectories in randomized uniform magnetic fields, we develop a “local solver” that demonstrates strong zero-shot transfer to previously unseen magnetohydrodynamic topologies. Central to addressing non-ideal laser-driven proton sources, we introduce a spectral out-of-distribution filter that rejects inputs outside the training energy envelope. By preventing extrapolation, the filter enables accurate reconstruction within the represented training domain while maintaining stable populated-cell reconstruction metrics across diverse spectral conditions. Furthermore, we introduce an a priori reconstruction-reliability indicator based on the in-distribution fraction of the source spectrum, which provides a practical estimate of reconstruction coverage before inference. This approach can be integrated with energy-resolved detector systems, such as stacked nuclear track detectors, establishing a practical and highly parallelizable framework for quantitative plasma diagnostics.
Catherine Jao, Chiung-Yin Chang, Kun-Han Lee et al.· APL Machine Learning· 0 citations
Monte Carlo simulation of calorimeter showers is a principal bottleneck for the High-Luminosity LHC, and diffusion models have emerged as fast, high-fidelity surrogates. Their denoising objective is purely statistical, however: a model can minimize it while placing the physics wrong. Existing physics-informed generative methods cannot close this gap, because they assume a closed-form law, a governing PDE residual or a hard per-sample constraint, that a shower does not supply: no per-sample PDE governs a stochastic cascade, and energy conservation fixes only one scalar per shower. Standard metrics ignore the correlation structure across calorimeter layers and voxels, comparing showers only in a physics feature space. We address both gaps. We introduce the Correlation Frobenius Distance (CFD), a single normalized score for correlation fidelity at layer-wise and voxel-wise scales. We then encode the soft per-sample structure available in a shower as two physics-aware auxiliary losses: a variance-stabilized voxel residual loss grounded in counting statistics, and a graph Laplacian loss over the detector geometry. We combine both with denoising through GradBlend, which anchors the step magnitude to the denoising gradient while letting the auxiliary steer its direction, yielding Lantern, a physics-guided diffusion surrogate. On CaloChallenge Dataset 2, injecting the physics losses through task-symmetric rules such as PCGrad, GradNorm, IMTL-G, and ConFIG inflates FPD by 2-100x relative to denoising alone, whereas GradBlend admits the same signal without regression and, with the Laplacian loss, Lantern improves both FPD and CFD. Our ablation on the auxiliary loss scheduler shows that the voxel residual loss, whose gradient conflicts with denoising, requires a terminal denoising-only phase to preserve shower fidelity, whereas the non-conflicting Laplacian loss is insensitive to the schedule.
Farzana Yasmin Ahmad, V. Venkataswamy, Geoffrey C. Fox· 0 citations