GRALIS-Report: auditable region-level attribution and structured clinical report generation for breast cancer histology
Abstract
Deep learning classifiers for breast cancer histology achieve expert-level accuracy but do not explain which tissue regions drove the diagnosis. We present GRALIS-Report , an attribution pipeline with two defining architectural properties: raw images never enter the report generation stage , and no information is fused across modalities in a latent space . Attribution signals—not pixels—are the sole input to pathology-oriented language; every modality transition is explicit, symbolic, and traceable to the stored structured attribution record, making every visual-totext inference auditable and independently verifiable. The system: (i) trains a DenseNet-121 classifier on BreaKHis via knowledge distillation (high internal accuracy on a patient-level split); (ii) applies GRALIS —which transforms the image into a semantic attribution signal : a per-superpixel importance score ϕ i computed via coalition-conditioned path integration—and (iii) converts that structured signal into a research SOAP-style report. Formal theoretical properties—canonical form, a priori convergence bound, and structural incompatibility of locality with exact completeness—are proved in a companion preprint (arXiv:2605.05480); the present paper is entirely experimental. On the BreaKHis test set (1,187 images), independent faithfulness benchmarks place GRALIS at rank 2 of 6 on pixel-level deletion AUC and rank 2 of 3 on ROAD MoRF AUC (the three methods for which ROAD was computed), with the largest MoRF-LeRF discrimination gap among the three evaluated methods. This ranking reflects a deliberate design trade-off: by operating at superpixel rather than pixel resolution, GRALIS sacrifices marginal pixel-level faithfulness relative to Integrated Gradients in exchange for region level spatial coherence, a pre-run Monte Carlo sample-size bound (not numerically instantiated in this paper), and a fully auditable attribution-to-report pipeline—properties that pixel-precise methods do not jointly provide. Cross-dataset evaluation on two held-out external subsets further characterises this trade-off: on IDC Breast Cancer (50 × 50 px patches, frozen backbone), GRALIS ranks first on both Deletion AUC and ROAD MoRF; on PatchCamelyon (96 × 96 px), GRALIS ranks sixth—a result plausibly associated with a mismatch between the fixed superpixel granularity ( n seg = 30) and the finer discriminative feature scale of lymph-node patches, although a dedicated n seg ablation would be required to test this explanation. A supplementary ExpiScore s profile is reported alongside these independent metrics; we caution that this metric shares authorship with the present work and should be weighted accordingly. The deterministic engine generated 1,187/1,187 syntactically complete reports with no execution failures; 1,175 (98.99%) corresponded to correct classifier predictions, with the 12 discordant cases identified retrospectively using test labels, operating fully offline. An expert perception discordance study ( N = 4 anatomopathologists, 60 cases) reveals marked inter-rater variability in perceived clinical utility, suggesting that perceived explanation utility may not constitute a stable ground truth in histopathology. This is reported as a methodological finding for the XAI evaluation community, not as evidence of clinical utility. No clinical efficacy claims are made.