ZID is introduced, which combines six standardized location- and dispersion-sensitive arms from a rank graph and Gaussian kernels (GPK at two bandwidths) and reports three linked outputs: an index for ranking departure magnitude, a permutation-value for testing distributional equality, and a signed dispersion readout for diagnosis.
Abstract
Generative models are commonly ranked by Fr\'echet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance obtain FID $24.7$ versus $58.6$ for held-out real images (lower is better). Moreover, FID and KID are scalar discrepancies that are unchanged when the two samples are exchanged and therefore do not encode the direction of a dispersion change: under-dispersion, as can occur in mode collapse, versus over-dispersion. We introduce \textbf{ZID} (\emph{Z-resolved Integrated Diagnostic}), which combines six standardized location- and dispersion-sensitive arms from a rank graph (RISE) and Gaussian kernels (GPK at two bandwidths). Rather than asking one scalar to serve incompatible roles, ZID reports three linked outputs: an index for ranking departure magnitude, a permutation $p$-value for testing distributional equality, and a signed dispersion readout for diagnosis. In controlled experiments, ZID detects a broad range of departures, and its score tracks increasing severity along the corresponding sweeps, including cases in which FID is flat or reversed. On DiT-XL/2 and SiT-XL/2 guidance sweeps, ZID detects departure from real data, and its signed readout labels the high-guidance diversity collapse as under-dispersion.
Equipping a classifier two-sample test with a gradient-boosted discriminator and decomposing it by controlled permutation into marginal, dependency, and numerical-categorical cross components, each read against a fully factorized reference that destroys all dependency while leaving every marginal intact, and against a real-data oracle.
Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected market statistics. These measures provide useful diagnostics but may not capture the joint temporal and cross-level structure of order-book trajectories. We introduce LOB-ID, an embedding-based framework that adapts the Fr\'echet Inception Distance (FID) and Monge Inception Distance (MIND) to LOB data. To obtain domain-specific embeddings, we train the DeepLOB architecture on four months of Level-2 order-book data for five equities. We show that LOB-ID is stable across time, instruments, and embedding checkpoints, and rises monotonically under controlled distortions. We then construct a moment-matching attack against FID and a deep-book perturbation that evades statistic-based evaluation. MIND remains substantially more sensitive to both distortions. Finally, we score five generative LOB models, spanning stochastic baselines and deep learning approaches, and find that LOB-ID ranks them in line with the joint temporal and cross-level structure each captures by construction.
Andreea Bacalum, Zhuohan Wang, Ollie Olby et al.· 0 citations
A marketplace review photograph is a document: platforms approve refunds on it, and generative models drove the cost of forging one to zero. We study that detection problem, so we build a detector and attach an attribution map as its evidence, then measure what that pair delivers on 186,527 images under controls designed to change our conclusions when something is wrong. Compression history, not synthesis, drives naive evaluation: our strongest model reaches 0.9999 PR-AUC (area under the precision-recall curve) on a product-disjoint split, yet falls to 0.7254 once we re-encode synthetics into the real class's format, while five public detectors move by at most 0.07. Aligning one class relocates the cue rather than removing it, and the repaired model then assigns native files a median probability of synthesis of 0.0004. One identical final encode for both classes repairs that, and a three-seed factorial credits the encoding change with the whole gain (+0.176 +- 0.009 PR-AUC). That encode equalises the last stage only: forensic features alone still separate the classes at 0.7145 against a base rate of 0.254. For evidence we test maps causally, against controls that never consult the detector. Whether an attribution ranking exists at all depends on whether the detector reacts to the image. On our first-fix detector, which calls 96 of 100 edited frames real, no map beats a random one. On the detector we selected, twelve of seventeen maps clear that control on edited images and eight on generated ones; perturbation leads both axes and no gradient-CAM variant shows a positive advantage. The trivial controls never clear it, and on generated images the centre prior is worse than random. Our ensembled regional map clears both axes and takes the top pixel AP at 12.4 s per map against 44.9 for occlusion. Clearing a detector-blind control is not yet a faithful explanation, and we demonstrate none.
Leonid Kuturin, Ilya Sotnikov, Mark Khusnutdinov et al.· 0 citations
The Complementary Evidence Guard (CEG), a detector-agnostic wrapper that preserves complementary evidence through a non-compensatory fusion of the base detector, level, and sharpness using only empirical in-distribution percentiles is introduced.
I. M. Jara, Cristian Rodriguez-Opazo, Stephen Gould et al.· 0 citations
This work analyses the linearity and quality of MGT representations and shows that simple linear probes outperform a wide range of detectors while being substantially more sample-efficient, and demonstrates the potential of linear probes as as robust and sample-efficient MGT detectors.
Gerrit Quaremba, Hanqi Yan, E. Black et al.· 0 citations
Diffusion models are increasingly fine-tuned for domain-specific image generation, yet fine-tuning strategies are usually selected with a single evaluation metric. This paper examines when the Fréchet Inception Distance (FID) and a text–image similarity score based on Contrastive Language–Image Pre-training (CLIP) disagree in the ranking of such strategies. The study evaluates 68 configurations built on Stable Diffusion 1.5 with low-rank adaptation: on each of four datasets that span style transfer and subject-driven personalization, a clustering-based curriculum and a random-order baseline are matched across five data-availability regimes, plus seven shared ablation and baseline controls. Three results stand out. The FID–CLIP relationship changes in character across datasets, from a strongly positive association to a sign reversal between the rank and linear correlations. The two metrics select different winners in most head-to-head comparisons, and the practical cost of following the wrong metric ranges from negligible on style transfer to severe on personalization. The random-order baselines are FID-optimal in most comparisons, training time tends to agree with the FID-optimal choice, and a lower training loss is a poor proxy for generation quality. Because each configuration is evaluated with a small generated sample of 30 to 40 images, FID is read throughout as a comparative diagnostic under a fixed protocol and its absolute values are not interpreted. Robustness checks with larger evaluation sets and multi-seed reruns confirm the large-gap conclusions, whereas the close-call cases prove fragile. A validation on Stable Diffusion XL gives initial evidence that these patterns are not specific to one architecture.