Skip to content

The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection

Sep 2026 · 0 citations · 62 references
Computer Science

TL;DR

This study re-run two recent methods, DEMO and NSReg, together with OUTPOST, a small first-order detector built for this study, and finds that a 0.002 tie band for hyperparameter selection lies below the paired standard error on all six graphs tested, even at ten seeds.

Abstract

Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers are copied from earlier papers, and most anomalies are minority classes relabeled as anomalous. We ask how much of the reported ranking these conventions decide. We re-run two recent methods, DEMO and NSReg, together with OUTPOST, a small first-order detector built for this study. All three use one protocol with identical seeds and splits on eight graphs (seven for the baselines, which cannot run on ogbn-mag), ten seeds each, and every run is scored under both the best-epoch rule and a deployable validation rule. Before the runs that test them, we registered 40 predictions. Three findings hold. First, the rule changes the leader: under the best-epoch rule, OUTPOST and NSReg each lead three of seven graphs, while under the validation rule, NSReg leads five. Second, the best-epoch bonus depends on how the benchmark was built: 0.045--0.080 AUC-ROC on the three small relabeled-class graphs and 0.002--0.014 on the three real fraud graphs. Third, pseudo-labeling in OUTPOST is worth 0.038--0.065 AUC-ROC on the same three graphs but gives no benefit on any real fraud graph. We also show that a 0.002 tie band for hyperparameter selection lies below the paired standard error on all six graphs tested, even at ten seeds. Twelve of our 40 predictions were falsified, and we report them. We close with a short reporting checklist.

View source

Similar papers

#artificial intelligence Review Oct 2026

Machine learning for journal entry testing: A type-aware evaluation of anomaly detectors under a review budget

Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall (FSR), which caps each type's credit at its budget share, t...

Jan Gronewald, Michel Scherer, Nijat Mehdiyev · 0 citations
#edge computing Preprint Aug 2026

Beat the Counter First: A Baseline for Temporal-Graph Anomaly Detectors

SimpleCount is proposed, a reference with no parameter fitting that selects one scalar feature per dataset from a fixed pool of counts, recencies, first-occurrence indicators, and count-derived transforms that matches or exceeds SLADE on three of six datasets and exceeds IsoForest on all six.

Omair Shafi Ahmed, Zohair Shafi · 0 citations
Open access 2026

When Calibrated Detectors Meet New Attacks: Per-Category Reliability of Machine-Learning Intrusion Detection Under Distribution Shift

Machine-learning intrusion detectors are usually reported with accuracy or F1 on a single train and test split, and their confidence scores are often read operationally as probabilities without an explicit calibration check. We test that assumption. We measure the reliability of the predicted probabilities of three cla...

Khalid Alalawi · 0 citations
#artificial intelligence Preprint Sep 2026

LLM-Generated Feature Pools for Time Series Anomaly Detection

A pool per domain is generated by prompting a multimodal LLM with in-context example windows from that domain by prompting a multimodal LLM with in-context example windows from that domain, and the generated pools match the hand-crafted one under matched selection, and the two cover different domains.

Youssef Attia El Hili, Malik Tiomoko, Corinne Ancourt · 0 citations
#artificial intelligence Preprint Oct 2026

Have an LLM Write Your Anomaly Detector: Autonomous Discovery of Compact, Interpretable Detectors for Time Series

Time-series anomaly detection trades off predictive accuracy, computational efficiency, and interpretability. We use a large language model not as the detector but as the author of one: an autonomous research loop in which the model repeatedly edits a single short NumPy program under a leakage-free objective, keeping t...

D. Berghaus · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.