This study re-run two recent methods, DEMO and NSReg, together with OUTPOST, a small first-order detector built for this study, and finds that a 0.002 tie band for hyperparameter selection lies below the paired standard error on all six graphs tested, even at ten seeds.
Abstract
Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers are copied from earlier papers, and most anomalies are minority classes relabeled as anomalous. We ask how much of the reported ranking these conventions decide. We re-run two recent methods, DEMO and NSReg, together with OUTPOST, a small first-order detector built for this study. All three use one protocol with identical seeds and splits on eight graphs (seven for the baselines, which cannot run on ogbn-mag), ten seeds each, and every run is scored under both the best-epoch rule and a deployable validation rule. Before the runs that test them, we registered 40 predictions. Three findings hold. First, the rule changes the leader: under the best-epoch rule, OUTPOST and NSReg each lead three of seven graphs, while under the validation rule, NSReg leads five. Second, the best-epoch bonus depends on how the benchmark was built: 0.045--0.080 AUC-ROC on the three small relabeled-class graphs and 0.002--0.014 on the three real fraud graphs. Third, pseudo-labeling in OUTPOST is worth 0.038--0.065 AUC-ROC on the same three graphs but gives no benefit on any real fraud graph. We also show that a 0.002 tie band for hyperparameter selection lies below the paired standard error on all six graphs tested, even at ten seeds. Twelve of our 40 predictions were falsified, and we report them. We close with a short reporting checklist.
Journal entry anomaly detectors are commonly evaluated on the full population with ROC-AUC, precision and recall, ignoring the review budget and which anomaly types are found. We propose a type-aware evaluation combining per-type recall, fair-share type recall (FSR), which caps each type's credit at its budget share, t...
Jan Gronewald, Michel Scherer, Nijat Mehdiyev· 0 citations
SimpleCount is proposed, a reference with no parameter fitting that selects one scalar feature per dataset from a fixed pool of counts, recencies, first-occurrence indicators, and count-derived transforms that matches or exceeds SLADE on three of six datasets and exceeds IsoForest on all six.
Machine-learning intrusion detectors are usually reported with accuracy or F1 on a single train and test split, and their confidence scores are often read operationally as probabilities without an explicit calibration check. We test that assumption. We measure the reliability of the predicted probabilities of three cla...
Khalid Alalawi· International Journal of Adv...· 0 citations
A pool per domain is generated by prompting a multimodal LLM with in-context example windows from that domain by prompting a multimodal LLM with in-context example windows from that domain, and the generated pools match the hand-crafted one under matched selection, and the two cover different domains.
Youssef Attia El Hili, Malik Tiomoko, Corinne Ancourt· 0 citations
Time-series anomaly detection trades off predictive accuracy, computational efficiency, and interpretability. We use a large language model not as the detector but as the author of one: an autonomous research loop in which the model repeatedly edits a single short NumPy program under a leakage-free objective, keeping t...
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026