Skip to content

BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

BENCHCOMPASS is introduced, a payment-domain benchmark whose construction pipeline builds scenario-grounded tasks from typed evidence packs, applies LLM-based quality checks, creates task-input attack variants, and reserves final item admission for domain experts.

Abstract

Payment operations are a critical financial infrastructure, but the value of large language models in this domain remains unclear because payment rules change quickly, evidence is fragmented, and decisions depend on transaction state, participant role, region, and payment rail. Existing benchmarks do not isolate whether failures come from missing payment-rule knowledge, poor use of supplied evidence, or brittleness under imperfect harness inputs. We introduce BENCHCOMPASS, a payment-domain benchmark whose construction pipeline builds scenario-grounded tasks from typed evidence packs, applies LLM-based quality checks, creates task-input attack variants, and reserves final item admission for domain experts. The release contains an expert-reviewed Pro benchmark covering payment knowledge, context-grounded scenario reasoning, and Attacked Open robustness, plus a lower-assurance Normal pool for inspection and future curation. Across 16 model variants, BENCHCOMPASS shows qualitatively different failure modes: missing parametric payment knowledge, incomplete reasoning over supplied rules, and failure to reject plausible but invalid workflows. The benchmark remains unsaturated: the best frontier model reaches 89.6% on Open Context-Grounded Reasoning and 81.7% under attacked inputs, while a representative 32B open-weight model reaches 69.8% and 42.6%. Benchmark data and code are available at https://github.com/ant-intl/BenchCompass.

View source

Similar papers

Review Aug 2026

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is introduced, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions, and shows that broad financial capability scores do not fully capture where models are r...

Su-Yang Zhong, Jingzhe Zhu, Qi Xu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, l...

Prof. S. B. Ghawate · 2 citations
Preprint Aug 2026

V-FiLLM: Verified Financial LLM Reasoning Benchmark

V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction is introduced, suggesting that targeted, low-cost adaptation is a promising direction for compositional reasoning in financial QA.

A. Larsen, Victoire Laurent, Aulia Kharis Rakhmasari et al. · 0 citations
Preprint Aug 2026

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

It is demonstrated that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.

Ofir Ben Shoham, Shrutendra Harsola, Vignesh T. Subrahmaniam et al. · 0 citations
Review Aug 2026

From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation

Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples. We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more...

Minh-Ha Nguyen, Ngoc-Ngo Quang Tran, Thuy Dung Nguyen et al. · 0 citations
Review Open access Sep 2026

A Reproducible Computational Pipeline for Modelling Sequential Decision Thresholds from Stopping-Rule Data

Stopping-rule experiments ask participants to move through ordered stages until they decide to act. These designs are useful for studying intervention thresholds, but raw survey exports are usually stored in wide form, while correct modelling requires scenario-level thresholds and stage-level at-risk rows. This paper p...

A. Kolesnikov · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.