ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, are introduced to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence to frame financial compliance evaluation as an audit of rule-grounded actions and evidence use.
Abstract
LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules reduce but do not eliminate rejected actions, and incentive or persona framing shifts behavior. A bridge study shows that trader rationales can mislead an independent monitor unless enforcement evidence is shown. In monitoring, simple structured baselines either match or exceed prompt-only LLMs. The results frame financial compliance evaluation as an audit of rule-grounded actions and evidence use, rather than a single compliance score.
What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on three synthetic settings that share one 24-day upward phase m...
This work presents OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents, and builds a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility.
Xinying Cai, Ming-Hao Guo, Jiahe Liu et al.· 0 citations
Large language model (LLM)-based financial agents use market information to generate trading actions together with textual reasoning. However, the resulting executable position transition may conflict with its supporting analysis, while the effect of intervening on that transition becomes observable only after the mark...
Agent benchmarks evaluate policy compliance but assume each policy determines a unique correct action. Natural-language policies can violate this assumption through silence, ambiguity, or contradiction, admitting multiple defensible readings that a single gold trajectory cannot capture. Auditing two $\tau^2$-bench doma...
This work proposes a four-layer framework (Policy, Engineering, Composition, Systemic) grounded in two distinct kinds of evidence, kept explicitly separate, and provides a 90-day implementation sequence spanning trading and payments/customer-facing systems.
This work develops a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system.
Henry L. Han· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.