Large language models are increasingly used to read earnings calls, investor-relations Q&A, guidance, and disclosure language. In this setting, supervised financial NLP benchmarks can become evidence for vendor selection, deployment approval, and model-risk records. Gold labels, however, do not make a benchmark score a...
Si-Di Chang, Pei-Ke Zhu, Yu-Xiao Chen et al.· IEEE Conference on Computati...· 0 citations
The result is a bounded rule for interpreting aggregate agent behavior: first establish exposure, then score change, and abstain when the trace cannot support the claim.
An executable contract linking observable support, statistical calibration, and justified claims is contributed an executable contract linking observable support, statistical calibration, and justified claims.
This single-workflow forensic case is an existence proof of a failure mode, not an estimate of its prevalence: existing human work supports an exploratory audit of synthetic proposals, but not LLM-judge operating characteristics, clinical validity, corpus prevalence, or robust inter-annotator agreement.
ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim, is introduced.
The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks, and contributes a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting.
Pei-Ke Zhu, Si-Di Chang· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.