Skip to content
Preprint

V-FiLLM: Verified Financial LLM Reasoning Benchmark

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction is introduced, suggesting that targeted, low-cost adaptation is a promising direction for compositional reasoning in financial QA.

Abstract

While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction. Trees are evaluated symbolically to obtain ground truth and rendered into natural-language questions, removing any model from the labeling loop, so items can be generated at arbitrary scale without annotation cost and without inheriting a generator's error rate. V-FiLLM exposes four independently controllable axes of difficulty including computation depth, expression breadth, financial concept complexity, and context size. By evaluating on open-source models, we find that accuracy falls up to 51% as reasoning depth increases, and up to 47% points under adversarial numerical perturbations, highlighting remaining challenges in robust financial reasoning over tables. We further show that lightweight LoRA fine-tuning on verified chain-of-thought traces improves accuracy from 81.1% to 85.6% on held-out problems and outperforms the base model by 5% points on FinQA (Chen et al., 2022a), s), suggesting that targeted, low-cost adaptation is a promising direction for compositional reasoning in financial QA.

View source

Similar papers

#natural language process... Preprint Sep 2026

Towards Robust Numerical Claim Verification

Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Usin...

Peter Røysland Aarnes, Vinay Setty · 0 citations
Preprint Aug 2026

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling

PluginEval is introduced, a benchmark constructed through a two-stage framework that systematically mitigates limitations of Large Language Models and evaluates five model families, including proprietary models and models with open weights, analyze their performance across difficulty levels and error categories, and va...

Dong Xu, Julius, Hanchi Dong et al. · 0 citations
#artificial intelligence Review Sep 2026

BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

BENCHCOMPASS is introduced, a payment-domain benchmark whose construction pipeline builds scenario-grounded tasks from typed evidence packs, applies LLM-based quality checks, creates task-input attack variants, and reserves final item admission for domain experts.

Si-Jie Dong, Wei-Feng Ren, Xuan-Wei Hu et al. · 0 citations
#artificial intelligence Review Sep 2026

EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models

EADC, a novel advanced evaluation benchmark of LLMs based on an AI compliance knowledge graph and AI compliance legal experts, is introduced, offering a rigorous, AI laws and regulations-aligned benchmark to safeguard high-level and deep compliance in the application of LLMs.

Yan Zhang, Rui-En Li, Yao-Yao Peng et al. · 0 citations
Preprint Aug 2026

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems

The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search. Prior process-level methods focus on the coherence and redundancy of chain-of-...

Shunwen Bai, Ziping Ma, Chaoyang Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CruxBench: A Benchmark of Information Discovery

Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whos...

Hui Dai, Li-Na Piao, Nick Merrill et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.