Skip to content
Review Open access

QFRS: quantitative finance reporting standards for forecasting, evaluation and trading claims

Aug 2026 · Artificial Intelligence Review · 0 citations

TL;DR

QFRS underpins a public state-of-the-art leaderboard, ensuring that only studies satisfying these standards are ranked, with the goal of shifting the literature from opaque, error-metric-driven results to transparent, economically meaningful and comparable benchmarks.

Abstract

Financial time-series forecasting lies between AI and market microstructure, but most studies optimise generic error metrics instead of risk-adjusted economic value under realistic frictions. Unlike NLP and vision, the field lacks a shared, reviewer-enforced standard for data handling and evaluation, leading to persistent problems such as data leakage, backtest overfitting and metric-chasing on RMSE/MAE. This paper introduces QFRS a novel, enforceable by reviewers and editors, seven-standard framework and checklist for evaluating and reporting financial asset forecasting and trading claims. QFRS covers quantitative studies on equities (stocks), forex, cryptocurrencies, rates, derivatives (futures, forwards, options, swaps), energy prices, and commodities (gold, oil and silver) and other asset classes. The seven standards specify an end-to-end experimental pipeline, covering (i) dataset construction, (ii) labelling, (iii) point-in-time feature engineering, (iv) leakage-free scaling or normalisation, (v) time-respecting data splits, (vi) evaluation metrics and (vii) cost and slippage-aware backtesting with explicit execution assumptions and decision rules mapping predictions to positions. To validate the standard’s diagnostic value, a compliance audit of Scopus-indexed forex forecasting papers published in 2025 is presented. None of these papers achieved full compliance across all seven standards, with economic backtesting (12.2%) and causal scaling (31.7%) recorded the lowest pass rates. QFRS underpins a public state-of-the-art leaderboard, ensuring that only studies satisfying these standards are ranked, with the goal of shifting the literature from opaque, error-metric-driven results to transparent, economically meaningful and comparable benchmarks. The accompanying leaderboard is available and updated regularly at http://mkhushi.github.io .

Read PDF

Similar papers

Open access Jul 2026

TRuE-XAI: causal and explainable ai framework for trustworthy corporate earnings growth forecasting

This study proposes TRuE-XAI (Transparent, Rule-based, and Explainable Artificial Intelligence), an integrated framework combining imbalance-aware ensemble learning, automated hyperparameter optimization, rule-based explainability, visual analytics, and causal inference for transparent earnings-growth forecasting.

G. Jamnal · 0 citations
Case report Aug 2026

An Open Benchmark for Evaluating Time Series Forecasting Methods Across Financial Markets

Accurate time series forecasts underpin asset pricing, risk management, monetary and macroprudential policy, and other applications. The set of available forecasting methods is expanding rapidly, driven by new machine learning models. This raises a practical question: Do these methods deliver real forecasting power gains on financial data? Since no single method is best across all data generating processes, the question can be answered only by direct evaluation on domain-specific data, and a fair comparison requires holding the data fixed across methods. Yet, new methods are rarely evaluated on the financial data commonly used in academic finance; when they are, each is assessed on its own dataset and cleaning conventions. Apparent method rankings entangle method skill with cleaning choices that themselves require domain expertise. We address this by assembling, in one place, canonical cleaning procedures for many financial datasets to hold the data fixed across forecasting experiments. We introduce a standardized open-source dataset covering equities, corporate bonds, U.S. Treasuries, foreign exchange, commodities, credit default swaps, options, five basis spread datasets, and bank and intermediary indicators, each cleaned per the canonical paper for that asset class. Holding the data fixed and evaluating roughly a dozen univariate methods without exogenous regressors, we find that asset returns remain near unforecastable across every method family and that hybrid and machine learning methods exhibit additional forecasting power on basis spreads and bank indicators.

J. Bejarano, Viren Desai, K. Keshava et al. · 0 citations
Preprint Aug 2026

FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction

Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and economic analysis. However, existing financial benchmarks largely focus on answer-level accuracy and often assume that relevant data are already provided, leaving the assessment of the intermediate process of indicator construction underexplored. In this work, we propose FinDeepIndicator, the first benchmark dedicated to evaluating Deep Research (DR) agents in end-to-end financial indicator construction. Specifically, FinDeepIndicator evaluates DR agents across four stages in indicator construction: formula specification, data collection, indicator calculation, and answer generation, and covers fundamental, technical, and macroeconomic indicators organized into 21 fine-grained sub-categories. It contains 3,350 curated question-answer (QA) pairs derived from both U.S. and Chinese markets, 10 years of historical financial data, and 800 listed companies. Extensive experiments on search-equipped Large Language Models (LLMs) and DR agents show that, while LLMs generally perform well in formula specification, their accuracy drops substantially during data retrieval and numerical execution. DR agents consistently outperform search-equipped LLMs, yet remain unreliable in realistic financial analysis settings. These findings provide insights for developing more capable and trustworthy DR agents in finance.

Chaoqun Yang, Fengbin Zhu, Xinyu Lin et al. · 0 citations
Open access Jul 2026

Adverse Price Excursion Risk Prediction for Margin-Call Early Warning in Forex Trading Using Attention-BiLSTM Across Currency Pairs

Leveraged foreign-exchange trading is exposed to rapid adverse price movements that can contribute to margin-call events, yet most prior studies emphasize price or direction forecasting rather than early risk classification. This study uses the Sample, Explore, Modify, Model, and Assess (SEMMA) framework to organize a time-aware experiment on hourly EUR/USD and GBP/USD data from 2010 to 2026. After cleaning, the datasets contained 99,987 and 99,985 observations, respectively. A direction-agnostic maximum adverse excursion proxy labeled whether either a hypothetical long or short position would experience at least 50 pips of adverse movement within five hours. Sixteen technical features were converted into 60-step sequences and evaluated using identical five-fold chronological walk-forward splits. XGBoost, BiLSTM, Attention-BiLSTM, TransformerEncoder, and PatchTST-Lite were compared using training-only scaling and validation-only threshold selection. XGBoost achieved the strongest mean F1-score and ROC-AUC on EUR/USD (0.3710 and 0.7927) and GBP/USD (0.4855 and 0.7847). Attention-BiLSTM remained competitive, with F1-scores of 0.3583 and 0.4825 and the highest mean recall on GBP/USD (0.6680). In a no-retraining EUR/USD-to-GBP/USD transfer test, it obtained an F1-score of 0.5152 and ROC-AUC of 0.7923. Five-fold Wilcoxon tests lacked sufficient resolution to establish superiority. The results support Attention-BiLSTM as a temporally attributable early-warning component, while XGBoost offers the best efficiency-performance trade-off.

Gea Natasya, Nandang Taufik WM · 0 citations
Open access Jul 2026

Predictive Model Based on Machine Learning to Determine Gold Price Fluctuation and Improve Trading Decisions

Gold’s price reflects currency, opportunity-cost, and safe-haven channels whose strength shifts across regimes, motivating an empirical, data-driven forecasting approach. This study develops a monthly gold price forecasting system for ASM sales-timing decisions in Peru (January 2020–June 2026) using macro-financial predictors including a geopolitical risk index and three U.S. monetary indicators, none of which were Granger-causal and were therefore excluded from the production set. After confirming non-stationarity and Johansen cointegration (four vectors), thirty-two model-feature-set combinations, including Elastic Net, Bayesian Ridge, and a PCA factor, were compared under strict temporal validation with bounded hyperparameter search. The selected model, Ridge regression on the CONTROL feature set, achieved a cross-validation MAPE of 2.29% and test MAPE of 3.62% (official)/3.15% (extended sensitivity window). It was benchmarked against random walk, historical mean, and exponential smoothing and evaluated via the Diebold–Mariano, Clark–West, encompassing, and Model Confidence Set tests (low-power caveats given the small sample). A dual-horizon Monte Carlo simulation, robust to heavy-tailed shocks, projected USD 4482/oz (December 2026) and USD 5106/oz (December 2027). A sales-timing backtest showed a statistically significant result (−0.67%) versus a passive strategy, indicating calibrated price information alone does not yet yield a reliable trading edge, supporting the model’s role as decision support rather than an autonomous trading signal.

Alexander Vladimir Velez Flores, Arturo Rafael Chayña Rodriguez, Wildor Jazmany Jara Vilca et al. · 0 citations
Aug 2026

Corporate Financial Distress Prediction in Vietnam Using Calibrated and Explainable Machine Learning

Corporate financial distress imposes sizable and persistent costs on shareholders, creditors, employees, and the broader economy. However, practical risk governance in emerging markets requires not only accurate risk ranking but also reliable probabilities that can be translated into monitoring thresholds and escalation actions. This study develops an early-warning framework for Vietnamese listed non-financial firms that targets decision-useful one-year-ahead distress probabilities. It evaluates whether calibrated and explainable machine-learning models remain operationally usable under temporal change. The analysis uses a firm-year panel of companies listed on the Ho Chi Minh City Stock Exchange and the Hanoi Stock Exchange over 2014–2024, comprising 7305 observations. A strict chronological design is implemented to emulate forward deployment: model estimation uses 2015–2019, tuning and probability calibration use 2020–2021, and final evaluation is conducted once on an out-of-time hold-out period of 2022–2024. During the test period, distress prevalence increases to 15.42%, compared with approximately 12% in earlier windows. A conventional probabilistic benchmark is compared with multiple machine-learning classifiers under an identical feature space and temporal protocol. An explanation layer is also applied to support governance-oriented interpretation. Out-of-time evidence indicates that tree-ensemble methods provide the strongest combination of ranking performance and probability accuracy, with the random forest providing the best out-of-time performance among the evaluated models, with reasonable discrimination and the lowest probability error, although the magnitude of the AUC improvement should be interpreted as moderate rather than exceptional. The random forest’s calibrated probabilities support monotonic risk stratification and transparent, capacity-constrained watchlists. Selecting the top 10% of firm-years by predicted risk captures 48.3% of distress events with 44.3% precision, corresponding to a 2.87-fold lift over the base rate. Incremental gains from the structural, market-implied distance-to-default proxy are limited once standard accounting and market measures are included. This suggests that routinely available accounting and market variables already capture most of the relevant distress information in this setting. Overall, the results support an out-of-time calibrated and explainable pipeline as a practical foundation for auditable monitoring and tiered escalation in Vietnam’s listed corporate sector.

Tuyen Le Nam, Tam Phan Huy · 0 citations