Skip to content

Automated Regulatory Compliance Question Answering in Financial Services with Domain-Adapted Retrieval-Augmented Generation

Sep 2026 · 0 citations · 59 references
Computer Science

TL;DR

This work studies whether a carefully domain-adapted retrieval-augmented generation pipeline closes the gap between compact and compact model quality in financial institutions under dense, frequently amended rulebooks.

Abstract

Financial institutions operate under dense, frequently amended rulebooks, and answering a compliance question correctly requires not only fluency but verifiable grounding in the authoritative text. Large language models are attractive for this task, yet the models that firms can realistically deploy on-premise are compact ones, and compact models hallucinate obligations. We study whether a carefully domain-adapted retrieval-augmented generation pipeline closes that gap. Our retriever is built in three stages on top of LegalBERT: entailment tuning that recasts question--passage matching as premise--hypothesis reconstruction, contrastive tuning with in-batch negatives, and score-level fusion with BM25. Our generator is a compact model (2B--12B parameters) served under 4-bit quantization, either prompted or adapted with retrieval-aware fine-tuning (RAFT) through LoRA. On ObliQA, a question-answering benchmark built from the Abu Dhabi Global Market rulebooks, the staged retriever raises Recall@10 from 0.256 to 0.774 and outperforms BM25 (0.678) and E5-large-v2 (0.758), the strongest general-purpose dense encoder we tested. RAFT-LoRA then improves the composite RePASs answer-quality score for every model we could adapt, with the largest gain on the weakest one. However, the adapted models do not transfer to Australian case-law questions, and a closed-book model that receives no passages at all scores within 0.011 RePASs of the full pipeline while producing answers that cite nothing and misstate obligations. The retrieval gain is therefore measured directly, the generation gain is a gain in RePASs rather than demonstrated grounding, and grounding itself requires an evaluation protocol that RePASs does not provide.

View source

Similar papers

Open access Sep 2026

TempFinRAG: Multimodal Temporal Retrieval-Augmented Generation for Point-in-Time Financial Question Answering

This work introduces TempFinRAG, a point-in-time evaluation protocol built from public filings and XBRL facts, and introduces TempFinQA, a point-in-time evaluation protocol built from public filings and XBRL facts, and evaluates the framework on complementary evidence-grounded, numerical, conversational, and multi-tabl...

Lanju Tao, Zheng-Ji Li, Ying-Rui Ji et al. · 0 citations
Conference Aug 2026

Retrieval-Augmented Fine-Tuning with Reasoning Distillation for Vietnamese Medical Question Answering

Medical question answering (QA) plays a crucial role in clinical decision support, yet robust performance requires models to effectively distinguish relevant evidence from topically similar distractors within retrieved contexts. Existing Vietnamese medical QA benchmarks, however, focus exclusively on zero-shot evaluati...

Nhan Phuoc Thanh Tran, P. Huynh, Trương Quốc Tuấn Trương et al. · 0 citations
#small language model Open access Aug 2026

Retrieval Granularity as Evidence Design in Small-Model RAG Question Answering: A Diagnostic HotpotQA Study

Results align with a diagnostic perspective on chunking: using evidence at a task-appropriate level of granularity can improve grounding, auditability, and answer quality, but the observed patterns should be interpreted within the HotpotQA distractor setting, fixed generator, and tested context budgets.

Wei-Mao Ke, Li-Xia Yang, Meng-Yang Xu · 0 citations
#artificial intelligence Preprint Sep 2026

IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA

We present the IGT system for PolyFiQA Task 2 of the FinMMEval Lab at CLEF 2026, a multilingual financial question answering task over English SEC filings and multilingual news articles (English, Chinese, Japanese, Spanish, Greek) for four companies. Our central observation is that the 344 development questions divide...

Yu-Wen Chiu · 1 citation
Preprint Sep 2026

FinRegQA-EU: Corruption-Based Preference Data for Grounded EU Financial Regulatory Question Answering

Large Language Models (LLMs) struggle with region-specific factual knowledge, particularly in financial regulation. While benchmarks such as CFinBench, provide broad coverage of financial knowledge in other regions, no comparable resource exists for European financial regulation. We close this gap by proposing an end-t...

Aulia Kharis Rakhmasari, Fang-Rong Yu, A. Hoyle et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks'financial statements

FinRAG-QA is introduced, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10 standardised indicators, grounded in 209 annual and Pillar 3 reports from 24 major European and U.S. banks spanning 2019-2023.

Arianna Miola, Bruno Spaccavento, Lorenzo Silotto et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.