Skip to content

Restrictive Training (QuAART): Question Answering for the Rapid Development of New Knowledge Extraction Pipelines

· 0 citations · 23 references

TL;DR

This work hypothesizes that using models trained only on generic question answering data (e.g. SQuAD) is a good starting point for domain specific entity extraction, and explores whether the addition of small amounts of training data can help lift model performance.

View source

Similar papers

Conference Open access 2025

Research on Automatic Question Answering System

This review aims to systematically sort out the technical framework of automatic question answering system, analyze its performance bottlenecks, and explore innovative solutions based on large language model and multimodal fusion.

Xuxin Peng · 2 citations
Open access 2026

LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations

LMEnt is released to support studies of knowledge in LMs, including knowledge representations, plasticity, editing, attribution, hallucinations, and learning dynamics, finding that entity co-occurrence and mention forms—which are difficult to study with existing tools—affect learning trends.

Daniela Gottesman, Alon Gilaie-Dotan, Ido Cohen et al. · 0 citations
Preprint Aug 2026

Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge

Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors is presented, describing both the gains and the general-capability cost of staged enterprise adaptation.

Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma et al. · 0 citations
Open access Aug 2026

Optimizing sample selection for large language model-based entity matching using AssistEM

AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.

John Bosco Mugeni, S. Lynden, Toshiyuki Amagasa et al. · 0 citations
Book Open access Aug 2026

RA-QGQA: A Question-Driven Pipeline for Corpus-Grounded Knowledge Graph Verification

RA-QGQA is presented, which recasts triple verification as a question-driven, corpus-grounded task, and demonstrates RA-QGQA as an interactive web system in which users import a KG and its source corpus, verify all triples in a single pass, and inspect the passages that justify its verdict.

Siyang Liu, Hong Duc Nguyen, Yunmiao Li et al. · 0 citations