Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Research on the construction of retrieval question and answer dataset and optimization of a two-stage retrieval model for the field of technical supervision for power systems

On-site management for technical supervision for power systems is implemented based on diversified unstructured documents including industrial technical specifications, corporate operational codes and fault incident summaries. Conventional keyword-based retrieval suffers from prominent drawbacks such as semantic mismatch, inconsistent alignment between regulatory clauses and search queries, as well as poor precision in fault case lookup. Targeting endto- end practical deployment of domain-specific retrieval tailored to technical supervision, this paper develops an integrated technical pipeline covering structured parsing of heterogeneous multi-source technical documents, automatic compilation of domain question-and-answer corpora, dual-layer human-machine quality control, and fine-tuning of embedding and reranking dual models. To begin with, hierarchical structured extraction rules are formulated for 6 categories of unstructured materials: industry technical standards, internal corporate circulars, implementation guidelines and detailed rules, alongside typical fault investigation reports. Large language models are leveraged to automatically generate paired datasets linking regulatory provisions with corresponding inquiry items, followed by noise reduction via a two-tier quality control mechanism combining preliminary automated screening and selective manual inspection. Next, domain-specific parameter fine-tuning is performed on Qwen3-Embedding-0.6B dense retriever and Qwen3-Reranker- 0.6B reranking model to build a two-stage retrieval framework consisting of coarse-grained dense candidate retrieval and subsequent refined result reranking. Comparative experiments are conducted on an in-house Q&A benchmark for technical supervision and publicly available universal power retrieval datasets. Quantitative results across Recall@k metrics verify that the proposed approach outperforms prevalent baseline alternatives comprehensively and achieves state-of-the-art performance within this niche application field.

Sai Zhang, Xiao Liang, Bochuan Song et al. · 0 citations