A schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard, and demonstrates generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages.
Abstract
We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving as an information model encoding domain knowledge, provides a unified, systematic, and consistent framework for extraction of hierarchical, nested information, with attributes of variable cardinality, and subsequent evaluation of the results. Information extraction from a document is performed in a single call to the model, in zero-shot mode. In the evaluation step, we introduce a path-based semantic matching algorithm to align the nested, variable-cardinality attributes in the extracted results with those in the gold standard. We use generative AI for semantic comparison of the extracted and gold standard values of an attribute, and introduce a rubric to classify the result of the comparison, according to domain-specific considerations, as an exact, semantic, useful, or non-match. We were able to extract 12 out of 14 attributes with an F1 score of $>$90\% from documents published by the health technology assessment organisation NICE, using the generative AI model Claude Opus 3. The time needed to extract the attributes from a document was $\sim$30 times lower than the time taken by a human domain expert. We further demonstrate generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages.
The rapid growth of unstructured textual data necessitates automated approaches for transforming such information into structured, machine-readable knowledge. Knowledge Graphs (KGs) provide an effective framework for representing entities and their relationships; however, existing methods often suffer from fragmented pipelines, limited semantic consistency, and challenges in handling domain-specific variations. This paper presents an intelligent and scalable approach for knowledge graph construction from semantically enriched keyword-based inputs derived from a context-aware extraction process. The proposed method employs a unified pipeline comprising entity identification and ontology-based linking, context-aware relation extraction, and structured triple generation in the form of subject-predicate-object (SPO) representations. The generated triples are futher transformed into RDF format and organized into a coherent knowledge graph, followed by refinement steps to ensure semantic consistency and structural integrity. The approach is evaluated on representative datasets, including PubMed abstracts, and demonstrates improved performance in triplet extraction, entity and relation accuracy, and graph-level quality metrics such as density, clustering coefficient, and modularity. Comparative analysis with baseline methods highlights the effectiveness of the proposed approach in generating coherent and semantically enriched knowledge graphs. Additionally, the system exhibits strong scalability and computational efficiency, making it suitable for large-scale and real-world applications acreoss diverse domains. Overall, the proposed approach effectively bridges the gap between unstructured text and structured knowledge representation, enabling reliable, scalable, and high-quality knowledge graph construction.
Avinash Gondal, Sunil Wankhade· Dandao Xuebao/Journal of Bal...· 0 citations
LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.
Boyang Zhang, Adrian Lyjak, Elizabeth Stewart et al.· 1 citation
The findings suggest that AI-based structured extraction may redefine how organisations formalise expertise, shifting from document-centric storage toward schema-driven knowledge architectures.
Dilyan Georgiev, E. Gourova· European Conference on Knowl...· 0 citations
A production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology, and improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect.
Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik· 0 citations
To address the problems of complex process knowledge sources, heterogeneous representations, dispersed semantic associations, and limited reusability in the domain of machining distortion of thin-walled parts, this study proposes a knowledge graph construction method for the workpiece machining distortion domain, together with an intelligent decision-making framework driven by the collaboration of knowledge graphs and large language models. First, a domain ontology model is established around core concepts, including workpiece objects, deformation-driving factors, analytical resources, analytical methods, and optimization knowledge, thereby providing a unified semantic foundation for domain knowledge organization. Second, considering the characteristics of domain texts, such as dense technical terminology, ambiguous entity boundaries, and complex relation expressions, a dual-channel knowledge extraction method integrating BERT-BiLSTM-CRF and Universal Information Extraction (UIE) is developed to achieve high-precision extraction of entities and relations from unstructured texts. Knowledge fusion is further carried out through cross-validation, entity disambiguation, coreference resolution, and semantic alignment, and the extracted knowledge is ultimately stored and organized in Neo4j. Furthermore, an intelligent decision-making framework based on the collaboration of knowledge graphs and large language models is constructed. In this framework, a LoRA-tuned Qwen model is employed for user intent recognition and key information extraction, RapidFuzz WRatio is adopted for similar-node retrieval, and local subgraph construction, Label Propagation-based community detection, Betweenness Centrality-based key-node analysis, and evidence fusion are integrated to support process recommendation and intelligent question answering. Based on the proposed framework, an intelligent decision-making system is further developed for process recommendation and intelligent question answering in machining distortion scenarios. Experimental results show that the proposed dual-channel knowledge extraction model achieves an F1-score of 0.88, demonstrating its effectiveness in knowledge acquisition for the machining distortion domain. The constructed knowledge graph contains 4639 entities and 5822 relations, enabling a systematic representation of machining distortion knowledge. Case studies further demonstrate that the proposed method can generate interpretable recommendation results under complex process constraints in real industrial query scenarios. Overall, the proposed approach provides a feasible pathway for the structured organization, intelligent retrieval, and decision support of workpiece machining distortion knowledge.
Deguo Yao, Zhaoze Sun, Jie Gao et al.· Applied System Innovation· 0 citations
Extracting reliable knowledge from unstructured materials literature remains a central bottleneck for data-driven and AI-enabled materials discovery. Large language models (LLMs) are reshaping this task by integrating multimodal document parsing, ontology-guided semantic grounding, structured extraction, and agentic verification into increasingly unified workflows. This review analyzes these developments through a Perception–Cognition–Action lens. At the perception layer, we examine how scientific document parsers, multimodal LLMs, table and chart readers, and optical chemical-structure-recognition systems convert visually rich papers into computable evidence. At the cognition layer, we discuss how ontologies and knowledge graphs constrain LLM outputs, support entity alignment, and reduce semantic ambiguity. At the action layer, we compare schema-based extraction, schema-free discovery, and agentic extraction as a control–coverage–autonomy spectrum rather than a simple succession of tools. We further argue that reliability is the decisive criterion for large-scale deployment, and synthesize failure modes, layered defenses, and evaluation protocols that connect source grounding, ontology constraints, physical verification, and human-in-the-loop review. By distinguishing demonstrated extraction capabilities from more speculative AI-scientist and self-driving-laboratory visions, this review provides a comparative and risk-aware account of how LLM-driven systems can produce evidence-linked, physically meaningful, and reusable materials knowledge.
Shuai Yang, Yimeng Wang, Qiong Tu et al.· Journal of Materials Informa...· 0 citations