Improving Reproducibility in Ontology-Based Data Generation: A Framework and Case Study
Abstract
The exponential growth of heterogeneous data across domains such as healthcare, cybersecurity, and industry has intensified the need for robust semantic organization mechanisms. Ontology-based data generation (OBDG) offers a powerful approach to automatically structure knowledge and enrich information sources through techniques ranging from traditional entity extraction to modern LLM-based generation. However, despite significant methodological advances, a critical reproducibility gap persists, undermining the reliability and adoption of OBDG systems. This paper addresses this challenge through three main contributions: (1) a systematic analysis of reproducibility barriers in current OBDG research, examining issues with resource accessibility, documentation, and evaluation standardization; (2) a structured four-step assessment framework covering accessibility, executability, comparability, and generalizability, accompanied by a practical checklist for authors; and (3) empirical validation through case studies on conversational interfaces and robotics systems. Our analysis reveals that only a minority of OBDG systems provide complete replication packages, with most offering partial or no access to essential resources. The proposed framework successfully identifies key reproducibility bottlenecks and provides actionable guidelines for improving transparency, supporting the development of more robust OBDG systems for real-world applications.