Skip to content

A Framework and Prototype for a Navigable Map of Datasets in Engineering Design and Systems Engineering

Mar 2026 · arXiv.org · Vol abs/2603.15722 · 0 citations · 57 references
Computer Science

TL;DR

A systematic framework for a Map of Datasets in EDSE is proposed, built upon a multi-dimensional taxonomy designed to classify engineering datasets by domain, lifecycle stage, data type, and format, enabling faceted discovery.

Abstract

The proliferation of data across the system lifecycle presents both a significant opportunity and a challenge for Engineering Design and Systems Engineering (EDSE). While this"digital thread"has the potential to drive innovation, the fragmented and inaccessible nature of existing datasets hinders method validation, limits reproducibility, and slows research progress. Unlike fields such as computer vision and natural language processing, which benefit from established benchmark ecosystems, engineering design research often relies on small, proprietary, or ad-hoc datasets. This paper addresses this challenge by proposing a systematic framework for a"Map of Datasets in EDSE."The framework is built upon a multi-dimensional taxonomy designed to classify engineering datasets by domain, lifecycle stage, data type, and format, enabling faceted discovery. An architecture for an interactive discovery tool is detailed and demonstrated through a working prototype, employing a knowledge graph data model to capture rich semantic relationships between datasets, tools, and publications. An analysis of the current data landscape reveals underrepresented areas ("data deserts") in early-stage design and system architecture, as well as relatively well-represented areas ("data oases") in predictive maintenance and autonomous systems. The paper identifies key challenges in curation and sustainability and proposes mitigation strategies, laying the groundwork for a dynamic, community-driven resource to accelerate data-centric engineering research.

View source

Similar papers

Open access Aug 2026

VESA: A Visualization-Enabled Search Application for Exploratory Dataset Discovery

The increasing complexity and scale of scientific datasets demand advanced tools for efficient discovery and exploration. Traditional search systems often fall short in addressing the multidimensional nature of data and their inherent relationships, limiting their effectiveness for exploratory search. This paper presents the visualization-enabled search application (VESA), a backend-agnostic visual analytics framework that supports interactive and multidimensional exploration of heterogeneous scientific metadata. VESA enables cross-repository data discovery through configurable adapter interfaces that map repository-specific metadata onto a minimal set of common discovery dimensions, enabling seamless ingestion and consistent processing. At the frontend, coordinated visualizations, including spatial, temporal, and relational views, support exploratory search and sensemaking. To demonstrate the functionality and usefulness of the system, a software prototype is developed and applied to Earth System Science repositories, showing how heterogeneous sources can be integrated and explored through a unified interface. The framework is evaluated against established design guidelines and further validated through an online user study. In addition, adapters for two data repositories are implemented, illustrating how different backends can be connected to the system. Results indicate positive user reception, highlighting VESA’s usability, low learning curve, and its potential to enhance data discovery workflows through interactive visual exploration.

Tobias Hecking, Hudaif Mohammad Malikathazham, A. Gerndt et al. · 0 citations
Open access Jul 2026

The PRIMA Thesaurus for Materials Science and Engineering

Materials science and engineering (MSE) is characterized by heterogeneous workflows, often coupled with limited research data management (RDM) practices. In particular, provenance metadata are often confined within individual electronic laboratory notebooks (ELNs), highlighting the need for semantic resources that support findable, accessible, interoperable, and reusable (FAIR) data, while remaining accessible to nonexperts. The presented Provenance information for materials science (PRIMA) Thesaurus addresses this challenge by providing a structured vocabulary for describing experimental and computational workflows, without relying on complex conceptual models or formal axiomatization. Its development, based on an iterative process involving domain experts, includes requirement analysis across multiple techniques, selection and harmonization of concepts, and alignment with existing community standards. The resulting terminology, structured in a few hierarchical layers, is implemented in Simple Knowledge Organization System (SKOS) to ensure flexibility and ease of integration. The applicability of PRIMA as a versatile resource capable of serving diverse scientific disciplines is demonstrated through use cases spanning theoretical and experimental condensed‐matter physics, metrology, surface science, and metallic biomaterial research. These examples illustrate the integration of concepts into various ELNs and metadata schemas, showing how this shared semantic layer supports consistent data description, improves discoverability, and enables cross‐platform interoperability.

R. Aversa, A. Boubnov, D. De Angelis et al. · 0 citations
Open access Jul 2026

A novel pipeline and benchmark for automated technical datasheets processing

Technical datasheets are fundamental for manufacturing tasks like design, engineering, procurement, and maintenance, but their varied, unstructured PDF formats make them time-consuming to compare and their data remains largely unexploitable by automated systems. This lack of data accessibility necessitates manual processing, which is both time-consuming and error-prone, ultimately reducing the overall usability of the information. Although recent research has addressed knowledge extraction from manufacturing process specifications and semi-structured web sources, no methods effectively overcome the limitation of extracting data from highly heterogeneous, unstructured PDF technical datasheets. Automating this process involves challenges that go well beyond conventional text extraction, requiring advanced semantic modeling to accurately identify, classify, and interrelate product attributes within structured and hierarchical formats (e.g., JSON-based schemas). The difficulty is further amplified in multi-product catalogues, where contextual ambiguity and interdependencies increase significantly. Current benchmarks do not adequately address these challenges, as their limited metrics fail to reflect the depth of semantic understanding necessary for industrial applicability. This paper seeks to address this gap by introducing a highly adaptive automated extraction pipeline and systematically comparing its performance with two alternative extraction strategies: a human-assisted annotation workflow and a zero-shot approach: Using three representative datasets, covering heterogeneous single products, multi-product catalogues, and a homogeneous single-company archive, we systematically benchmark each method’s accuracy, efficiency, and robustness. Our findings quantify the trade-offs between human effort, automation, and large language models performance, providing a data-driven analysis of modern models’ capabilities and limitations in handling the unique complexities of industrial documentation and offering a methodology for developing more scalable and reliable solutions for semantic extraction from industrial documentation.

Lorenzo Cutrupi, A. Laborde, D. Knüttel et al. · 0 citations
Aug 2026

Balancing Richness and Reliability: An Explore-Construct-Verify Framework for API Knowledge Graph Construction

This work proposes Explore-Construct-Verify (ECV), a three-stage framework for API KG construction using large language models (LLMs), which preserves LLMs’ ability to discover domain-specific knowledge while enabling efficient post-hoc validation.

Yanbang Sun, Qing Huang, Zhenchang Xing et al. · 0 citations
Open access Jul 2026

Semantic Modeling in Materials Science and Engineering With Platform MaterialDigital Core Ontology 3.0

Materials Science and Engineering (MSE) increasingly relies on data‐intensive, automated, and distributed workflows that span synthesis, manufacturing, characterization, design, and simulation. These settings require machine‐actionable representations of materials and processes that remain interoperable across laboratories, software stacks, and organizations. Therefore, Platform MaterialDigital Core Ontology (PMDco) 3.0 is introduced as a mid‐level ontology that provides a semantic framework for the processing–structure–properties paradigm in MSE. PMDco 3.0 adopts an architecture aligned with the Basic Formal Ontology that enables a logically consistent classification of fundamental MSE concepts and the explicit representation of intrinsic material properties, contextual roles and functions, and related information artifacts. The work outlines the technical curation approach that supports sustainable ontology evolution through reproducible builds, automated release generation, and systematic validation workflows. Representative semantic patterns are presented as reusable building blocks for consistent modeling and data mapping, including material object duality, intensive versus extensive qualities, role and function assignment, immaterial entities for spatial context, process modeling across production, assay, and computation, and the separation of requirements from observations via set points and measurements. PMDco 3.0 is intended to serve as a community‐driven anchor for interoperable domain and application ontologies and scalable semantic interoperability in MSE.

Markus Schilling, P. von Hartrott, Jörg Waitelonis et al. · 0 citations
Preprint Aug 2026

ModBench: A Pipeline for Building Modelica Benchmark Datasets Mined from Library Repositories

Research on equation-based cyber-physical systems modeling languages, such as Modelica, is constrained by the lack of curated benchmark datasets. This limits empirical insight into the evolution and development of models. We address this gap with ModBench, a pipeline that mines Git repositories of Modelica libraries to produce benchmark datasets of model snapshots. The pipeline (1) filters repository commits to retain human-authored, Modelica-relevant revisions; (2) extracts simulation-eligible classes; and (3) builds canonical representations of Modelica classes. For empirical validation, we applied ModBench to the Modelica Standard Library (MSL) and report the resulting dataset, spanning the full commit history (since Modelica language v3), with 85,562 distinct class snapshots, and links enabling traceability to original models and Git metadata. The dataset, its API, and the data generation pipeline are publicly available to support future research on model evolution analysis, compiler testing, and automated model repair or generation.

Masoud Sadrnezhaad, Martin Sjölund, A. Pop et al. · 1 citation

Related blog posts