This work presents a pipeline that transforms flat Europeana records into an FDO-compliant knowledge graph structured with CIDOC-CRM, and addresses the core technical challenge of automating the FDO-prescribed distinction between values that must become PID references (resolvable entities) and those that may remain literals (terminal leaves).
Abstract
The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to produce a fully machine-actionable graph in which every reference is resolvable. The Europeana Data Model was designed long before the FDO specification, and it stores most metadata values as plain text. This serves human browsing well enough, but gives an automated agent nothing to follow across records or collections. We present a pipeline that transforms flat Europeana records into an FDO-compliant knowledge graph structured with CIDOC-CRM. Following the FDO specification, we model every heritage entity as a discrete FDO with its own PID, type, profile, and metadata layer. The core technical challenge is automating the FDO-prescribed distinction between values that must become PID references (resolvable entities) and those that may remain literals (terminal leaves such as notes, measurements, and dates). We address this with a large language model that classifies each metadata value, routes it to a controlled vocabulary (Getty AAT, Wikidata, VIAF, PeriodO), and links it to a shared entity FDO. We evaluate using 637 archaeological records from five Europeana providers, processing each with the LLM. The pipeline links 86% of metadata slots, resolving 58.5% of values Europeana had not already enriched. It also merges cross-lingual surface forms that byte-identical matching keeps apart, where 17 of 33 such merges are correct on manual review. Graph connectivity does not separate this from string matching; what distinguishes the FDO graph is that every node is typed and resolvable.
This paper describes the ingestion and ontology-tagging layer that turns a validated extraction stream into a knowledge graph of 537,157 entities and 2,198,567 relationships drawn from 98,795 government documents, and describes a record-identity ladder that decides sameness from identifier columns, name columns, displa...
Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik· 0 citations
Software is a first-class scientific object, yet validated links between source code and the scholarly record remain largely absent from the Linked Open Data (LOD) cloud, isolating archived artefacts from semantic discovery. This paper presents an end-to-end reconciliation pipeline that harvests, validates, and models...
Camillo Carlo Pellizzari di San Girolamo, Francesco Tosoni· 0 citations
The Semantic Tree Advanced module is presented, a specialized hierarchical navigation component for the ResearchSpace platform addressing performance and usability limitations in browsing RDF knowledge graphs.
Remo Grillo, Lucia Giagnolini, P. Bonora· 0 citations
The format, its schema mechanism, and its packaging are introduced: economical enough for an agent to process, and provable enough for its answer to be trusted for enterprise document AI.
J. Paoli· Balisage Series on Markup Te...· 0 citations
Census-style address records contain heterogeneous structures that must be decomposed into fine-grained fields before they can support linkage, geocoding, and administrative processing. This paper describes the design and implementation of a three-stage Planner–Manager–Worker pipeline using a locally deployed open-weig...
Adeeba Tarannum, Muzakkiruddin Ahmed Mohammed, Shames Al Mandalawi et al.· Knowledge· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.