Skip to content
Review

Validation-Driven Automation in a Federated XML Ingest Pipeline: The NCBI Bookshelf Case

· Balisage Series on Markup Technologies · Vol 31 · 0 citations

TL;DR

This work demonstrates how validation can serve not only as a quality assurance mechanism, but as a central organizing principle for workflow automation, enabling scalable and reliable content management across a distributed XML ecosystem.

Abstract

The NCBI Bookshelf, managed by the National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), ingests and makes publicly accessible a wide range of biomedical books and reports contributed by diverse submitters who deposit heterogeneous source formats, including Word, PDF, and XML. Ensuring data integrity across this content has historically required substantial manual intervention due to variability in input formats, metadata quality, and distributed system dependencies. This paper describes the design and implementation of a validation-driven automation framework for Bookshelf that supports scalable conversion, ingestion, processing, and release of content within a federated architecture. The system integrates rule-based validation, workflow orchestration, and identifier-based reconciliation across multiple systems, including content management, XML processing pipelines, PubMed indexing, and Open Access dataset services. Central to the approach is the use of layered validation, including Schematron/XPath and business-rule enforcement, to control workflow transitions, prevent duplicate or inconsistent content, and ensure the completeness and consistency of metadata. Validation is applied at multiple stages in the lifecycle—from submission and conversion through preview and release—and is exposed to submitters and informaticists through actionable feedback mechanisms. The pipeline supports multiple conversion and ingestion pathways, including direct XML submission, PDF-to-XML conversion via tagging vendors, and Word-based authoring workflows. Despite this diversity, all workflows converge on a common processing model governed by validation rules and state transitions tracked in an external workflow system. The paper also describes mechanisms for controlled self-service publishing, automated generation of derived content structures, and cross-system integration through identifier reconciliation. Operational reporting and provenance tracking further support auditability and continuous improvement. This work demonstrates how validation can serve not only as a quality assurance mechanism, but as a central organizing principle for workflow automation, enabling scalable and reliable content management across a distributed XML ecosystem.

View source

Similar papers

From a FAIR Workflow Blueprint to an Executable Data Harmonisation Pipeline: Exploiting the BnF SPARQL Service to Produce and Disseminate an Early Modern Research-Oriented Bibliographic Dataset

The design of a workflow blueprint compliant with FAIR principles and grounded in Open Science practices, developed for the retrieval, harmonisation, and publication of curated bibliographic metadata is introduced.

Arianna Moretti, Iiro Tiihonen, Jonas Fischer · 0 citations

KGpipe: Generation of Pipelines for Data Integration into Knowledge Graphs

The results show that structured RDF pipelines currently provide the most stable integration behavior, whereas JSON and text pipelines remain more sensitive to errors in mapping, extraction, and linking.

Marvin Hofer, Erhard Rahm · 1 citation
Review Open access Sep 2026

Accelerating metadata annotation in collaborative research centers: A hybrid AI workflow for biomedical entities

Collaborative Research Centers rely on FAIR-compliant, richly structured metadata, yet manual annotation is a major bottleneck. We implemented a search-augmented large language model (LLM) workflow within a local research data management system to pre-annotate biomedical entities, using human-in-the-loop verifica...

M. Watter, F. Engel, Aref Kalantari et al. · 0 citations
Preprint Aug 2026

Strengthening LargeRDFBench for Interoperable Federated SPARQL Evaluation

This work strengthens an already valuable community resource by aligning its artifacts with the RDF standards, broadening the set of engines that can be fairly and reproducibly compared, and raises the question of how the results of federated queries under automatic source selection can be made reproducible.

Bryan-Elliott Tam, Muhammad Saleem, Ruben Taelman · 0 citations

MetaFAIR Ecosystem for SSH Digital Libraries: Dataset Packaging, Validation, and SPARQL Publication

The paper discusses how this approach contributes to strategic goals such as sustainability, interoperability, and responsible data stewardship, and illustrates how digital library infrastructures can evolve into data-centric platforms that support reuse, collaboration, and long-term value creation in the SSH domain.

Emiliano Degl'Innocenti, Francesco Pinna, Alessia Spadi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.