Validation-Driven Automation in a Federated XML Ingest Pipeline: The NCBI Bookshelf Case
TL;DR
This work demonstrates how validation can serve not only as a quality assurance mechanism, but as a central organizing principle for workflow automation, enabling scalable and reliable content management across a distributed XML ecosystem.
Abstract
The NCBI Bookshelf, managed by the National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), ingests and makes publicly accessible a wide range of biomedical books and reports contributed by diverse submitters who deposit heterogeneous source formats, including Word, PDF, and XML. Ensuring data integrity across this content has historically required substantial manual intervention due to variability in input formats, metadata quality, and distributed system dependencies. This paper describes the design and implementation of a validation-driven automation framework for Bookshelf that supports scalable conversion, ingestion, processing, and release of content within a federated architecture. The system integrates rule-based validation, workflow orchestration, and identifier-based reconciliation across multiple systems, including content management, XML processing pipelines, PubMed indexing, and Open Access dataset services. Central to the approach is the use of layered validation, including Schematron/XPath and business-rule enforcement, to control workflow transitions, prevent duplicate or inconsistent content, and ensure the completeness and consistency of metadata. Validation is applied at multiple stages in the lifecycle—from submission and conversion through preview and release—and is exposed to submitters and informaticists through actionable feedback mechanisms. The pipeline supports multiple conversion and ingestion pathways, including direct XML submission, PDF-to-XML conversion via tagging vendors, and Word-based authoring workflows. Despite this diversity, all workflows converge on a common processing model governed by validation rules and state transitions tracked in an external workflow system. The paper also describes mechanisms for controlled self-service publishing, automated generation of derived content structures, and cross-system integration through identifier reconciliation. Operational reporting and provenance tracking further support auditability and continuous improvement. This work demonstrates how validation can serve not only as a quality assurance mechanism, but as a central organizing principle for workflow automation, enabling scalable and reliable content management across a distributed XML ecosystem.