A Confound-Annotated Curriculum Dataset with Parsed Prerequisite Logic for the Universities of the United Arab Emirates
Abstract
A machine-readable, course-level corpus of the curricula of the universities of the United Arab Emirates, assembled from their published course catalogs. The corpus comprises twenty-two institutions across 56 catalog editions, 52,802 course records, and 33,945 parsed prerequisite relations, of which 7,030 (20.7%) are disjunctive alternatives rather than mandatory obligations. Its distinguishing property is that prerequisites are parsed into conjunctive-normal form, so that the alternatives a catalog states with the word "or" are preserved as boolean structure rather than flattened into a list of mandatory courses; the group index and alternative flag in the edge file recover the full conjunctive-normal structure. Every record is annotated with the measurement confounds that make document-derived curriculum data misleading if they are ignored, namely notation drift, selective disclosure, subject-code renumbering, and prerequisite-operator ambiguity, each exposed as a filterable field. One institution, the United Arab Emirates University, is covered by an eleven-edition panel spanning the decade from 2015-2016 to 2025-2026. All twenty-two institutions are represented at the course level. The verbatim course-description prose is not redistributed; its availability, language, and length are recorded in the course table, and its semantic content is provided as non-reproducing sentence embeddings. A pre-registered sampled correctness audit, included with the deposit, places course-code agreement at 100%, credit and prerequisite agreement in the mid-to-high nineties, and substantive title accuracy near 99%. The deposit includes the analysis code that computes curricular complexity under both the standard all-conjunctive reading and the alternative-aware reading, the integrity-verification script, and the full audit bundle.