Skip to content
Review Open access

SP 9.02 Robustness of Retrospective Surgical Complication Grading for Time-Series Data Capture: An Inter-Observer Agreement Analysis Using Clavien-Dindo Classification

Aug 2026 · British Journal of Surgery · 0 citations

TL;DR

It is demonstrated that a supervised dual junior rater approach yields substantial-to-excellent inter-observer agreement in non-structural clinical data coding, however, agreement is likely over-estimated when only the highest grade is recorded over an extended interval (e.g. 30 days).

Abstract

Non-structural clinical data, such as post-operative complications, are susceptible to inter-observer disagreement, undermining data validity. This study assesses the validity of multi-timepoint Clavien-Dindo complication grading in a single-procedure cohort of Whipple resections. Complication grading within 7, 14, 30 and 90 days for 130 Whipple resections were independently coded by two final-year medical students and validated by a senior clinician. Inter-observer agreement was assessed using Cohen’s Kappa (κ). Disagreements were analysed by category (within minor, between major and minor, and within major grades). The senior clinician reviewed disagreements and identified potential systematic grading errors. Across 1040 coding episodes (130 patients, two coders, four time intervals) agreement results showed: κ (days 1-7) = 0.77(95% CI, 0.66-0.88); κ (days 8-14) = 0.91(0.85-0.97); κ (days 15-30) = 0.88(0.81-0.95); κ (days 31-90) = 0.73(0.59-0.88). Recording the highest complication grade within 30 days reduced disagreements from 32 to 15, κ (days 1-30) = 0.84(0.50-1). Most disagreements occurred within minor grades (I–II). Disagreement between minor and major grades (≤II vs ≥IIIa) was only observed in days 1-7. This study demonstrates that a supervised dual junior rater approach yields substantial-to-excellent inter-observer agreement in non-structural clinical data coding. However, agreement is likely over-estimated when only the highest grade is recorded over an extended interval (e.g. 30 days). For shorter time-series intervals, this approach is more prone to inconsistencies, highlighting the need for diagnostic criteria-based, automated data collection system to ensure data robustness for research and care quality improvement.

Read PDF

Similar papers

Jul 2026

Strong Correlation but Moderate Agreement: Comparison of Clavien-Dindo and Clavien-Madadi Classification Systems in Pediatric Percutaneous Nephrolithotomy.

Despite strong correlation, CD and CM classifications are not interchangeable and the CD system may overestimate complication severity in pediatric patients due to its anesthesia-based grading criteria, whereas the CM system appears more aligned with pediatric clinical practice.

Ibrahim Topcu, Resul Çiçek, Bulut Dural et al. · 0 citations
Open access Jul 2026

Clinical-Radiological Heterogeneity Within Intermediate Spinal Instability Neoplastic Scores (7-12): Factors Associated with Instrumented Stabilization in a Surgical Cohort.

How the intermediate Spinal Instability Neoplastic Score (SINS 7-12) category was operationalized in a real-world surgical spine oncology practice and identify preoperative factors associated with instrumented stabilization are described are described.

K. Krystkiewicz, Magdalena Orzechowska, Aleksander Kowal et al. · 0 citations
Review Aug 2026

Utility of the Modified Clavien-Dindo-Sink (mCDS) Grading System for Classifying Casting Complication Severity in Early Onset Scoliosis Patients.

The utility of the mCDS system for grading complications of EOS casting is assessed and it is hypothesized that, with modifications, it would be a valid system for assessing these complications.

Elinor Stern, Elizabeth Kappler, Makayla Hart et al. · 0 citations
Open access Jul 2026

Agreement and discordance between the modified thoracolumbar injury classification and severity score and the thoracolumbar AOSpine injury score in guiding surgical decision-making for thoracolumbar fractures: a comparative study.

For burst fractures with intervertebral disc involvement, mTLICS tends to recommend surgery more often, reflecting a difference in classification behaviour rather than a demonstrated clinical advantage, and mTLICS and TL AOSIS show substantial concordance as decision-support tools for thoracolumbar fractures.

Han Zhang, Junwei Feng, Hui-Bin Luo et al. · 0 citations
#small language model Open access Aug 2026

Evaluation of Small and Large Language Models for Calculation of the ASA Score and Charlson Comorbidity Index in Orthopedic Surgical Patients: A Retrospective Concordance Analysis

Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction.

Marco di Maio, G. Stopper, Vincenzo Di Matteo et al. · 0 citations