Skip to content
#software testing Open access

Utility of Bland-Altman Plot in the Assessment of Inter-observer Variability for Internal Quality Control in Semen Analysis: A Cross-sectional Study

Sep 2026 · Journal of Clinical and Diagnostic Research · 0 citations

TL;DR

In a low throughout laboratory, fresh semen is a useful sample for daily quality control and Bland-Altman plot is an easy and effective tool for monitoring internal quality control when there are two assessors, obviating the need for complex statistical tests.

Abstract

Introduction: Assessment of inter-observer variability is an essential component of quality control in semen analysis. The authors compared Bland-Altman (BA) plot, Intraclass correlation coefficient and Student’s paired t-test to determine which was the most feasible method for statistical analysis of internal quality control in their laboratory and the reason for the same. Aim: To assess inter-observer variability in sperm concentration and motility in fresh samples using Bland-Altman plot, Student’s paired t-test and Intraclass Correlation Coefficient (ICC), as a part of internal quality control. Materials and Methods: A cross-sectional observational study was conducted in the South of India, Puducherry, over a period of six months from 1st January 2020 to 30th June 2020. As a part of internal quality control two assessors independently analysed sperm concentration, progressive and non progressive motility and immotile spermson two aliquots of the samples tested by the manual method. Inter-observer variability was analysed using Bland-Altman plot, Students paired t-test and ICC. Data was analysed using Microsoft Excel® (2016), GraphPad Prism version 9, Mangold ICC calculator software and online ICC calculator at the website http://vassarstats.net/index.html Results: Nineteen men were included in the study. The ICC coefficient showed good correlation for sperm concentration, progressive motility, immotile sperms with values of 0.77, 0.84, 0.95 respectively and moderate correlation for non progressive motility with ICC=0.72. The p-value of Student’s paired t-test was above 0.05 for all parameters. There was no significant difference between the assessors for sperm concentration and motility using ICC and Student’s paired t-test, although the Student’s paired t-test does not measure agreement. BlandAltman plot showed an occasional outlier for both parameters. The authors attributed different pockets of sampling as the reason for outlier in sperm concentration. In motility testing, the two outliers were a result of delayed reporting by one of the assessors resulting in decline in the progressive motile and an increase in the non progressive motile and immotile sperms. Conclusion: In a low throughout laboratory, fresh semen is a useful sample for daily quality control. Bland-Altman plot is an easy and effective tool for monitoring internal quality control when there are two assessors, obviating the need for complex statistical tests. Adopting simple yet effective methods for internal quality control, would encourage more laboratories to come into the ambit of quality assurance.

Read PDF

Similar papers

Open access Aug 2026

Inter-Rater Reliability And Agreement Of The Diagnostic Assessment Scale For Kushtha In Papulosquamous Skin Disease: A Cross-Sectional Two-Rater Study

Background and objectives: The Diagnostic Assessment Scale for Kushtha (DASK) is a newly developed 26-item ordinal instrument that renders the classical threefold Ayurvedic examination as a severity score and a dosha attribution. No reliability estimate has been published. This study estimated its inter-rater reliability, agreement, and the measurement error attaching to an individual score. Methods: Sixty consecutive patients with consultant-confirmed papulosquamous disease (43 psoriasis, 17 lichen planus) were each assessed independently by two postgraduate-qualified Ayurvedic physicians in a single clinical session, giving a fully crossed design of 3,120 item scores. Reliability was estimated as the intraclass correlation coefficient ICC(2,1); agreement as the standard error of measurement (SEM), minimal detectable change (MDC95) and Bland–Altman limits; and item-level agreement as quadratic weighted kappa with Gwet’s AC2. Internal consistency and categorical agreement were also assessed, following the GRRAS guidelines. Results: No data were missing. ICC(2,1) was 0.952 (95% confidence interval 0.92–0.97) for the total score and 0.945, 0.920 and 0.920 for the Vata, Pitta and Kapha subscales. Between-patient variance accounted for 95.2% of total variance and the rater effect for 0.0–0.6%. SEM was 2.60 points and MDC95 7.21 points on the 0–104 scale, and all 26 items reached substantial or almost-perfect agreement (weighted kappa 0.732–0.963). The categorical outputs were less reliable than the scores generating them: severity band kappa 0.800 and dosha attribution kappa 0.640 (0.37–0.85), every disagreement occurring at a cut-point or a narrow margin. Kapha was reproduced consistently yet was not internally coherent (alpha 0.43 and 0.42). Conclusion: DASK measures the severity of papulosquamous Kushtha reliably enough for group comparison and, against a threshold of 8 points, for individual monitoring. Agreement on its categorical outputs was lower than on the continuous scores from which they are derived, and the Kapha subscale was reproduced consistently between raters without cohering internally.

A. M, N. S, Saritha T · 0 citations
Open access Aug 2026

Inter- and intra-rater reliability of linear scoring in warmblood breeding.

BACKGROUND Reliable phenotypes are essential for successful animal breeding programmes. Subjective phenotypic assessment may lead to under- or overestimation of heritability estimates and thereby affect the power of genome-based analyses. In Warmblood horses, conformation traits are commonly assessed using linear scoring, but its reliability has not been sufficiently quantified. AIMS/OBJECTIVES In this study, we investigated inter- and intra-rater reliability of linear scoring of conformation traits in German warmblood horses. METHODS Professional equine breeding experts (e.g. breeding judges) completed an online questionnaire in which they scored 17 linear conformation traits in 34 warmblood stallions using side-view images. The participants were asked to repeat the questionnaire after about one month. The images, sourced from the archive of the Hanoverian State Stud Celle, showed stallions born between 1929 and 1993. Reliability was assessed using intraclass correlation coefficients (ICCs). RESULTS Mean linear scores were close to zero. Low average absolute deviations indicated limited use of the -3 to +3 scale. Overall, inter-rater reliability between the 11 participants was lower than intra-rater reliability. Trait-specific patterns were identified across both analyses: The trait Head achieved moderate to good inter- and intra-rater agreement (inter-rater: Round 1 ICC = 0.51; 95% Confidence Interval (CI) [0.38, 0.65], Round 2 ICC = 0.50; 95% CI [0.37, 0.65], intra-rater: mean ICC = 0.71 ± 0.23), whereas traits such as Shoulder Length and Carpal Joints showed ICCs close to zero and CIs including zero. CONCLUSION Our findings imply the possible need for improved standardisation, training, or alternative phenotyping approaches to enhance reliability in equine conformation assessment.

A. Weigt, A. Brockmann, J. Tetens et al. · 0 citations
Review Open access Aug 2026

Interobserver and intraobserver variability of fetal and maternal Doppler measurements: systematic review and meta‐analysis

To evaluate the variability and reproducibility of umbilical artery (UA), fetal middle cerebral artery (MCA) and uterine artery (UtA) Doppler ultrasound measurements in pregnancy.

L. I. Prins, N. El Guili, C. Naaktgeboren et al. · 0 citations
Open access Jul 2026

Assessment of Analytical Quality in Clinical Laboratory Using Sigma Metrics

Background: Analytical quality in clinical laboratories is crucial for generating reliable test results that directly influence diagnosis and patient management. Traditional indicators, such as precision and accuracy, provide only partial assessment. Six sigma metrics offer a comprehensive, quantitative approach by integrating total allowable error (TEa), bias, and imprecision to evaluate the overall performance of analytical methods. Objectives: To assess the analytical performance of routine biochemical analytes using six sigma metrics and classify analytes according to sigma performance, and to identify analytes that require method improvement. Methods: A retrospective observational study was conducted at Biochemistry department DRPGMC, Tanda, Himachal Pradesh, India, using Internal Quality Control data and External Quality Assessment (EQA) results of 6 months from a clinical biochemistry laboratory. Imprecision (coefficient of variation [CV %]) was calculated from daily quality control (QC) data, and bias (%) was derived from EQA peer-group mean values. TEa% values were adopted from the Clinical Laboratory Improvement Amendments (CLIA) guidelines. Sigma metrics were calculated using the formula: σ = TEa ˗ ∣Bias∣ ÷ CV. Analytes were categorized into high (≥6σ), moderate (3–5.9σ), and low (<3σ) performance groups to guide QC rule selection. Results: Sigma metrics varied across analytes and required the TEa criteria applied. When assessed using the CLIA-88 TEa limits, triglycerides and high-density lipoprotein cholesterol (HDL-C) demonstrated high sigma performance (≥6σ), indicating excellent analytical precision. Moderate sigma performance (3–5.9σ) was observed for glucose, uric acid, alanine aminotransferase, aspartate aminotransferase, alkaline phosphatase, total protein, cholesterol (at level 3), and calcium (at level 3), necessitating multi-rule quality control strategies. In contrast, urea, creatinine, albumin, and phosphorus exhibited poor analytical performance with sigma values <3σ, indicating the need for improving the method. However, when sigma metrics were recalculated using the more stringent CLIA-2025 TEa limits, a further decline in analytical performance was observed. Uric acid, liver enzymes, and total protein demonstrated sigma values <3σ under the revised criteria, whereas triglycerides (at level 3) and HDL-C consistently maintained high sigma performance (≥6σ) despite the narrower allowable error limits. Conclusion: Six sigma assessments provided a comprehensive and quantitative measure of analytical quality in a clinical laboratory. Incorporating sigma metrics into routine quality assurance enhances reliability, optimizes QC protocols, and strengthens patient safety.

Anita Devi, N. Dogra, Mimosa Das · 0 citations
Open access Jul 2026

Reproducibility and inter-observer variability of the internal jugular vein ultrasonographic assessment: a multicenter cross-sectional study

The AP-IJV max and the CSA-IJV max measurements showed excellent inter-observer reproducibility, suggesting the use of point-of-care ultrasound protocols for non-invasive volume assessment, particularly at the neck’s base or cricoid level.

N. Parenti, E. Guidetti, Davide Allegri et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.