Skip to content
#software testing Open access

Utility of Bland-Altman Plot in the Assessment of Inter-observer Variability for Internal Quality Control in Semen Analysis: A Cross-sectional Study

Sep 2026 · Journal of Clinical and Diagnostic Research · 0 citations

TL;DR

In a low throughout laboratory, fresh semen is a useful sample for daily quality control and Bland-Altman plot is an easy and effective tool for monitoring internal quality control when there are two assessors, obviating the need for complex statistical tests.

Abstract

Introduction: Assessment of inter-observer variability is an essential component of quality control in semen analysis. The authors compared Bland-Altman (BA) plot, Intraclass correlation coefficient and Student’s paired t-test to determine which was the most feasible method for statistical analysis of internal quality control in their laboratory and the reason for the same. Aim: To assess inter-observer variability in sperm concentration and motility in fresh samples using Bland-Altman plot, Student’s paired t-test and Intraclass Correlation Coefficient (ICC), as a part of internal quality control. Materials and Methods: A cross-sectional observational study was conducted in the South of India, Puducherry, over a period of six months from 1st January 2020 to 30th June 2020. As a part of internal quality control two assessors independently analysed sperm concentration, progressive and non progressive motility and immotile spermson two aliquots of the samples tested by the manual method. Inter-observer variability was analysed using Bland-Altman plot, Students paired t-test and ICC. Data was analysed using Microsoft Excel® (2016), GraphPad Prism version 9, Mangold ICC calculator software and online ICC calculator at the website http://vassarstats.net/index.html Results: Nineteen men were included in the study. The ICC coefficient showed good correlation for sperm concentration, progressive motility, immotile sperms with values of 0.77, 0.84, 0.95 respectively and moderate correlation for non progressive motility with ICC=0.72. The p-value of Student’s paired t-test was above 0.05 for all parameters. There was no significant difference between the assessors for sperm concentration and motility using ICC and Student’s paired t-test, although the Student’s paired t-test does not measure agreement. BlandAltman plot showed an occasional outlier for both parameters. The authors attributed different pockets of sampling as the reason for outlier in sperm concentration. In motility testing, the two outliers were a result of delayed reporting by one of the assessors resulting in decline in the progressive motile and an increase in the non progressive motile and immotile sperms. Conclusion: In a low throughout laboratory, fresh semen is a useful sample for daily quality control. Bland-Altman plot is an easy and effective tool for monitoring internal quality control when there are two assessors, obviating the need for complex statistical tests. Adopting simple yet effective methods for internal quality control, would encourage more laboratories to come into the ambit of quality assurance.

Read PDF

Similar papers

Open access Aug 2026

Inter-Rater Reliability And Agreement Of The Diagnostic Assessment Scale For Kushtha In Papulosquamous Skin Disease: A Cross-Sectional Two-Rater Study

Background and objectives: The Diagnostic Assessment Scale for Kushtha (DASK) is a newly developed 26-item ordinal instrument that renders the classical threefold Ayurvedic examination as a severity score and a dosha attribution. No reliability estimate has been published. This study estimated its inter-rater reliability, agreement, and the measurement error attaching to an individual score. Methods: Sixty consecutive patients with consultant-confirmed papulosquamous disease (43 psoriasis, 17 lichen planus) were each assessed independently by two postgraduate-qualified Ayurvedic physicians in a single clinical session, giving a fully crossed design of 3,120 item scores. Reliability was estimated as the intraclass correlation coefficient ICC(2,1); agreement as the standard error of measurement (SEM), minimal detectable change (MDC95) and Bland–Altman limits; and item-level agreement as quadratic weighted kappa with Gwet’s AC2. Internal consistency and categorical agreement were also assessed, following the GRRAS guidelines. Results: No data were missing. ICC(2,1) was 0.952 (95% confidence interval 0.92–0.97) for the total score and 0.945, 0.920 and 0.920 for the Vata, Pitta and Kapha subscales. Between-patient variance accounted for 95.2% of total variance and the rater effect for 0.0–0.6%. SEM was 2.60 points and MDC95 7.21 points on the 0–104 scale, and all 26 items reached substantial or almost-perfect agreement (weighted kappa 0.732–0.963). The categorical outputs were less reliable than the scores generating them: severity band kappa 0.800 and dosha attribution kappa 0.640 (0.37–0.85), every disagreement occurring at a cut-point or a narrow margin. Kapha was reproduced consistently yet was not internally coherent (alpha 0.43 and 0.42). Conclusion: DASK measures the severity of papulosquamous Kushtha reliably enough for group comparison and, against a threshold of 8 points, for individual monitoring. Agreement on its categorical outputs was lower than on the continuous scores from which they are derived, and the Kapha subscale was reproduced consistently between raters without cohering internally.

A. M, N. S, Saritha T · 0 citations
Open access Aug 2026

Inter- and intra-rater reliability of linear scoring in warmblood breeding.

BACKGROUND Reliable phenotypes are essential for successful animal breeding programmes. Subjective phenotypic assessment may lead to under- or overestimation of heritability estimates and thereby affect the power of genome-based analyses. In Warmblood horses, conformation traits are commonly assessed using linear scoring, but its reliability has not been sufficiently quantified. AIMS/OBJECTIVES In this study, we investigated inter- and intra-rater reliability of linear scoring of conformation traits in German warmblood horses. METHODS Professional equine breeding experts (e.g. breeding judges) completed an online questionnaire in which they scored 17 linear conformation traits in 34 warmblood stallions using side-view images. The participants were asked to repeat the questionnaire after about one month. The images, sourced from the archive of the Hanoverian State Stud Celle, showed stallions born between 1929 and 1993. Reliability was assessed using intraclass correlation coefficients (ICCs). RESULTS Mean linear scores were close to zero. Low average absolute deviations indicated limited use of the -3 to +3 scale. Overall, inter-rater reliability between the 11 participants was lower than intra-rater reliability. Trait-specific patterns were identified across both analyses: The trait Head achieved moderate to good inter- and intra-rater agreement (inter-rater: Round 1 ICC = 0.51; 95% Confidence Interval (CI) [0.38, 0.65], Round 2 ICC = 0.50; 95% CI [0.37, 0.65], intra-rater: mean ICC = 0.71 ± 0.23), whereas traits such as Shoulder Length and Carpal Joints showed ICCs close to zero and CIs including zero. CONCLUSION Our findings imply the possible need for improved standardisation, training, or alternative phenotyping approaches to enhance reliability in equine conformation assessment.

A. Weigt, A. Brockmann, J. Tetens et al. · 0 citations
Review Open access Aug 2026

Interobserver and intraobserver variability of fetal and maternal Doppler measurements: systematic review and meta-analysis.

OBJECTIVES To evaluate the variability and reproducibility of umbilical artery (UA), fetal middle cerebral artery (MCA) and uterine artery (UtA) Doppler ultrasound measurements in pregnancy. METHODS A systematic search of MEDLINE, EMBASE and the Cochrane Library (CENTRAL) was conducted from inception to 17 June 2024. Studies were included if they reported intra- or interobserver variability of UA, MCA or UtA Doppler ultrasound measurements in singleton or multiple pregnancies, using the intraclass correlation coefficient (ICC), concordance correlation coefficient, Cohen's kappa or limits of agreement. Data extraction was performed using standardized data-extraction forms. A meta-analysis was conducted using a random-effects model to account for heterogeneity, and variability estimates were pooled using Fisher's Z-transformation. Reproducibility was categorized according to the True Reproducibility of Ultrasound Techniques (TRUST) criteria. RESULTS In total, 2426 records were screened. Of these, 26 studies including 2457 patients met the inclusion criteria; eight studies described UA Doppler, 11 described MCA Doppler and 11 described UtA Doppler results; two of the studies described cerebroplacental ratio results. The pooled intraobserver ICC for UA pulsatility index (PI) was 0.88 (95% CI, 0.77-0.94) and the pooled interobserver ICC was 0.68 (95% CI, 0.55-0.79). For MCA-PI, these values were 0.85 (95% CI, 0.69-0.93) and 0.80 (95% CI, 0.62-0.90), respectively, and for mean UtA-PI they were 0.90 (95% CI, 0.85-0.93) and 0.84 (95% CI, 0.78-0.89), respectively. Findings across studies suggested generally poor-to-moderate reproducibility according to the strict TRUST criteria. CONCLUSIONS Our results emphasize that there is room for improvement regarding the reproducibility of Doppler ultrasound in routine obstetric practice. By standardizing ultrasound measurement protocols, advancing technologies such as artificial-intelligence-supported measurement applications and training programs, as well as obtaining multiple measurements per session and performing quality audits when outliers are detected, the field could move towards more accurate and reliable fetal assessments. © 2026 The Author(s). Ultrasound in Obstetrics & Gynecology published by John Wiley & Sons Ltd on behalf of International Society of Ultrasound in Obstetrics and Gynecology.

L. I. Prins, N. El Guili, C. Naaktgeboren et al. · 0 citations
Open access Jul 2026

Assessment of Analytical Quality in Clinical Laboratory Using Sigma Metrics

Background: Analytical quality in clinical laboratories is crucial for generating reliable test results that directly influence diagnosis and patient management. Traditional indicators, such as precision and accuracy, provide only partial assessment. Six sigma metrics offer a comprehensive, quantitative approach by integrating total allowable error (TEa), bias, and imprecision to evaluate the overall performance of analytical methods. Objectives: To assess the analytical performance of routine biochemical analytes using six sigma metrics and classify analytes according to sigma performance, and to identify analytes that require method improvement. Methods: A retrospective observational study was conducted at Biochemistry department DRPGMC, Tanda, Himachal Pradesh, India, using Internal Quality Control data and External Quality Assessment (EQA) results of 6 months from a clinical biochemistry laboratory. Imprecision (coefficient of variation [CV %]) was calculated from daily quality control (QC) data, and bias (%) was derived from EQA peer-group mean values. TEa% values were adopted from the Clinical Laboratory Improvement Amendments (CLIA) guidelines. Sigma metrics were calculated using the formula: σ = TEa ˗ ∣Bias∣ ÷ CV. Analytes were categorized into high (≥6σ), moderate (3–5.9σ), and low (<3σ) performance groups to guide QC rule selection. Results: Sigma metrics varied across analytes and required the TEa criteria applied. When assessed using the CLIA-88 TEa limits, triglycerides and high-density lipoprotein cholesterol (HDL-C) demonstrated high sigma performance (≥6σ), indicating excellent analytical precision. Moderate sigma performance (3–5.9σ) was observed for glucose, uric acid, alanine aminotransferase, aspartate aminotransferase, alkaline phosphatase, total protein, cholesterol (at level 3), and calcium (at level 3), necessitating multi-rule quality control strategies. In contrast, urea, creatinine, albumin, and phosphorus exhibited poor analytical performance with sigma values <3σ, indicating the need for improving the method. However, when sigma metrics were recalculated using the more stringent CLIA-2025 TEa limits, a further decline in analytical performance was observed. Uric acid, liver enzymes, and total protein demonstrated sigma values <3σ under the revised criteria, whereas triglycerides (at level 3) and HDL-C consistently maintained high sigma performance (≥6σ) despite the narrower allowable error limits. Conclusion: Six sigma assessments provided a comprehensive and quantitative measure of analytical quality in a clinical laboratory. Incorporating sigma metrics into routine quality assurance enhances reliability, optimizes QC protocols, and strengthens patient safety.

Anita Devi, N. Dogra, Mimosa Das · 0 citations
Open access Aug 2026

Evaluating Systematic and Proportional Bias in Point-of-care Glucose Testing: A Correlation and Bland-Altman Analysis

Point-of-care testing (POCT) for blood glucose provides rapid results and is widely used in clinical practice; however, concerns remain regarding its accuracy and agreement with laboratory auto-analyser methods. The purpose of this study was to compare glucose measurements obtained using POCT and a laboratory auto-analyser in a secondary healthcare centre within a low-resource setting. A comparison study was conducted using 120 paired blood glucose measurements obtained simultaneously by a POCT glucometer and a laboratory auto-analyser. Descriptive statistics were used to summarise glucose values. Pearson correlation and linear regression analyses assessed the relationship between methods, while agreement was evaluated using Bland-Altman analysis. Statistical significance was set at P < .05. The mean glucose concentration measured by POCT was 6.59 ± 2.27 mmol/L, while that measured by the auto-analyser was 6.33 ± 2.93 mmol/L. A strong positive correlation was observed between the two methods ( r = 0.971, P < .001). Linear regression analysis yielded the equation auto-analyser = 1.249 × POCT − 1.905 ( R ² = 0.943), indicating proportional bias. Bland-Altman analysis demonstrated a mean bias of −0.26 mmol/L, with 95% limits of agreement ranging from −2.03 to +1.50 mmol/L. Greater variability in differences was observed at higher glucose concentrations. POCT glucose measurements show excellent correlation with laboratory auto-analyser results but exhibit systematic and proportional bias, particularly at higher glucose levels. While POCT is suitable for rapid glucose assessment and monitoring, laboratory auto-analyser methods remain the preferred reference for diagnostic and critical clinical decision-making.

Ufuoma Ohwo, K. Digban, C. Okafor · 0 citations
Open access Jul 2026

Reproducibility and inter-observer variability of the internal jugular vein ultrasonographic assessment: a multicenter cross-sectional study

This study evaluates inter-observer variability in measuring internal jugular vein (IJV) ultrasonographic parameters using images obtained from medical ward patients. A cross-sectional study was conducted across 4 Italian hospitals. After brief training, 8 expert sonographers and 11 novices measured at the supraclavicular, cricoid, and submandibular levels, the anteroposterior expiratory maximum diameter (AP-IJV max), maximum cross-sectional area (CSA-IJV max), aspect ratio (AP-IJV max to latero-lateral diameter), and collapsibility index (IJV-c) on 60 IJV ultrasound images from 10 patients. Inter-observer reliability was assessed using the intraclass correlation coefficient. No significant differences were observed between experts and novices for AP-IJV max, CSA-IJV max, or aspect ratio across the different views. The lowest inter-examiner variability was found for AP-IJV max and CSA-IJV max at the supraclavicular level. Aspect ratio showed moderate reliability among experts but poor reproducibility among novices. The IJV-c demonstrated the highest inter-examiner variability in both groups. To our knowledge, this is the first multicenter study aimed at evaluating the method's reliability of different non-invasive acoustic measurements and windows to identify IJV dimensions better to use mainly in case of hypovolemia or fluid challenge. The AP-IJV max and the CSA-IJV max measurements showed excellent inter-observer reproducibility, suggesting the use of point-of-care ultrasound protocols for non-invasive volume assessment, particularly at the neck’s base or cricoid level.

N. Parenti, E. Guidetti, Davide Allegri et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.