Psychometric properties of the Armenian and Georgian versions of the PHQ-9, GAD-7, and WHO-5 well-being index
Abstract
Depressive and anxiety disorders are highly prevalent among healthcare workers. A survey was conducted in Armenia, Georgia, Moldova, and Ukraine to map their mental health as an extension of a previous study to improve the sustainability of health systems. The Patient Health Questionnaire (PHQ-9), the Generalized Anxiety Disorder scale (GAD-7), and the WHO 5-item well-being index (WHO-5) were used as screening tools. These were neither available nor validated in Armenian or Georgian. Assessment of their psychometric properties is needed to validate their applicability. We conducted a cross-sectional survey study. Translations were coordinated by the World Health Organization (WHO) Regional and Country Offices. A pilot study helped identify translation errors. The survey was open from May to September 2025. We used Cronbach’s α and McDonald’s ω to test for internal consistency, exploratory and confirmatory factor analyses for construct validity, Pearson’s correlation for discriminant validity, and Multiple-Group Confirmatory Factorial Analyses for measurement invariance. We analysed 3,364 valid responses from Armenian and Georgian doctors and nurses. Analyses indicated overall good internal consistency with Cronbach’s α and McDonald’s ω ranging between 0.8 and 0.9. The factor loadings ranged between 0.38 and 0.76, and the Comparative Fit Indices (CFI), Tucker-Lewis indices (TLI), Root-Mean-Square Error of Approximation (RMSEA), and Standardised Root-Mean-Squared Residual (SRMR) indices were appropriate even across measurement invariance tests. The strong positive correlation ( r = 0.75–0.78) between the PHQ-9 and GAD-7 scores suggested convergent validity, whereas a negative correlation ( r = − 0.59 to -0.51) between these scores and the WHO-5 score indicated divergent validity. The Armenian and Georgian versions of the PHQ-9, GAD-7, and WHO-5 showed generally acceptable reliability and partial support for construct validity within this specific sample of healthcare workers. Although these findings provide a foundation for future research, the reliance on a specific sample of healthcare professionals and the absence of validation against a diagnostic gold standard limit their immediate clinical applicability. Consequently, further validation in representative populations is required before these instruments can be broadly implemented in routine practice.