Predicting Pharmacy Closures in Uganda Using Machine Learning a Risk-Based Approach to Regulatory Prioritization
No Thumbnail Available
Date
2026-09-25
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Uganda Christian University
Abstract
Pharmacy closures can disrupt access to medicines and may indicate underlying regulatory or operational weaknesses within Uganda’s pharmaceutical sector. Although the National Drug Authority routinely collects pharmacy licensing and inspection data, these data are largely used for reactive oversight rather than proactive prediction of closure risk.
This study identified factors associated with pharmacy closure and empirically evaluated supervised machine-learning models for predicting pharmacy closure in Uganda to support risk-based regulatory prioritization. A retrospective analytical design was used based on secondary regulatory data from licensed retail and wholesale pharmacies registered between 2016 and 2024. The analytical dataset comprised 19,636 inspection records from 5,149 unique pharmacies. Descriptive analysis characterized closure patterns by time, geography, license status, inspection outcomes, and pharmacy category. Exploratory analysis examined regulatory and inspection variables associated with closure, while Logistic Regression, Decision Tree, Random Forest, and Extreme Gradient Boosting (XGBoost) models were evaluated for predicting closure within 180 days after inspection.
Model development used a chronological, non-random train/test split, with the oldest 80% of inspection records used for training and the most recent 20% retained as an independent test set. Group-aware cross-validation was applied within the training data using pharmacy identifier to reduce information leakage arising from repeated inspections. Model performance was assessed using ROC-AUC, PR-AUC, precision, recall, F1-score, and Precision@k.
XGBoost achieved the strongest group-aware cross-validation performance, with a ROC-AUC of 0.989 and PR-AUC of 0.948. On the independent chronological test set, Random Forest achieved the highest ROC-AUC (0.907), while XGBoost achieved the highest PR-AUC (0.214), with corresponding ROC-AUC and PR-AUC values of 0.902 and 0.214, respectively. XGBoost also achieved the highest recall (98.0%) and a joint-highest F1-score (0.375) at the 0.5 classification threshold.
The findings demonstrate that routinely collected regulatory data contain predictive signals for pharmacy closure, although model performance varied across evaluation criteria and between cross-validation and chronological testing. Inspection history, license-related variables, and geographic characteristics contributed important predictive information. Overall, XGBoost provided the strongest combination of cross-validation and independent test performance and was selected as the preferred model for risk prediction. The findings suggest that machine-learning-derived risk scores could support targeted inspection planning and proactive regulatory prioritization, while model outputs should remain subject to professional regulatory judgement. Prospective validation, improved data quality, calibration, fairness assessment, and monitoring for model drift are recommended before operational deployment.
