Master of Science in Data Science and Analytics

Permanent URI for this collectionhttps://hdl.handle.net/20.500.11951/1204

Browse

Recent Submissions

Now showing 1 - 20 of 30
  • Item
    A Multi-Class Machine Learning Customer Classification For Personalised Voice Bundle Recommendations: A Case Study of Airtel Uganda Limited
    (Uganda Christian University, 2026-09-14) Bisimbeko, Remmy
    Airtel Uganda Limited relies on Average Revenue Per User (ARPU)-banded segmentation to deliver voice bundle promotions, yielding conversion rates below 3 percent and wasting an estimated 20 percent of the marketing budget on untargeted communications. The absence of a personalised, data-driven recommendation system represents both a commercial ine"ciency and an academic gap this study addresses. This study developed and validated a machine learning recommendation framework using real transactional data from 1,458,900 Airtel Uganda prepaid subscribers (February–April 2026). Four classification models — Logistic Regression, Decision Tree, Random Forest, and XGBoost — were evaluated following a preprocessing pipeline of median imputation, capped SMOTE oversampling, and behavioural feature engineering. Three statistical tests were applied to evaluate relationships between customer attributes and bundle segment: Spearman Rank Correlation to measure the strength and direction of monotonic associations between numeric features and the target; one-way ANOVA to test whether feature means di!er significantly across the eight value segment groups; and Chi-Square to assess the independence of categorical features from the target variable. Statistical tests confirmed 34 of 35 features significantly associated with bundle segment (p < 0.05); UPGRADE_RATE was the only non-significant feature. Feature importance analysis identified TOTAL_REVENUE (18.01%), SPEND_PER_MINUTE (6.57%), CALLS (5.97%), VOICE_INTENSITY (5.87%), and MOBILE_MONEY (4.64%) as the top five predictors. XGBoost achieved the highest F1-Score of 0.661 on the 8-class hold-out test set — a 5.4-fold improvement over a random baseline. A simulated pilot projected a four- to five-fold improvement in conversion rate (3% to 12–15%) and a 9 percent ARPU uplift (simulated pilot projection). Three strategic recommendations were formulated: a VOICE_INTENSITY-triggered recommendation engine, enriched segmentation incorporating device type and purchase channel, and a formal A/B test with model governance. This is the first empirically validated ML bundle recommendation framework built on real Ugandan telecom data. Keywords: Machine Learning, Telecommunications, Voice bundle recommendation, Customer behavioural attributes, Predictive Analytics.
  • Item
    Multi-horizon predictive modeling of HIV treatment adherence using clinical and social determinants in Uganda
    (Uganda Christian University, 2026-06-05) Jacob Nyonyintono
    Treatment interruption among people living with HIV (PLHIV) remains a critical challenge to achieving sustained viral suppression and optimal treatment outcomes, particularly in sub- Saharan Africa. In Uganda, despite significant progress in scaling up antiretroviral therapy (ART), a substantial proportion of patients experience interruptions in treatment, contributing to viral rebound, increased transmission risk, and drug resistance. The adoption of the Multi- Month Dispensing (MMD) model, while improving access and convenience, reduces routine patient-provider contact and may delay early identification of patients at risk of disengagement. This study aimed to develop and evaluate a machine learning-based predictive framework for identifying patients at risk of treatment interruption under Uganda’s MMD system. A quantitative research design was employed using retrospective data extracted from Electronic Medical Records (EMRs) of 8,788 patients receiving ART. Data preprocessing involved cleaning, feature engineering, and handling missing values, followed by the development of machine learning models using a structured pipeline. The Random Forest algorithm was selected as the primary model due to its ability to capture complex nonlinear relationships and its robustness in handling imbalanced clinical datasets. The models were trained and evaluated across three temporal prediction windows (30, 60, and 90 days) using performance metrics including precision, recall, F1-score, and ROC-AUC. Particular emphasis was placed on precision to ensure reliable identification of high-risk patients while minimizing false positive classifications. The findings demonstrate that integrating clinical, demographic, and behavioral factors improves the predictive ability of machine learning models in identifying treatment interruption risk. Key predictors included viral load status, CD4 count, duration on ART, and selected socio-demographic characteristics. The developed framework enables the generation of individualized risk scores and prediction of likely interruption periods, supporting proactive patient management. This study contributes to the growing body of evidence on the application of machine learning in HIV care and highlights the importance of incorporating multidimensional determinants in predictive modelling. The proposed model provides a practical tool for early risk identification and targeted intervention, with potential to enhance patient retention and improve treatment outcomes within Uganda’s HIV care system.
  • Item
    Machine Learning Decomposition of Gender and Rural–Urban Wage Gaps: Evidence from Uganda National Household Surveys
    (2026-09-30) Tumushiime, Bob Robert
    Uganda's gender wage gap stands at about 32 percent (UN Women, 2024), with rural–urban differences of similar size. Published decomposition studies for Uganda have relied largely on the 2002/03 survey round and a linear wage equation that assumes constant returns to characteristics. This study examines whether a flexible functional form, applied to recent multi-round data, alters the share of each gap attributed to measured characteristics. Four Uganda National Household Survey rounds (2012/13–2023/24) were harmonised into a pooled sample of 14,963 formal wage earners aged 14–64. Ordinary least squares (OLS), LASSO, Random Forest and XGBoost were trained on a stratified 70/15/15 split, and SHapley Additive exPlanations (SHAP) were used to attribute predictions to worker characteristics. Predictions from the best-performing model replaced the linear fitted values in an Oaxaca–Blinder decomposition estimated for each round. XGBoost achieved the highest test R² (0.410), followed by Random Forest (0.399) and the two linear models (0.354); pooling all four rounds raised it from 0.286. Education was the leading predictor. The rural–urban gap was largely explained under both methods: 49–59 percent under OLS and 68–84 percent under the machine-learning decomposition. For the gender gap, OLS produced a negative explained share in three of four rounds, while the machine-learning decomposition attributed 14–57 percent to measured characteristics. The explained share of the gender gap is thus sensitive to the functional form of the wage model, which calls for caution in interpreting single-round linear estimates. The approach can be reproduced on successive survey rounds to track wage inequality over time. The unexplained component is reported as a residual and is not interpreted as a measure of discrimination.
  • Item
    Interpretable Predictive Modeling of Postpartum Surgical Site Infections
    (Uganda Christian University, 2026-07-11) Emmanuel Nkurunziza
    Background: Postpartum surgical site infections (SSIs) are a major cause of maternal morbidity in low-resource settings, yet early detection is limited by reactive, symptom-based surveillance systems. Ma chine learning (ML) and explainable artificial intelligence (XAI) offer potential tools for proac tive SSI risk identification using routinely collected clinical data. Methods: A retrospective dataset of 8,918 caesarean deliveries from three regional referral hospitals in Uganda (SSI prevalence: 8.1%) was analysed. After rigorous cleaning, leakage-free feature engineering, imputation, and encoding, four ML models Logistic Regression, Random Forest, Support Vector Machine (RBF), and XGBoost were trained and evaluated under extreme class imbalance using SMOTENC. Model performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and precision–recall curves. SHAP explainability was applied to identify globally and locally influential predictors. Results: Overall discriminatory performance was modest due to limited predictive signal in routine clin ical records. Logistic Regression achieved the highest sensitivity (recall = 0.9931; ROC-AUC = 0.6528), whereas XGBoost produced the highest accuracy (0.7932) but poor recall (0.2639). SHAP analysis highlighted preoperative showering, skin antiseptic preparation, obstructed labour, and prolonged surgery duration as key contributors to increased SSI risk. Conclusion: Explainable ML models are feasible for early SSI risk identification in low-resource settings but are constrained by sparse routine data. Logistic Regression may serve as a high-sensitivity clinical screening tool. Improved performance will require enriched clinical datasets, inclu sion of intraoperative variables, and multi-centre validation to support practical deployment in maternal health surveillance.
  • Item
    A Machine Learning Framework for Flood Risk Classification and Early warning in River Catchment Areas
    (Uganda Christian University, 2026-09-29) Martha Frances Namakula
    Flooding remains one of the most destructive natural hazards in Uganda, with the River Man afwa catchment area on the slopes of Mount Elgon experiencing recurrent and severe flood events that displace communities, destroy infrastructure, and disrupt livelihoods. Existing flood forecasting in Uganda relies heavily on global systems such as GloFAS and ECMWF, which operate at low spatial resolution and are often unreliable for small catchments, limiting their effectiveness for local early warning and disaster preparedness. This study developed a machine learning framework for flood risk classification in the River Manafwa catchment, using historical rainfall data (CHIRPS satellite, 1989–2024) and observed river discharge data (Min istry of Water and Environment, 1989–2024). Hydrological features were engineered, including antecedent rainfall accumulations (1, 3, 5, and 7-day windows), lagged and rolling-mean dis charge variables, discharge differencing and acceleration terms, and seasonal cyclical indicators. A locally calibrated four-level flood risk classification (Normal, Mild, Advanced, Extreme) was developed based on discharge percentile thresholds (Q75, Q90, and Q95). Five machine learn ing models, Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, and MLP Neural Network, were trained and evaluated using Macro F1-score, Extreme F1-scores, and AUC-ROC, with SMOTE applied to address class imbalance in the training data. Gradient Boosting achieved the strongest overall performance, recording the highest Macro F1-score and Extreme-class F1-score following SMOTE application, and was selected as the most suitable model for this study. SHAP analysis identified short-term lagged discharge, particularly the previous day’s discharge, as the dominant indicator of flood risk across all classes, while same day rainfall and short-term cumulative rainfall accumulations retained comparatively greater influence specifically within the Extreme risk class. The model was validated against three documented historical flood events (June 2019, October 2019, and September 2021), with the model successfully flagging Extreme risk conditions several days ahead of officially reported flood dates. The resulting framework offers a low-cost, replicable, data-driven approach to strengthen flood early warning systems for the Manafwa catchment and other flood-prone rivers in Uganda, supporting disaster preparedness agencies including URCS, UNMA, and the OPM’s Department of Relief and Disaster Preparedness.
  • Item
    Provenance-aware sentiment analysis and multi-criteria influencer selection: a decision-support framework demonstrated for Uganda Airlines
    (Uganda Christain University, 2026-10-06) Jordan Micheal Senyondo
    Brands increasingly rely on creators to reach audiences, yet follower scale alone does not represent brand fit, audience relevance, engagement quality, or the emotional context of public discourse. This thesis develops and demonstrates a provenance-aware decision- support framework that combines descriptive sentiment context with evidence-aware, multi-criteria influencer ranking. Uganda Airlines provides the management scenario, while the empirical baseline is an observed legacy aviation-discourse dataset rather than a complete record of Uganda Airlines’ own social-media activity. After tweet-ID deduplication, the observed baseline contains 14,485 unique tweets collected from 17 to 24 February 2015. Of these records, 62.70% are negative, 21.19% neutral, and 16.11% positive. A separate scenario layer contains 125,200 simulated records covering 2015–2026. The resulting 139,685-record union is therefore described as simulation-augmented; its longitudinal patterns are scenario results and not indepen- dently observed trends. The influencer campaign panel contains 683 observed and 60,000 simulated rows across 95 name–platform profiles representing 37 distinct influencer names. Sixty-four profiles meet the default eligibility threshold of at least five observed posts. Influencer selection is operationalised through the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS). The ranking criteria are value alignment, audience fit, sentiment context, engagement quality, reach, and cost e"ciency. The integrated top ten has no profiles in common with a reach-only top ten and two in common with a follower-count-only top ten, showing that the implemented ordering di!ers from scale-only selection. Criterion ablations identify engagement quality and value alignment as the most influential preferences in the current ordering. These are deterministic ranking results, not campaign-e!ect estimates. The simulated layer is reserved for stress testing rather than increasing empirical confidence. The thesis contributes (i) a reusable, values-aware framework for linking social listening to influencer selection; (ii) a transparent method for separating observed evidence from simulation-based scenario analysis; and (iii) a working dashboard that exposes provenance, evidence thresholds, criterion contributions, and ranking sensitivity. The findings support managerial exploration and monitoring, but do not establish causal campaign uplift or externally validated real-world predictive accuracy. Keywords: social listening; sentiment analysis; simulation augmentation; brand–creator alignment; influencer selection; TOPSIS; decision-support dashboard.
  • Item
    Enhancing food security forecasting in Uganda: a geospatial Analytics framework with crop yield integration
    (Uganda Christian University, 2026-09-30) Arinaitwe, Philip
    Food security is a crucial pillar for human well being, societal stability and national development. Food security can be defined as a situation when all people, at all times, have physical and economic access to sufficient, safe and nutritious food that meets their dietary needs and food preferences for an active and healthy life. In Uganda, nearly 20 percent of the population lives below the poverty line and over 16 million Ugandans face food insufficiency due to various factors like unreliable climate, economic instability, population pressure, conflict and inadequate forecasting mechanisms. Recent research studies have increasingly turned to predictive modelling using machine learning for food security forecasting, integration of climate data, socio-economic indicators and innovative geospatial analytics. Whereas advances in machine learning and geospatial analytics have led to significant improvements in climatic modelling for food security, they often fall short by excluding actual agricultural outputs. Without integrating crop yield data, forecasts may misestimate food availability yet it can be a key driver of food crises. Particularly for Uganda, where over 70% of the population is employed by agriculture, the omission of crop yield data from forecasting models can result in inaccurate predictions/forecasts. To address this gap, this study developed and evaluated a geospatial machine learning framework for forecasting food security in Uganda by integrating crop yield data from FAOSTAT with multi-source environmental, market, conflict, demographic and historical Integrated Food Security Phase Classification (IPC) data from 2007 to 2020. A hybrid model combining a Long Short-Term Memory (LSTM) temporal encoder and an XGBoost classifier was trained and validated on 44, 082 district-month observations. The hybrid LSTM–XGBoost model achieved a macro F1-score of 0.75 and an overall accuracy of 93.8% on the test set. Importantly, the model demonstrated exceptional performance in identifying severe food crises (IPC 3+), achieving a recall of 0.82, a ROC-AUC of 0.998 and a PR-AUC of 0.852. An ablation study confirmed that incorporating crop yield data significantly improved predictive power, increasing the IPC 3+ F1-score from 0.688 to 0.759. Furthermore, model comparisons demonstrated that a country-specific model trained purely on Ugandan observations significantly outperformed a consolidated regional model shown by a macro F1 score of 0.747 achieved by the country-specific model as compared to the consolidated regional model with a macro F1-score of 0.725. These results provide robust empirical evidence that integrating crop yield data into geospatial frameworks enhances food security early warning systems which enables timely, targeted interventions to mitigate severe hunger.
  • Item
    A Predictive Analytics Framework for Case Backlog Management in Uganda’s Judiciary: An Explainable Machine Learning Approach
    (Uganda Christian University, 2026-10-02) Bbossa, Isaac Sserunkuma
    The Uganda judicial system is faced with a persistent and rising backlog of cases, whereby 26.32% of cases have remained pending for more than two years, negatively impacting access to justice and socio-economic development. Current systems such as the Electronic Court Case Management Information System (ECCMIS) can only make use of descriptive reporting techniques, thus imple- menting a reactive administrative paradigm. This research bridges the critical gap in the area of predictive capability through developing an explainable machine learning model to preemptively pre- dict the backlog risk of cases. Adopting the Design Science Research method, the research adopted a quantitative experimental design. Making use of secondary data from the National Court Case Census (2025), the research included substantial data pre-processing, feature engineering, and com- paring various machine learning classifiers including Logistic Regression, Random Forest, Gradient Boosting, and K-Nearest Neighbors to carry out binary classification task of backlog prediction. The best performing classifier was Random Forest with an ROC AUC of 0.885 and F1 Score of 0.698, significantly outperforming other classifiers. The explainable AI (XAI) analysis indicated that Claim/offence (case complexity), Court name (institutional capacity), and case status (procedural stage) were the most predictive features. This further confirms the hypothesis that backlog risk is essentially due to the legal and institutional dynamics rather than merely time dynamics. The analysis succeeded in developing an empirical predictive model that proved the potency of ensemble models such as Random Forest in capturing the complex non-linear dynamics of judicial backlog in Uganda. The results allow shifting from reactive to proactive case management strategy. This thesis ends up with some practical recommendations regarding the implementation of a predictive warning system, dynamic resource planning, and targeted procedural reform.
  • Item
    The Effect of Exchange Rate Dynamics on Food Prices in Urban Uganda: An Integrated Machine Learning and Econometrics Analysis of Kampala Markets
    (Uganda Christian University, 2026-09-15) Econia, Racheal
    Food price volatility remains a major policy concern in urban Uganda because changes in staple food prices directly affect household welfare, food security, and the purchasing power of lowincome consumers. This study investigated the effect of UGX/USD exchange rate dynamics on food prices in Kampala markets using a multi-method analytical framework that combines econometric analysis, causal inference, non-linear modelling, frequency-domain analysis, machine learning, and model deployment. The study used World Food Programme food-price data for Uganda and official Bank of Uganda exchange-rate data, merged on a monthly basis to create an analytical panel covering January 2006 to September 2025. The analysis began with data preprocessing and exploratory data analysis to examine foodprice trends, commodity-level volatility, and consistency between implied and official exchange rates. Causal and dynamic relationships were then evaluated using regression analysis, Granger causality testing, transfer entropy, DCC-GARCH, Markov-switching regression, quantile regression, and wavelet coherence. Predictive performance was assessed using traditional regression models, Random Forest, XGBoost, Prophet, and deep learning models, with model accuracy evaluated using R2, RMSE, and MAE. Finally, selected models were operationalized through a Flask API and Progressive Web Application frontend to demonstrate practical food-price prediction. The findings show that exchange rate movements contain meaningful information for understanding and forecasting urban food prices, but the relationship is non-linear, commodityspecific, and time-varying. Causal evidence indicates a predominantly one-directional influence from exchange rates to food prices, while regime and quantile results show that pass-through effects intensify during high-volatility and high-price episodes. Wavelet analysis further shows strong short-term co-movement at the 3–6 month horizon, with weaker associations at longer horizons. In predictive modelling, XGBoost achieved the strongest panel-level performance with an R2 of 0.7944, RMSE of 646.25, and MAE of 401.60, followed closely by Random Forest with an R2 of 0.7815. These results demonstrate that machine learning models using exchangerate and commodity information can substantially improve food-price prediction compared with linear baselines. The study contributes theoretically by extending exchange rate pass-through analysis to urban food markets, methodologically by integrating causal, non-linear, time-frequency, and predictive approaches, and practically by deploying a web-based decision-support tool for commodity price prediction. The findings suggest that food-price stabilization policies in Uganda should combine exchange-rate monitoring with commodity-specific market intelligence and early warning systems.
  • Item
    Machine Learning Forecasting for Electricity Demand Using Climate and Socioeconomic Data in Sub-Saharan Africa
    (Uganda Christian University, 2026-09-30) Lutalo, Lordin
    Electricity demand forecasting is important for energy planning, resource allocation, and infrastructure development. In Sub-Saharan Africa, forecasting electricity demand is particularly difficult because electricity usage is driven by rapid population growth, economic changes, urbanisation, and climate variability, while data quality and availability remain inconsistent across countries. Although previous studies showed that climate and socioeconomic variables are associated with electricity demand, many focused on short-term forecasting, single-country analysis, or datasets from developed regions. This study examined how far integrated climate and socioeconomic data improves annual electricity demand forecasting across 47 Sub-Saharan African countries. Country-level panel time-series data was collected from publicly available sources including the World Bank, Our World in Data, and the ERA5 climate reanalysis dataset, covering the period from 2000 to 2021. The methodology included data integration and preprocessing, normality testing, temporal feature engineering, a two-stage feature selection process, and the training and comparison of six forecasting models: Linear Regression, Fixed Effects Panel Regression, Dynamic Fixed Effects Panel Regression, Random Forest, XGBoost, and LightGBM. Model performance was evaluated using Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), Mean Absolute Percentage Error (MAPE), and the coefficient of determination (R^2). A decision-support dashboard was also developed using Streamlit to visualise electricity demand forecasts, trends, and key demand drivers. The results showed that lagged electricity demand was by far the strongest predictor of future demand, with the one-year lag achieving a near-perfect Pearson correlation of 0.999 with the target variable. Among structural climate and socioeconomic predictors, population size and its three-year rolling mean were the strongest predictors of electricity demand, followed by the three-year rolling mean of temperature, the interaction between population and electricity access, and the temperature-urbanisation interaction. Linear Regression on the full multivariate feature set achieved the best overall test set performance, with a log-scale MAPE of 6.01% and an R^2 of 0.996 (note: all metrics are on the log-transformed demand scale, not in original TWh units), and was selected as the final model for the dashboard. The study confirms that annual electricity demand in Sub-Saharan Africa is extremely persistent, lagged demand provides most of the predictive power in the model. Climate and socioeconomic variables contain substantial structural information about demand patterns when historical demand is unavailable, but contribute little additional one-year-ahead predictive accuracy once lagged demand is already included. This distinction between predictive accuracy and structural explanation is the core finding of the study.
  • Item
    Predicting Pharmacy Closures in Uganda Using Machine Learning a Risk-Based Approach to Regulatory Prioritization
    (Uganda Christian University, 2026-09-25) Fiona Mbabazi
    Pharmacy closures can disrupt access to medicines and may indicate underlying regulatory or operational weaknesses within Uganda’s pharmaceutical sector. Although the National Drug Authority routinely collects pharmacy licensing and inspection data, these data are largely used for reactive oversight rather than proactive prediction of closure risk. This study identified factors associated with pharmacy closure and empirically evaluated supervised machine-learning models for predicting pharmacy closure in Uganda to support risk-based regulatory prioritization. A retrospective analytical design was used based on secondary regulatory data from licensed retail and wholesale pharmacies registered between 2016 and 2024. The analytical dataset comprised 19,636 inspection records from 5,149 unique pharmacies. Descriptive analysis characterized closure patterns by time, geography, license status, inspection outcomes, and pharmacy category. Exploratory analysis examined regulatory and inspection variables associated with closure, while Logistic Regression, Decision Tree, Random Forest, and Extreme Gradient Boosting (XGBoost) models were evaluated for predicting closure within 180 days after inspection. Model development used a chronological, non-random train/test split, with the oldest 80% of inspection records used for training and the most recent 20% retained as an independent test set. Group-aware cross-validation was applied within the training data using pharmacy identifier to reduce information leakage arising from repeated inspections. Model performance was assessed using ROC-AUC, PR-AUC, precision, recall, F1-score, and Precision@k. XGBoost achieved the strongest group-aware cross-validation performance, with a ROC-AUC of 0.989 and PR-AUC of 0.948. On the independent chronological test set, Random Forest achieved the highest ROC-AUC (0.907), while XGBoost achieved the highest PR-AUC (0.214), with corresponding ROC-AUC and PR-AUC values of 0.902 and 0.214, respectively. XGBoost also achieved the highest recall (98.0%) and a joint-highest F1-score (0.375) at the 0.5 classification threshold. The findings demonstrate that routinely collected regulatory data contain predictive signals for pharmacy closure, although model performance varied across evaluation criteria and between cross-validation and chronological testing. Inspection history, license-related variables, and geographic characteristics contributed important predictive information. Overall, XGBoost provided the strongest combination of cross-validation and independent test performance and was selected as the preferred model for risk prediction. The findings suggest that machine-learning-derived risk scores could support targeted inspection planning and proactive regulatory prioritization, while model outputs should remain subject to professional regulatory judgement. Prospective validation, improved data quality, calibration, fairness assessment, and monitoring for model drift are recommended before operational deployment.
  • Item
    An Explainable Transformer-Based Vision-Language Model For Multimodal Consistency Verification
    (none, 2026-09-23) Mugimba Kakure Jude
    The rapid growth of multimodal data has underscored the necessity for robust methods to verify semantic alignment between images and their associated textual descriptions. This problem, termed multimodal consistency verification, asks whether a given image and text pair are semantically compatible or contradictory. Unlike conventional image–text retrieval, which ranks many candidates for a query, consistency verification is a pairwise decision on a single pair. While large-scale vision-language models demonstrate strong performance in retrieval and zero-shot tasks, explicit consistency verification, particularly with integrated explainability, remains underexplored. We present an explainable dual-encoder vision-language architecture that utilizes a Vision Transformer (ViT) and Bidirectional Encoder Representations from Transformers (BERT) to encode images and multi-caption documents, respectively. These representations are projected into a shared embedding space and aligned via a symmetric InfoNCE contrastive objective. We quantify consistency using cosine similarity, employing a decision threshold to distinguish between consistent and inconsistent pairs. The model incorporates visual attention maps from the ViT and token importance scores from the BERT encoder that provide qualitative evidence for its classification decisions. For evaluation, a proxy benchmark dataset was constructed from the Flickr30K dataset by aggregating the five captions associated with each image into a single multi-caption document and generating consistent, randomly mismatched, and Facebook AI Similarity Search (FAISS) retrieved hard-negative pairs. The inconsistent pairs were generated algorithmically rather than through extensive human annotation of semantic contradiction type. The model was tested across balanced, imbalanced, and hard-negative scenarios and compared with Contrastive Language–Image Pre-training (CLIP) as a baseline in all scenarios. To isolate the contributions of the transformer-based encoders, the Vision Transformer was replaced by Residual Network-50 (ResNet-50) to evaluate the impact of using a convolutional rather than an attention-based image encoder, whereas BERT was replaced by Term Frequency–Inverse Document Frequency (TF-IDF) representations to assess the value of contextual language modelling against a non contextual text baseline. The results show that the proposed model excels at identifying random image and textual mismatches on this proxy benchmark, achieving an ROC-AUC between 0.989 and 0.999, and maintains strong ranking quality under an imbalanced 90/10 class distribution (ROC-AUC 0.9894). However, performance declines substantially when faced with hard negatives (ROC-AUC approximately 0.70), indicating that semantically similar mismatches remain a challenge for dual-encoder architectures. The ablation results confirmed that both the transformer-based vision and language encoders are important for performance relative to the convolutional and TF-IDF alternatives. Qualitative analysis showed that attention maps and token rankings highlight influential image regions and tokens; these visualisations were not evaluated with quantitative faithfulness tests and are not claimed as causal explanations. The study is therefore limited to a public multi-caption photograph benchmark with algorithmically constructed negatives, and the findings should not be generalised to domains whose images, documents, or contradiction types differ from this setting.
  • Item
    A Data Science Approach to Measuring Cognitive Offloading and Short-Term Independent Problem Solving in AI-Assisted Tasks: A Two-Session Experimental Study in Kampala, Uganda
    (Uganda Christian University, Mukono, 2026-09-05) Richard Wambede
    This study investigates the impact of Artificial Intelligence (AI) assistance on human cognitive offloading and short-term independent problem-solving performance. Utilizing a two-session experimental study with a mixed-methods factorial design conducted in Kampala, Uganda, the research evaluates how cognitive reliance on LLMs and automated tools influences critical thinking, problem retention and subsequent unassisted execution. To systematically capture these behavioral dynamics, original data-driven metrics were established, including the Cognitive Offloading Index (COI) and the Human Engagement Score (HES). The findings provide empirical insights into balancing algorithmic support with independent skill retention, offering practical guidelines for human-computer interaction frameworks, educational policy and the deliberate design of AI system guardrails
  • Item
    Predictive analysis of pediatric in hospital malaria mortality using machine learning for early clinical intervention
    (Uganda Christian University, 2026-06-17) Dennis Arthur Nyanzi
    Malaria remains a major public health challenge in Uganda especially among the children. Pediatric patients are more vulnerable because of weaker immune systems and faster dis-ease progression. Delayed assessment increases mortality rates. This study aimed at using machine learning to predict mortality among children admitted with malaria. An experi-mental computational research design was adopted on retrospective secondary data. Five supervised machine learning models were developed: Decision Tree Classifier, Random Forest Classifier, Naive Bayes, Logistic Regression and Gradient Boosting model. These models were developed to predict mortality among pediatric malaria patients. Models were evaluated using precision, recall, F1 score, Confusion Matrix, Learning curves and Re-ceiver Operating Characteristic Area Under Curve. K fold cross validation was applied during training and Grid search was used for hyperparameter tuning. The results showed that the Naive Bayes classifier performed best followed by the Logistic Regression model having achieved the highest receiver operating characteristic area under curve of 0.899 and 0.885 respectively. The models were able to accurately predict the likelihood of mortality among children admitted with malaria through a combination of both clinical and labora-tory factors. Machine learning demonstrated potential in early identification of high risk pediatric malaria patients for timely clinical intervention.
  • Item
    Predicting employment outcomes for youth with disabilities in economic empowerment programs: a machine learning approach
    (Uganda Christian University, 2026-06-12) Leonard Akoch
    Youth with disabilities face significant barriers to employment, including discrimination, limited access to education, and inaccessible workplaces which contribute to high unemployment rates and social exclusion. To improve the effectiveness of economic empowerment programs for this demographic, this study developed a machine learning model to predict employment outcomes for youth with disabilities. Drawing from data of 895 youth with disabilities from the Bunyoro subregion of Uganda who participated in economic empowerment programs, encompassing demographic data, disability types, intervention details, and employment status at follow‐up, we trained and evaluated several machine learning models. Among these were ensemble methods such as Random Forest, XGBoost, Gradient Boosting, and Stacking Ensemble. The Stacking Ensemble achieved the best performance with an accuracy of 97.21%, a precision of 92.73%, a recall of 98.08%, and an F1‐score of 95.22% in predicting improved employment status. The key factors driving employment success were soft skills training, the provision of start‐up kits, and the duration of the interventions. This research addresses the critical need to improve the effectiveness of economic empowerment initiatives developed to support youth with disabilities. The findings can inform other programs with similar contexts, contributing to broader development efforts and potentially inspiring the adoption of predictive modeling in other social programs targeting marginalized groups.
  • Item
    Predicting client retention in an urban HIV clinic – a machine learning approach
    (Uganda Christian University, 2025-05) Jonathan Melvin Ikapule
    Retention in HIV care is critical to viral suppression, improved health outcomes, and reduced transmission; however, retention rates remain suboptimal in urban Uganda, with some studies reporting rates below 60%. This study aimed to identify retention predictors and develop a machine learning model to predict retention among people living with HIV (PLHIV) using routinely collected patient-level data. A retrospective cohort study was conducted using data from electronic medical records (EMR) from three urban HIV clinics in Kampala (January 2021 - December 2023). Clients who died or were transferred out were excluded, yielding 22,213 clients. Data included demographic, clinical, and visit-related variables, as well as engineered features like duration on antiretroviral therapy, distance to clinic, and viral suppression history. Retention was defined as attending a scheduled appointment within 90 days. Six classification algorithms were trained and evaluated using a 70:30 split and SMOTE (a technique to balance data). Accuracy, precision, recall, and F1 score assessed model performance. XGBoost outperformed other models, achieving an accuracy of 88% and an F1 score of 0.85. Key predictors, identified using SHAP values for feature importance, included duration on ART, weight, age, baseline CD4, distance to the clinic, and ART adherence. These findings demonstrate the feasibility of using EMR data and machine learning to support data-driven decision-making in HIV programs. Machine learning models integrated into EMR systems can enable real-time identification of clients at risk of disengaging from care, guiding targeted interventions. This study highlights the potential of data science to improve HIV service delivery, although further validation in diverse contexts is needed. Keywords: Antiretroviral Therapy, Classification, EMR, Retention, SHAP, SMOTE, Supervised Learning, XGBoost, Urban Clinic, Uganda.
  • Item
    A Machine Learning approach for identifying at risk pupils and recommending support strategies: a case study of primary schools in Mukono District, Uganda
    (Uganda Christian University, 2026-05-28) Charles Jovans Galiwango
    Academic vulnerability and pupil dropout remain persistent challenges in Ugandan primary education, despite high enrollment rates. Current school support systems are often reactive, intervening only after academic failure has occurred. This study developed a predictive early warning system to proactively identify pupils at risk of academic failure in Mukono District, Uganda. A mixed-methods approach was used, analysing structured records of pupils from Primary 4–6 and conducting interviews with teachers and administrators. The study first identified key behavioural and socioeconomic predictors of academic risk through statistical analysis. Four machine learning models were then evaluated and compared to determine the most effective approach for predicting vulnerability. The analysis revealed that behavioural indicators, specifically disciplinary issues, incomplete homework, and poor attendance, were the strongest predictors of academic risk. Among the models tested, Logistic Regression proved most suitable, achieving a recall of 0.833 and ROC-AUC of 0.941 on unseen test data, while providing interpretable predictions crucial for educational settings. Based on these findings, a three-tiered intervention framework was developed, classifying pupils by risk level and linking specific risk factors to tailored support strategies. The study concludes that a simple, interpretable predictive model using routinely collected school data can effectively identify vulnerable pupils early. The proposed framework offers Ugandan primary schools a practical, proactive tool for targeted intervention, shifting support from crisis management to prevention. This research contributes a feasible, evidence-based approach to enhancing educational equity and retention in resource-constrained settings.
  • Item
    Predictive maintenance of centrifugal water pumps using machine learning: a case study of National Water and Sewerage Corporation
    (Uganda Christian University, 0026-05-28) Quinton Ssebaggala
    This thesis explores Effective predictive maintenance strategies for Centrifugal water pumps, focusing on Uganda’s National Water and Sewerage Corporation (NWSC) and other similar large-scale water providers, aiming to improve water supply reliability for over 21 million people and reduce 185,000 annual customer complaints caused by 70% pump failures and 8-12 hours of operational downtime. However, despite advances in machine learning, tailored predictive maintenance approaches for water pumps in Uganda are understudied. Thus, this study developed a based predictive maintenance model for the centrifugal pumps using real-time operational data from National Water and Sewerage Corporation (NWSC) (N=13 pumps from the 3 pump stations i.e. (Gunhill, Katosi and uyenga)were analyzed. This study presents a comprehensive Machine-learning based predictive maintenance framework for estimating pump failure. The process integrates data preprocessing, extraction of statistical time-domain condition indicators, and evaluation of 5 machine learning algorithms; XGBoost, LightGBM, CatBoost, Random Forest, and a Voting Ensemble [applied to shift maintenance from a monthly health check to real-time monitoring] providing deeper insights into pump availability and health for future years. The primary objective was to accurately classify the pumps’ operational status into five distinct states: CHANGE, CRITICAL, OFF, OPERATIONAL, and WARNING. The results demonstrate that Extreme Gradient Boosting (XGBoost) model achieved superior predictive performance yielding an accuracy of 74% in detecting failure within pumps before more damage was done. Thus, leveraging of Machine Learning for Predictive maintenance enabled National Water and Sewerage Corporation to detect any anomalies in the Centrifugal pumps like; inconsistencies in flow rates, pressure fluctuations, vibration abnormalities etc. which helped reduce on the maintenance costs from (10-40%), reduce on equipment failure (70-75%), reduced on downtime (35% -45%) and lastly, increased on production capacity by(25%) thus improving on the well-being of the people in Uganda and promoting of SDG 6( Clean water and Sanitation). Keywords: Predictive Maintenance, Machine Learning, Centrifugal Pumps, National Water and Sewerage Corporation, Arduino, Classification Models
  • Item
    Predicting final CGPA using pre-admission data: proactive insights for academic excellence at Uganda Christian University
    (Uganda Christian University, 2026-06-29) Simon Fred Lubambo
    Higher education institutions increasingly rely on data-driven approaches to improve student support, academic planning, and decision-making. However, the adoption of predictive analytics in Sub-Saharan African universities remains limited, despite the availability of admission records that could inform early academic guidance. This study developed and evaluated a machine learning model for predicting students’ final Cumulative Grade Point Average (CGPA) using pre-admission academic and demographic data from Uganda Christian University. The study employed a quantitative research design using historical student records extracted from the university’s Management Information System. The dataset included O-Level and A-Level academic performance indicators, demographic attributes, and programme-related variables. Several machine learning models were trained and evaluated, with the Random Forest Regressor selected as the best-performing model after hyperparameter optimisation. Model performance was assessed using Mean Absolute Error, Root Mean Squared Error, and the coefficient of determination. To support transparency and responsible use, SHAP-based interpretability, sensitivity analysis, and subgroup fairness evaluation were incorporated. The findings showed that final CGPA can be predicted from admission-time data with moderate but useful accuracy. Prior academic performance, particularly average O-Level grade, weighted A-Level performance, and UCE credits, emerged as the strongest predictors of final CGPA. Fairness analysis across gender, campus, and academic level indicated generally consistent model performance, although continued monitoring is necessary for underrepresented groups. The study further demonstrated how predicted CGPA bands and programme-fit simulations can support proactive academic advising, early identification of students requiring support, and evidence-informed programme guidance. The study concludes that interpretable and fair machine learning models can provide practical value in Ugandan higher education when used as decision-support tools rather than deterministic placement mechanisms. By using pre-admission data already available within institutional systems, universities can strengthen academic advising, improve student support, and promote more evidence-based planning.
  • Item
    Data-driven precision public health: leveraging machine learning to track and reduce zero-dose and partially vaccinated children in Nakifuma, Uganda
    (Uganda Christian University, 2026) Kenneth Michael Ogwok
    Despite global progress, 14.3 million infants remain zero-dose (ZD) and 5.6 million are partially vaccinated (PV) worldwide (World Health Organization, 2024). In Uganda, where full immunization coverage stands at only 54% (Uganda Bureau of Statistics, 2022), precision public health approaches are urgently needed. This study applies data science to develop a community-level risk profiling framework in a resource-limited Ugandan setting. This study aimed to: (1) identify socio-demographic, health system, and behavioral factors distinguishing ZD, PV, and fully immunized (FI) children; (2) develop and validate machine learning (ML) models predicting vaccination status; and (3) propose data-driven interventions to increase FI coverage.A mixed-methods, cross-sectional study sampled 115 children and their caregivers under five in Nakifuma Sub-county. For objective one, 35 variables were analyzed using chi-square and Mann-Whitney U tests to identify significant predictors. For objective two, four supervised ML algorithms were trained on a stratified 70:30 split and evaluated using precision, recall, F1-score, and AUC. For objective three, validated model-derived risk scores informed targeted, parish-level interventions.The presence of ZD children (10.4%) was associated with negative attitudes of health workers (p=0.013), waiting time >60 minutes (p=0.021), importance of vaccines (p=0.018), and non-parent caregivers (p=0.026). The presence of PV children (40.9%) was associated with increasing child age (p<0.001) and vaccine stock-out (p=0.031), while FI children (48.7%) possessed vaccination cards (p=0.005). The best-performing algorithm was Random Forest, with an F1-score of 0.97 for ZD, 0.74 for PV, and 0.94 for FI. The clustering of ZD/PV children beyond 2 km from health facilities was used for designing a three-tier intervention matrix for sensitizing health workers, supply chain interventions, and SMS reminders.ML models were effective in triaging zero-dose, partially vaccinated, and fully immunized children. The precision public health strategy has immense scope for achieving 90% full immunization by 2030 in Uganda.