Master of Science in Data Science and Analytics
Permanent URI for this collectionhttps://hdl.handle.net/20.500.11951/1204
Browse
Browsing Master of Science in Data Science and Analytics by Title
Now showing 1 - 20 of 30
- Results Per Page
- Sort Options
Item A Data Science Approach to Measuring Cognitive Offloading and Short-Term Independent Problem Solving in AI-Assisted Tasks: A Two-Session Experimental Study in Kampala, Uganda(Uganda Christian University, Mukono, 2026-09-05) Richard WambedeThis study investigates the impact of Artificial Intelligence (AI) assistance on human cognitive offloading and short-term independent problem-solving performance. Utilizing a two-session experimental study with a mixed-methods factorial design conducted in Kampala, Uganda, the research evaluates how cognitive reliance on LLMs and automated tools influences critical thinking, problem retention and subsequent unassisted execution. To systematically capture these behavioral dynamics, original data-driven metrics were established, including the Cognitive Offloading Index (COI) and the Human Engagement Score (HES). The findings provide empirical insights into balancing algorithmic support with independent skill retention, offering practical guidelines for human-computer interaction frameworks, educational policy and the deliberate design of AI system guardrailsItem A Data-Driven NLP Skills Gap Analysis of Uganda’s TVET Curriculum and its Effects on Graduate Employability(Uganda Christian University, 2025-09-23) Patrick AtuheThis thesis evaluates the outcomes of the revisions to Uganda’s Technical and Vocational Education and Training (TVET) curriculum, focusing on graduate employability. The study applies data science methodologies, particularly Natural Language Processing (NLP), to assess how well the current curriculum aligns with industry needs. Data was collected from 350 TVET graduates, feedback from 50 employers who assessed over 1,250 graduates, and 30 stakeholders analyzed the curriculum. An NLP-based recommendation system was developed using TF-IDF and cosine similarity to quantify alignment between skills taught and those required in the workforce. Findings reveal significant gaps in digital skills, technical preparedness, and alignment with evolving industry expectations. Employers reported a 68% deficiency in digital competencies, with a mean curriculum-employer similarity score of 0.42. The NLP system achieved an F1-score of 0.87, outperforming manual reviews in skill-gap identification. The study provides actionable recommendations for curriculum reform, including the integration of digital tools, periodic review mechanisms, and the use of real-time feedback loops from the industry. These insights contribute to national development goals such as Uganda Vision 2040 by enhancing TVET effectiveness and workforce readiness.Item A Machine Learning Approach for Accurate Valuation of Imports in Uganda(Uganda Christian University, 2025-09-30) Paul SentongoAccurate customs valuation is central to revenue mobilization, trade compliance, and economic stability in Uganda, where import duties contribute nearly one-third of domestic tax revenue. Yet persistent inefficiencies in conventional valuation methods such as reliance on importer-declared invoice values, outdated price databases, and manual adjudication have resulted in systemic undervaluation, mis invoicing, and annual revenue losses exceeding USD 200 million. This thesis investigates the potential of machine learning (ML) to transform customs valuation by developing and deploying predictive models trained on more than 70,000 import declaration records from Uganda Revenue Authority’s ASYCUDA system (2020–2024). Three supervised ML algorithms; Random Forest, Extreme Gradient Boosting (XGBoost), and Artificial Neural Networks (ANN) were implemented following a rigorous pipeline that included exploratory data analysis, feature engineering, and model optimization. All models demonstrated strong predictive performance (R² >0.93), with Random Forest achieving near-perfect accuracy(R² = 0.997, MAE = UGX 560.35, RMSE = UGX 1,868.23). Compared to Uganda’s current average based approach (MAE = UGX124,797.76), this represents a 99.55% reduction in error, underscoring the transformative capacity of ML for valuation precision. Beyond model benchmarking, the study contributes technically by operationalizing the Random Forest model into a Streamlit based prototype web application, offering real-time decision support for customs officers. Empirically, it provides the first quantified evidence of ML’s potential to address valuation fraud and inefficiencies in Uganda. Practically, it establishes a replicable frame work for low-resource settings, integrating ML with existing trade platforms such as ASYCUDA. The findings have significant policy implications: adopting ML-driven valuation can curtail revenue leakages, enhance compliance with WTO Customs Valuation Agreements, and support Uganda’s Vision 2040 and National Development Plan III goals for domestic revenue mobilization. Limitations such as reliance on secondary data, exclusion of informal trade, and simulation based deployment highlight opportunities for future research. These include incorporating regional datasets, exploring explainable AI techniques (e.g., SHAP, LIME) to improve transparency, and piloting ML integration within operational customs systems. This thesis thus advances the discourse on AI in public sector modernization, demonstrating that machine learning is not merely a technical innovation but a strategic enabler for fiscal sustainability, trade integrity, and digital transformation in Uganda’s customs administration.Item A Machine Learning approach for identifying at risk pupils and recommending support strategies: a case study of primary schools in Mukono District, Uganda(Uganda Christian University, 2026-05-28) Charles Jovans GaliwangoAcademic vulnerability and pupil dropout remain persistent challenges in Ugandan primary education, despite high enrollment rates. Current school support systems are often reactive, intervening only after academic failure has occurred. This study developed a predictive early warning system to proactively identify pupils at risk of academic failure in Mukono District, Uganda. A mixed-methods approach was used, analysing structured records of pupils from Primary 4–6 and conducting interviews with teachers and administrators. The study first identified key behavioural and socioeconomic predictors of academic risk through statistical analysis. Four machine learning models were then evaluated and compared to determine the most effective approach for predicting vulnerability. The analysis revealed that behavioural indicators, specifically disciplinary issues, incomplete homework, and poor attendance, were the strongest predictors of academic risk. Among the models tested, Logistic Regression proved most suitable, achieving a recall of 0.833 and ROC-AUC of 0.941 on unseen test data, while providing interpretable predictions crucial for educational settings. Based on these findings, a three-tiered intervention framework was developed, classifying pupils by risk level and linking specific risk factors to tailored support strategies. The study concludes that a simple, interpretable predictive model using routinely collected school data can effectively identify vulnerable pupils early. The proposed framework offers Ugandan primary schools a practical, proactive tool for targeted intervention, shifting support from crisis management to prevention. This research contributes a feasible, evidence-based approach to enhancing educational equity and retention in resource-constrained settings.Item A Machine Learning Framework for Flood Risk Classification and Early warning in River Catchment Areas(Uganda Christian University, 2026-09-29) Martha Frances NamakulaFlooding remains one of the most destructive natural hazards in Uganda, with the River Man afwa catchment area on the slopes of Mount Elgon experiencing recurrent and severe flood events that displace communities, destroy infrastructure, and disrupt livelihoods. Existing flood forecasting in Uganda relies heavily on global systems such as GloFAS and ECMWF, which operate at low spatial resolution and are often unreliable for small catchments, limiting their effectiveness for local early warning and disaster preparedness. This study developed a machine learning framework for flood risk classification in the River Manafwa catchment, using historical rainfall data (CHIRPS satellite, 1989–2024) and observed river discharge data (Min istry of Water and Environment, 1989–2024). Hydrological features were engineered, including antecedent rainfall accumulations (1, 3, 5, and 7-day windows), lagged and rolling-mean dis charge variables, discharge differencing and acceleration terms, and seasonal cyclical indicators. A locally calibrated four-level flood risk classification (Normal, Mild, Advanced, Extreme) was developed based on discharge percentile thresholds (Q75, Q90, and Q95). Five machine learn ing models, Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, and MLP Neural Network, were trained and evaluated using Macro F1-score, Extreme F1-scores, and AUC-ROC, with SMOTE applied to address class imbalance in the training data. Gradient Boosting achieved the strongest overall performance, recording the highest Macro F1-score and Extreme-class F1-score following SMOTE application, and was selected as the most suitable model for this study. SHAP analysis identified short-term lagged discharge, particularly the previous day’s discharge, as the dominant indicator of flood risk across all classes, while same day rainfall and short-term cumulative rainfall accumulations retained comparatively greater influence specifically within the Extreme risk class. The model was validated against three documented historical flood events (June 2019, October 2019, and September 2021), with the model successfully flagging Extreme risk conditions several days ahead of officially reported flood dates. The resulting framework offers a low-cost, replicable, data-driven approach to strengthen flood early warning systems for the Manafwa catchment and other flood-prone rivers in Uganda, supporting disaster preparedness agencies including URCS, UNMA, and the OPM’s Department of Relief and Disaster Preparedness.Item A Multi-Class Machine Learning Customer Classification For Personalised Voice Bundle Recommendations: A Case Study of Airtel Uganda Limited(Uganda Christian University, 2026-09-14) Bisimbeko, RemmyAirtel Uganda Limited relies on Average Revenue Per User (ARPU)-banded segmentation to deliver voice bundle promotions, yielding conversion rates below 3 percent and wasting an estimated 20 percent of the marketing budget on untargeted communications. The absence of a personalised, data-driven recommendation system represents both a commercial ine"ciency and an academic gap this study addresses. This study developed and validated a machine learning recommendation framework using real transactional data from 1,458,900 Airtel Uganda prepaid subscribers (February–April 2026). Four classification models — Logistic Regression, Decision Tree, Random Forest, and XGBoost — were evaluated following a preprocessing pipeline of median imputation, capped SMOTE oversampling, and behavioural feature engineering. Three statistical tests were applied to evaluate relationships between customer attributes and bundle segment: Spearman Rank Correlation to measure the strength and direction of monotonic associations between numeric features and the target; one-way ANOVA to test whether feature means di!er significantly across the eight value segment groups; and Chi-Square to assess the independence of categorical features from the target variable. Statistical tests confirmed 34 of 35 features significantly associated with bundle segment (p < 0.05); UPGRADE_RATE was the only non-significant feature. Feature importance analysis identified TOTAL_REVENUE (18.01%), SPEND_PER_MINUTE (6.57%), CALLS (5.97%), VOICE_INTENSITY (5.87%), and MOBILE_MONEY (4.64%) as the top five predictors. XGBoost achieved the highest F1-Score of 0.661 on the 8-class hold-out test set — a 5.4-fold improvement over a random baseline. A simulated pilot projected a four- to five-fold improvement in conversion rate (3% to 12–15%) and a 9 percent ARPU uplift (simulated pilot projection). Three strategic recommendations were formulated: a VOICE_INTENSITY-triggered recommendation engine, enriched segmentation incorporating device type and purchase channel, and a formal A/B test with model governance. This is the first empirically validated ML bundle recommendation framework built on real Ugandan telecom data. Keywords: Machine Learning, Telecommunications, Voice bundle recommendation, Customer behavioural attributes, Predictive Analytics.Item A Predictive Analytics Framework for Case Backlog Management in Uganda’s Judiciary: An Explainable Machine Learning Approach(Uganda Christian University, 2026-10-02) Bbossa, Isaac SserunkumaThe Uganda judicial system is faced with a persistent and rising backlog of cases, whereby 26.32% of cases have remained pending for more than two years, negatively impacting access to justice and socio-economic development. Current systems such as the Electronic Court Case Management Information System (ECCMIS) can only make use of descriptive reporting techniques, thus imple- menting a reactive administrative paradigm. This research bridges the critical gap in the area of predictive capability through developing an explainable machine learning model to preemptively pre- dict the backlog risk of cases. Adopting the Design Science Research method, the research adopted a quantitative experimental design. Making use of secondary data from the National Court Case Census (2025), the research included substantial data pre-processing, feature engineering, and com- paring various machine learning classifiers including Logistic Regression, Random Forest, Gradient Boosting, and K-Nearest Neighbors to carry out binary classification task of backlog prediction. The best performing classifier was Random Forest with an ROC AUC of 0.885 and F1 Score of 0.698, significantly outperforming other classifiers. The explainable AI (XAI) analysis indicated that Claim/offence (case complexity), Court name (institutional capacity), and case status (procedural stage) were the most predictive features. This further confirms the hypothesis that backlog risk is essentially due to the legal and institutional dynamics rather than merely time dynamics. The analysis succeeded in developing an empirical predictive model that proved the potency of ensemble models such as Random Forest in capturing the complex non-linear dynamics of judicial backlog in Uganda. The results allow shifting from reactive to proactive case management strategy. This thesis ends up with some practical recommendations regarding the implementation of a predictive warning system, dynamic resource planning, and targeted procedural reform.Item A Text-based Poultry Health System: An Interactive Disease Detection and Prescription Recommendations(Uganda Christian University, 2025-09-23) Ritah NakimuliPoultry farming is vital to Uganda’s economy, providing income for many rural households. However, broiler chicken farmers struggle with early disease detection and management, leading to significant flock losses and financial hardship. Although advanced diagnostic tools exist, they are often too expensive and complicated for small-scale farmers in rural areas to access. This research presents a multilingual, symptom-based poultry disease prediction system, a lightweight, mobile-friendly machine learning solution that addresses the limitations of existing diagnostic tools. By allowing farmers to input observable symptoms like bird behavior, droppings, and flock age through a simple text-based interface, it eliminates the need for costly equipment, lab tests, or other traditional methods. Several machine learning algorithms were tested to identify the best method for disease prediction, including SVM, Random Forest, XGBoost, and KNN. KNN and SVM performed best, each achieving 96% accuracy and 97% precision, with Random Forest close behind. XGBoost performed poorly, with only 11% accuracy. Although SVM matched KNN in accuracy, it struggled with real-world probability calibration. KNN, on the other hand, provided reliable and interpretable confidence scores, making it the preferred choice for deployment. The final application is deployed using the Streamlit framework, enabling seamless access across desktop and mobile browsers. It provides real-time disease predictions, along with tailored prescriptions and prevention strategies. Additional features include a QR code for easy sharing, which enhances both the user experience and accessibility. This project bridges the gap between advanced AI and the practical realities of low-resource agricultural settings.Item An Explainable Transformer-Based Vision-Language Model For Multimodal Consistency Verification(none, 2026-09-23) Mugimba Kakure JudeThe rapid growth of multimodal data has underscored the necessity for robust methods to verify semantic alignment between images and their associated textual descriptions. This problem, termed multimodal consistency verification, asks whether a given image and text pair are semantically compatible or contradictory. Unlike conventional image–text retrieval, which ranks many candidates for a query, consistency verification is a pairwise decision on a single pair. While large-scale vision-language models demonstrate strong performance in retrieval and zero-shot tasks, explicit consistency verification, particularly with integrated explainability, remains underexplored. We present an explainable dual-encoder vision-language architecture that utilizes a Vision Transformer (ViT) and Bidirectional Encoder Representations from Transformers (BERT) to encode images and multi-caption documents, respectively. These representations are projected into a shared embedding space and aligned via a symmetric InfoNCE contrastive objective. We quantify consistency using cosine similarity, employing a decision threshold to distinguish between consistent and inconsistent pairs. The model incorporates visual attention maps from the ViT and token importance scores from the BERT encoder that provide qualitative evidence for its classification decisions. For evaluation, a proxy benchmark dataset was constructed from the Flickr30K dataset by aggregating the five captions associated with each image into a single multi-caption document and generating consistent, randomly mismatched, and Facebook AI Similarity Search (FAISS) retrieved hard-negative pairs. The inconsistent pairs were generated algorithmically rather than through extensive human annotation of semantic contradiction type. The model was tested across balanced, imbalanced, and hard-negative scenarios and compared with Contrastive Language–Image Pre-training (CLIP) as a baseline in all scenarios. To isolate the contributions of the transformer-based encoders, the Vision Transformer was replaced by Residual Network-50 (ResNet-50) to evaluate the impact of using a convolutional rather than an attention-based image encoder, whereas BERT was replaced by Term Frequency–Inverse Document Frequency (TF-IDF) representations to assess the value of contextual language modelling against a non contextual text baseline. The results show that the proposed model excels at identifying random image and textual mismatches on this proxy benchmark, achieving an ROC-AUC between 0.989 and 0.999, and maintains strong ranking quality under an imbalanced 90/10 class distribution (ROC-AUC 0.9894). However, performance declines substantially when faced with hard negatives (ROC-AUC approximately 0.70), indicating that semantically similar mismatches remain a challenge for dual-encoder architectures. The ablation results confirmed that both the transformer-based vision and language encoders are important for performance relative to the convolutional and TF-IDF alternatives. Qualitative analysis showed that attention maps and token rankings highlight influential image regions and tokens; these visualisations were not evaluated with quantitative faithfulness tests and are not claimed as causal explanations. The study is therefore limited to a public multi-caption photograph benchmark with algorithmically constructed negatives, and the findings should not be generalised to domains whose images, documents, or contradiction types differ from this setting.Item Contextualising AI Ethics in Uganda’s Microcredit With Adaptive Sensitive Re-weighting(Uganda Christian University, 2025-08-12) Emmanuel IsabiryeThis research tackles the pressing ethical concerns of using Artificial Intelligence (AI) in Uganda’s microcredit sector, namely to develop an Adaptive Sensitive Reweighting (ASR) model to mitigate algorithmic bias and promote equitable access to credit. Traditional credit scoring models - and AI algorithms trained on Western-biased data - discriminate against marginalized groups because they are based on formal financial records, reinforcing structural disadvantages. By iterative engagement with Ugandan policymakers, lenders, borrowers, and AI experts, we identify the most significant ethical concerns and specify context-specific fairness metrics. The ASR approach adaptively adjusts weights for sensitive features like collateral values and transaction history during model training to enhance fairness. Experimental outcomes on a typical credit scoring dataset demonstrate ASR’s success: the inclusion rate of disadvantaged borrowers is enhanced by 15% with predictive accuracy maintained, and significant improvements on key fairness metrics. The research provides actionable policy recommendations on implementing ASR-based AI systems in Uganda’s microfinance sector to drive financial inclusion and sustainable development. This study contributes to emerging Majority World scholarship on AI ethics by demonstrating the necessity of situating ethical frameworks and valuing stakeholder perspectives to develop equitable, inclusive AI systems. Our findings offer valuable insights for policymakers, microfinance institutions, and AI practitioners who aim to implement responsible AI in Uganda’s Microcredit sector.Item Data-driven Analysis and Prediction of Human Rights Violations Against Human Rights Defenders: A Case Study of Eastern Africa(Uganda Christian University, 2025-09-29) Esther Asiimire BagombekaDespite the growing availability of big data and machine learning, human rights monitoring in the region remains largely dependent on retrospective reports, eyewitness testimonies, and qualitative assessments, which lack the ability to anticipate future violations. The absence of real- time data processing and predictive analytics limits the ability of policymakers and advocacy groups to implement proactive intervention strategies. As a result, human rights organizations often respond reactively, only after violations occur, rather than deploying preemptive measures to protect HRDs. In this research, a quantitative research design was adopted, utilising a cross-sectional approach to analyse patterns in human rights violations. Data was collected from recognized human rights organisations, human rights databases, and global news agencies. The research employed descriptive analytics to identify trends, K-Means clustering to categorize high-risk regions, and predictive modeling to forecast future violations. Seasonal Autoregressive Integrated Moving Average (SARIMA) was used to model long-term seasonal trends, while Recurrent Neural Networks (RNN) captured short-term fluctuations and nonlinear patterns in the data. The Predictive Human Rights Violations Model (PHRVM) emerged as the most effective, balancing structural seasonality and real-time variations, resulting in higher accuracy and improved forecasting reliability compared to individual models. The findings revealed that human rights violations followed distinct temporal and geo- graphic trends, peaking around election periods, protest seasons, and government crackdowns. While the PHRVM outperformed other forecasting methods during training (MAE : 0.081, RMSE : 0.087), testing revealed a slight increase in prediction error, with MAE rising to 0.684 and RMSE increasing to 1.109. A paired t-test confirmed that the model significantly outperformed a naïve baseline forecast (p < 0.05), validating its predictive capability. This research concluded that human rights violations follow recognizable patterns, making it possible to anticipate high-risk periods and optimize protection efforts for HRDs. This helps policymakers, and advocacy groups to anticipate risks and implement preventive measures before violations escalate. The PHRVM’s success shows the potential of AI-driven forecasting in social science research, offering a more systematic approach to tracking civic space restric- tions. However, for predictive models to be more effective in real-world applications, further refinement is needed, including the integration of real-time data sources such as social media monitoring, remote sensing technologies, and expanded human rights reporting networks. Strengthening these capabilities will enhance model accuracy, responsiveness, and impact, ensuring that human rights organisations can move from reactive responses to preventative protection strategies.Item Data-driven precision public health: leveraging machine learning to track and reduce zero-dose and partially vaccinated children in Nakifuma, Uganda(Uganda Christian University, 2026) Kenneth Michael OgwokDespite global progress, 14.3 million infants remain zero-dose (ZD) and 5.6 million are partially vaccinated (PV) worldwide (World Health Organization, 2024). In Uganda, where full immunization coverage stands at only 54% (Uganda Bureau of Statistics, 2022), precision public health approaches are urgently needed. This study applies data science to develop a community-level risk profiling framework in a resource-limited Ugandan setting. This study aimed to: (1) identify socio-demographic, health system, and behavioral factors distinguishing ZD, PV, and fully immunized (FI) children; (2) develop and validate machine learning (ML) models predicting vaccination status; and (3) propose data-driven interventions to increase FI coverage.A mixed-methods, cross-sectional study sampled 115 children and their caregivers under five in Nakifuma Sub-county. For objective one, 35 variables were analyzed using chi-square and Mann-Whitney U tests to identify significant predictors. For objective two, four supervised ML algorithms were trained on a stratified 70:30 split and evaluated using precision, recall, F1-score, and AUC. For objective three, validated model-derived risk scores informed targeted, parish-level interventions.The presence of ZD children (10.4%) was associated with negative attitudes of health workers (p=0.013), waiting time >60 minutes (p=0.021), importance of vaccines (p=0.018), and non-parent caregivers (p=0.026). The presence of PV children (40.9%) was associated with increasing child age (p<0.001) and vaccine stock-out (p=0.031), while FI children (48.7%) possessed vaccination cards (p=0.005). The best-performing algorithm was Random Forest, with an F1-score of 0.97 for ZD, 0.74 for PV, and 0.94 for FI. The clustering of ZD/PV children beyond 2 km from health facilities was used for designing a three-tier intervention matrix for sensitizing health workers, supply chain interventions, and SMS reminders.ML models were effective in triaging zero-dose, partially vaccinated, and fully immunized children. The precision public health strategy has immense scope for achieving 90% full immunization by 2030 in Uganda.Item Detection of Banana Fusarium Wilt & Black Sigatoka: A Deep Learning Approach for Smallholder Farms in Central Uganda(Uganda Christian University, 2025-10-20) Peter MulindwaBananas are a vital food and income source in Central Uganda, yet their cultivation is severely threatened by destructive diseases such as Fusarium Wilt and Black Sigatoka. Smallholder farmers, who form the backbone of Uganda’s agricultural sector, rely heavily on manual disease identification methods, which are time-consuming, error-prone, and largely ineffective for early intervention. This thesis proposes a hybrid deep learning approach that integrates Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), and Gray Level Co-occurrence Matrix (GLCM) texture features to provide accurate, efficient, and scalable detection of banana leaf diseases using image-based classification. A dataset of over 17,000 annotated banana leaf images was sourced from the Lacuna Banana project. Rigorous preprocessing, including resizing, normalisation, and augmentation, was applied to enhance model robustness. Texture features extracted through GLCM were combined with spatial features learned by CNN and ViT models to improve classification sensitivity. Several models were developed and evaluated, including a custom CNN, InceptionV3 with transfer learning, and a ViT-based architecture. Evaluation metrics such as accuracy, precision, recall, and F1-score were used to assess model performance. The Vision Transformer outperformed other individual models with 99% classification accuracy, while the proposed hybrid model achieved a balanced accuracy of 98%, with substantial precision and recall across all disease categories. The integration of GLCM features significantly improved the detection of texture-specific diseases like Black Sigatoka. This research contributes a robust, interpretable, and field-deployable AI-based diagnostic tool that aligns with Uganda’s national goals for data-driven agricultural development and the food security-related Sustainable Development Goals (and sustainable farming.Item Enhancing food security forecasting in Uganda: a geospatial Analytics framework with crop yield integration(Uganda Christian University, 2026-09-30) Arinaitwe, PhilipFood security is a crucial pillar for human well being, societal stability and national development. Food security can be defined as a situation when all people, at all times, have physical and economic access to sufficient, safe and nutritious food that meets their dietary needs and food preferences for an active and healthy life. In Uganda, nearly 20 percent of the population lives below the poverty line and over 16 million Ugandans face food insufficiency due to various factors like unreliable climate, economic instability, population pressure, conflict and inadequate forecasting mechanisms. Recent research studies have increasingly turned to predictive modelling using machine learning for food security forecasting, integration of climate data, socio-economic indicators and innovative geospatial analytics. Whereas advances in machine learning and geospatial analytics have led to significant improvements in climatic modelling for food security, they often fall short by excluding actual agricultural outputs. Without integrating crop yield data, forecasts may misestimate food availability yet it can be a key driver of food crises. Particularly for Uganda, where over 70% of the population is employed by agriculture, the omission of crop yield data from forecasting models can result in inaccurate predictions/forecasts. To address this gap, this study developed and evaluated a geospatial machine learning framework for forecasting food security in Uganda by integrating crop yield data from FAOSTAT with multi-source environmental, market, conflict, demographic and historical Integrated Food Security Phase Classification (IPC) data from 2007 to 2020. A hybrid model combining a Long Short-Term Memory (LSTM) temporal encoder and an XGBoost classifier was trained and validated on 44, 082 district-month observations. The hybrid LSTM–XGBoost model achieved a macro F1-score of 0.75 and an overall accuracy of 93.8% on the test set. Importantly, the model demonstrated exceptional performance in identifying severe food crises (IPC 3+), achieving a recall of 0.82, a ROC-AUC of 0.998 and a PR-AUC of 0.852. An ablation study confirmed that incorporating crop yield data significantly improved predictive power, increasing the IPC 3+ F1-score from 0.688 to 0.759. Furthermore, model comparisons demonstrated that a country-specific model trained purely on Ugandan observations significantly outperformed a consolidated regional model shown by a macro F1 score of 0.747 achieved by the country-specific model as compared to the consolidated regional model with a macro F1-score of 0.725. These results provide robust empirical evidence that integrating crop yield data into geospatial frameworks enhances food security early warning systems which enables timely, targeted interventions to mitigate severe hunger.Item Forecasting Emerging Skill Demands with Machine Learning to Inform Curriculum Development in Uganda’s Higher Education(Uganda Christian University, 2025-09-24) Denis WanyamaRapid technological advancement and evolving industry demands have widened the skill gap in Uganda’s labor market. Higher education institutions often struggle to keep pace with these changes, leading to mismatches between graduate competencies and employer expectations. This study uses machine learning techniques to forecast emerging skill demands and inform the development of data-driven curricula in Ugandan universities. Drawing on more than one million job postings from 2021 to 2023, the research applies natural language processing (NLP), time series forecasting (ARIMA and Holt-Winters), and clustering algorithms to analyze labor market trends. Exploratory Data Analysis (EDA) revealed high-demand skills, while Holt-Winters outperformed ARIMA (MAE: 9.05 vs. 23.87), capturing the seasonal nature of skill fluctuations. Key findings indicate a growing demand for roles such as interaction designers, network administrators, user experience professionals, and social media managers. In-demand technical skills include Python, Google Analytics, CSS, Tableau, AWS, and Sketch. The increasing emphasis on digital literacy and soft skills underscores the need for more flexible and adaptive curricula. This study offers actionable recommendations for curriculum reform, including integrating technical skills, developing continuous learning pathways, and enhancing academic-industry collaboration. By applying machine learning to labor market analysis, the research equips universities, policymakers, and stakeholders with the information needed to align higher education with the demands of Uganda’s evolving digital economy. Keywords: Machine Learning, Labor Market Trends, Skill Forecasting, Higher Education Curriculum, Uganda, ARIMA, Holt-Winters, Data Science, skill mismatch.Item Improving Employee Retention by Predicting Employee Attrition using Machine Learning Techniques :Case Study: Centenary Bank Ltd Uganda(Uganda Christian University, 2025-10-13) Andrew Ronnie EngirotEmployee retention is a critical factor in the success and sustainability of organizations, ensuring that valuable human capital remains engaged, satisfied, and motivated over the long term. High turnover rates can significantly disrupt productivity, damage organizational culture, and inflate operational costs, underscoring the importance of retaining top talent. Continuity in operations is maintained when employees feel valued and supported, fostering a positive work environment where collaboration and high performance are encouraged. In contrast, frequent turnover can lead to instability and decreased morale, ultimately hindering productivity and organizational cohesion. Retaining top talent not only maintains operational continuity but also provides a competitive edge in the marketplace. Organizations with strong employee retention rates attract prospective hires more effectively and are better positioned to develop deep expertise within their workforce, contributing to long-term success. In today’s digital age, where social media and online reviews can quickly shape a company’s reputation, prioritizing employee satisfaction and well-being enhances brand image and appeals to both job seekers and consumers. From a financial perspective, employee retention contributes to significant cost savings. The expenses associated with recruiting, hiring, and training new employees are substantial. By retaining existing employees, organizations can allocate resources more efficiently, as long-term employees are typically more productive and require less supervision. This not only reduces direct costs but also improves overall organizational efficiency and effectiveness. HR analytics has emerged as a powerful tool in predicting and enhancing employee retention. By adopting a data-driven approach, HR analytics involves the collection, analysis, and interpretation of data related to employee behaviors and performance to inform strategic decisions. This approach combines HR-specific data, such as employee demographics, performance metrics, and engagement surveys, with financial and operational data to generate comprehensive insights into workforce trends. One of the key roles of HR analytics in employee retention is identifying predictors or drivers of turnover. By analyzing historical turnover data alongside various HR metrics, organizations can detect patterns and trends indicating employees are at risk of leaving. Predictive models employing algorithms and machine learning techniques analyze large datasets to forecast potential attrition, enabling proactive measures to address these risks. Additionally, HR analytics facilitates sentiment analysis and engagement surveys to assess employee satisfaction and pinpoint areas requiring improvement. Techniques such as natural language processing (NLP) and text analytics allow for the examination of unstructured data from employee feedback, performance reviews, and social media, providing deep insights into employee sentiment and morale. Beyond prediction, HR analytics informs the development of targeted retention strategies. By understanding the underlying factors contributing to attrition, organizations can implement personalized development opportunities, improve communication between managers and employees, and adjust compensation and benefits packages to better align with employee expectations. These tailored interventions aim to enhance engagement, satisfaction, and ultimately, retention. The objectives of this project are to enhance employee retention by leveraging machine learning techniques to predict and mitigate employee attrition. Through the analysis of historical employee data and the application of predictive modeling, the project seeks to identify key factors contributing to turnover and develop actionable insights for proactive retention strategies. The project aims to build predictive models that accurately anticipate staff attrition by examining historical data on demographics, job categories, performance measures, and other relevant variables. Furthermore, the study intends to identify critical organizational factors that predict employee attrition. By analyzing the output of predictive models, the research aims to pinpoint specific risk factors associated with higher turnover rates, such as job dissatisfaction, inadequate compensation, or lack of career advancement opportunities. Based on these insights, the project will formulate targeted intervention strategies to address identified risk factors and reduce employee churn. Recommendations will focus on enhancing employee retention, engagement, and satisfaction. The effectiveness of these interventions will be evaluated by monitoring key performance indicators such as employee satisfaction ratings, attrition rates, and retention metrics. The goal is to assess the impact of implemented strategies on workforce stability and organizational performance over time. Additionally, the project aims to establish a framework for the ongoing evaluation and refinement of retention strategies and predictive models. Through continuous data analysis and feedback mechanisms, the project seeks to iteratively enhance the effectiveness of retention initiatives and improve the accuracy of predictive models, ensuring that retention efforts evolve to meet the organization’s changing needs.Item Interpretable Predictive Modeling of Postpartum Surgical Site Infections(Uganda Christian University, 2026-07-11) Emmanuel NkurunzizaBackground: Postpartum surgical site infections (SSIs) are a major cause of maternal morbidity in low-resource settings, yet early detection is limited by reactive, symptom-based surveillance systems. Ma chine learning (ML) and explainable artificial intelligence (XAI) offer potential tools for proac tive SSI risk identification using routinely collected clinical data. Methods: A retrospective dataset of 8,918 caesarean deliveries from three regional referral hospitals in Uganda (SSI prevalence: 8.1%) was analysed. After rigorous cleaning, leakage-free feature engineering, imputation, and encoding, four ML models Logistic Regression, Random Forest, Support Vector Machine (RBF), and XGBoost were trained and evaluated under extreme class imbalance using SMOTENC. Model performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and precision–recall curves. SHAP explainability was applied to identify globally and locally influential predictors. Results: Overall discriminatory performance was modest due to limited predictive signal in routine clin ical records. Logistic Regression achieved the highest sensitivity (recall = 0.9931; ROC-AUC = 0.6528), whereas XGBoost produced the highest accuracy (0.7932) but poor recall (0.2639). SHAP analysis highlighted preoperative showering, skin antiseptic preparation, obstructed labour, and prolonged surgery duration as key contributors to increased SSI risk. Conclusion: Explainable ML models are feasible for early SSI risk identification in low-resource settings but are constrained by sparse routine data. Logistic Regression may serve as a high-sensitivity clinical screening tool. Improved performance will require enriched clinical datasets, inclu sion of intraoperative variables, and multi-centre validation to support practical deployment in maternal health surveillance.Item Machine Learning Decomposition of Gender and Rural–Urban Wage Gaps: Evidence from Uganda National Household Surveys(2026-09-30) Tumushiime, Bob RobertUganda's gender wage gap stands at about 32 percent (UN Women, 2024), with rural–urban differences of similar size. Published decomposition studies for Uganda have relied largely on the 2002/03 survey round and a linear wage equation that assumes constant returns to characteristics. This study examines whether a flexible functional form, applied to recent multi-round data, alters the share of each gap attributed to measured characteristics. Four Uganda National Household Survey rounds (2012/13–2023/24) were harmonised into a pooled sample of 14,963 formal wage earners aged 14–64. Ordinary least squares (OLS), LASSO, Random Forest and XGBoost were trained on a stratified 70/15/15 split, and SHapley Additive exPlanations (SHAP) were used to attribute predictions to worker characteristics. Predictions from the best-performing model replaced the linear fitted values in an Oaxaca–Blinder decomposition estimated for each round. XGBoost achieved the highest test R² (0.410), followed by Random Forest (0.399) and the two linear models (0.354); pooling all four rounds raised it from 0.286. Education was the leading predictor. The rural–urban gap was largely explained under both methods: 49–59 percent under OLS and 68–84 percent under the machine-learning decomposition. For the gender gap, OLS produced a negative explained share in three of four rounds, while the machine-learning decomposition attributed 14–57 percent to measured characteristics. The explained share of the gender gap is thus sensitive to the functional form of the wage model, which calls for caution in interpreting single-round linear estimates. The approach can be reproduced on successive survey rounds to track wage inequality over time. The unexplained component is reported as a residual and is not interpreted as a measure of discrimination.Item Machine Learning Forecasting for Electricity Demand Using Climate and Socioeconomic Data in Sub-Saharan Africa(Uganda Christian University, 2026-09-30) Lutalo, LordinElectricity demand forecasting is important for energy planning, resource allocation, and infrastructure development. In Sub-Saharan Africa, forecasting electricity demand is particularly difficult because electricity usage is driven by rapid population growth, economic changes, urbanisation, and climate variability, while data quality and availability remain inconsistent across countries. Although previous studies showed that climate and socioeconomic variables are associated with electricity demand, many focused on short-term forecasting, single-country analysis, or datasets from developed regions. This study examined how far integrated climate and socioeconomic data improves annual electricity demand forecasting across 47 Sub-Saharan African countries. Country-level panel time-series data was collected from publicly available sources including the World Bank, Our World in Data, and the ERA5 climate reanalysis dataset, covering the period from 2000 to 2021. The methodology included data integration and preprocessing, normality testing, temporal feature engineering, a two-stage feature selection process, and the training and comparison of six forecasting models: Linear Regression, Fixed Effects Panel Regression, Dynamic Fixed Effects Panel Regression, Random Forest, XGBoost, and LightGBM. Model performance was evaluated using Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), Mean Absolute Percentage Error (MAPE), and the coefficient of determination (R^2). A decision-support dashboard was also developed using Streamlit to visualise electricity demand forecasts, trends, and key demand drivers. The results showed that lagged electricity demand was by far the strongest predictor of future demand, with the one-year lag achieving a near-perfect Pearson correlation of 0.999 with the target variable. Among structural climate and socioeconomic predictors, population size and its three-year rolling mean were the strongest predictors of electricity demand, followed by the three-year rolling mean of temperature, the interaction between population and electricity access, and the temperature-urbanisation interaction. Linear Regression on the full multivariate feature set achieved the best overall test set performance, with a log-scale MAPE of 6.01% and an R^2 of 0.996 (note: all metrics are on the log-transformed demand scale, not in original TWh units), and was selected as the final model for the dashboard. The study confirms that annual electricity demand in Sub-Saharan Africa is extremely persistent, lagged demand provides most of the predictive power in the model. Climate and socioeconomic variables contain substantial structural information about demand patterns when historical demand is unavailable, but contribute little additional one-year-ahead predictive accuracy once lagged demand is already included. This distinction between predictive accuracy and structural explanation is the core finding of the study.Item Multi-horizon predictive modeling of HIV treatment adherence using clinical and social determinants in Uganda(Uganda Christian University, 2026-06-05) Jacob NyonyintonoTreatment interruption among people living with HIV (PLHIV) remains a critical challenge to achieving sustained viral suppression and optimal treatment outcomes, particularly in sub- Saharan Africa. In Uganda, despite significant progress in scaling up antiretroviral therapy (ART), a substantial proportion of patients experience interruptions in treatment, contributing to viral rebound, increased transmission risk, and drug resistance. The adoption of the Multi- Month Dispensing (MMD) model, while improving access and convenience, reduces routine patient-provider contact and may delay early identification of patients at risk of disengagement. This study aimed to develop and evaluate a machine learning-based predictive framework for identifying patients at risk of treatment interruption under Uganda’s MMD system. A quantitative research design was employed using retrospective data extracted from Electronic Medical Records (EMRs) of 8,788 patients receiving ART. Data preprocessing involved cleaning, feature engineering, and handling missing values, followed by the development of machine learning models using a structured pipeline. The Random Forest algorithm was selected as the primary model due to its ability to capture complex nonlinear relationships and its robustness in handling imbalanced clinical datasets. The models were trained and evaluated across three temporal prediction windows (30, 60, and 90 days) using performance metrics including precision, recall, F1-score, and ROC-AUC. Particular emphasis was placed on precision to ensure reliable identification of high-risk patients while minimizing false positive classifications. The findings demonstrate that integrating clinical, demographic, and behavioral factors improves the predictive ability of machine learning models in identifying treatment interruption risk. Key predictors included viral load status, CD4 count, duration on ART, and selected socio-demographic characteristics. The developed framework enables the generation of individualized risk scores and prediction of likely interruption periods, supporting proactive patient management. This study contributes to the growing body of evidence on the application of machine learning in HIV care and highlights the importance of incorporating multidimensional determinants in predictive modelling. The proposed model provides a practical tool for early risk identification and targeted intervention, with potential to enhance patient retention and improve treatment outcomes within Uganda’s HIV care system.
