Browsing by Author "Deogratius Ssemwanga"
Now showing 1 - 2 of 2
- Results Per Page
- Sort Options
Item Employing Phylogenetic Tree Shape Statistics to Resolve the Underlying Host Population Structure(BMC Bioinformatics, 2021-11) Hassan W Kayondo; Alfred Ssekagiri; Grace Nabakooza; Nicholas Bbosa; Deogratius Ssemwanga; Pontiano Kaleebu; Samuel Mwalili; John M Mango; Andrew J Leigh Brown; Roberto A Saenz; Ronald Galiwango; John M KitayimbwaBackground: Host population structure is a key determinant of pathogen and infectious disease transmission patterns. Pathogen phylogenetic trees are useful tools to reveal the population structure underlying an epidemic. Determining whether a population is structured or not is useful in informing the type of phylogenetic methods to be used in a given study. We employ tree statistics derived from phylogenetic trees and machine learning classification techniques to reveal an underlying population structure. Results: In this paper, we simulate phylogenetic trees from both structured and non-structured host populations. We compute eight statistics for the simulated trees, which are: the number of cherries; Sackin, Colless and total cophenetic indices; ladder length; maximum depth; maximum width, and width-to-depth ratio. Based on the estimated tree statistics, we classify the simulated trees as from either a non-structured or a structured population using the decision tree (DT), K-nearest neighbor (KNN) and support vector machine (SVM). We incorporate the basic reproductive number ([Formula: see text]) in our tree simulation procedure. Sensitivity analysis is done to investigate whether the classifiers are robust to different choice of model parameters and to size of trees. Cross-validated results for area under the curve (AUC) for receiver operating characteristic (ROC) curves yield mean values of over 0.9 for most of the classification models. Conclusions: Our classification procedure distinguishes well between trees from structured and non-structured populations using the classifiers, the two-sample Kolmogorov-Smirnov, Cucconi and Podgor-Gastwirth tests and the box plots. SVM models were more robust to changes in model parameters and tree size compared to KNN and DT classifiers. Our classification procedure was applied to real -world data and the structured population was revealed with high accuracy of [Formula: see text] using SVM-polynomial classifier.Item Phylogenetic Networks and Parameters Inferred from HIV Nucleotide Sequences of High-Risk and General Population Groups in Uganda: Implications for Epidemic Control(MDPI, 2021-05) Nicholas Bbosa; Deogratius Ssemwanga; Rebecca N Nsubuga; Noah Kiwanuka; Bernard S Bagaya; John M Kitayimbwa; Alfred Ssekagiri; Gonzalo Yebra; Pontiano Kaleebu; Andrew Leigh-BrownPhylogenetic inference is useful in characterising HIV transmission networks and assessing where prevention is likely to have the greatest impact. However, estimating parameters that influence the network structure is still scarce, but important in evaluating determinants of HIV spread. We analyzed 2017 HIV pol sequences (728 Lake Victoria fisherfolk communities (FFCs), 592 female sex workers (FSWs) and 697 general population (GP)) to identify transmission networks on Maximum Likelihood (ML) phylogenetic trees and refined them using time-resolved phylogenies. Network generative models were fitted to the observed degree distributions and network parameters, and corrected Akaike Information Criteria and Bayesian Information Criteria values were estimated. 347 (17.2%) HIV sequences were linked on ML trees (maximum genetic distance ≤4.5%, ≥95% bootstrap support) and, of these, 303 (86.7%) that consisted of pure A1 (n = 168) and D (n = 135) subtypes were analyzed in BEAST v1.8.4. The majority of networks (at least 40%) were found at a time depth of ≤5 years. The waring and yule models fitted best networks of FFCs and FSWs respectively while the negative binomial model fitted best networks in the GP. The network structure in the HIV-hyperendemic FFCs is likely to be scale-free and shaped by preferential attachment, in contrast to the GP. The findings support the targeting of interventions for FFCs in a timely manner for effective epidemic control. Interventions ought to be tailored according to the dynamics of the HIV epidemic in the target population and understanding the network structure is critical in ensuring the success of HIV prevention programs.
