Analysis evaluates machine learning methods on microbiome data, highlighting the effectiveness of random forest algorithms.
The application of next-generation sequencing (NGS) technologies has enabled the identification of both culturable and non-culturable microorganisms in blood samples, revealing their potential roles in systemic infections and immune responses. However, the complexity and high dimensionality of microbiome data present significant challenges for analysis. In this study, it was evaluated the performance of various machine learning (ML) algorithms, including logistic regression, random forest (RF), decision tree, and support vector machines (SVM), in classifying 16S rRNA gene sequencing data of blood microbiota into cultured and uncultured groups. The dataset used in this study, obtained from Kalfin and Panaiotov, consists of 16S rRNA gene sequences from a total of 18,093 OTUs and 62 observations, including control samples. After excluding the six control samples, 56 samples from target sequencing of cultured and non-cultured blood samples of healthy individuals were analyzed. Results show that the random forest (RF) algorithm exhibits the highest classification performance, successfully distinguishing between cultured and uncultured blood microbiota. In the study, the potential of ML techniques in microbiome research was evaluated and the effectiveness and accuracy of these techniques in the analysis of microbiome data were investigated.
No takes yet. Share an insight, caveat, or question.
Akay et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: