Classification Of Stroke Risk Based On Machine Learning: A Comparative Study Of Naive Bayes And Decision Tree
Keywords:
stroke risk classification, naive bayes, decision tree, comparative study, disease predictionAbstract
Abstract—Stroke is the leading cause of death in Indonesia, accounting for 19.42% of total mortality, with prevalence increasing by 56% from 7 per 1,000 population in 2013 to 10.9 per 1,000 population in 2018 based on Riskesdas data. Beyond its impact on morbidity and mortality, stroke is a catastrophic disease with the third-highest healthcare expenditure in Indonesia, reaching IDR 5.2 trillion in 2023. This condition underscores the urgency of early stroke risk detection to reduce mortality rates and the economic burden on the healthcare system. This study aims to implement and compare the performance of the Naive Bayes and Decision Tree algorithms in stroke risk classification. The dataset used contains 4,981 patient records with 10 stroke risk features, including age, gender, hypertension, heart disease, marital status, occupation type, residence type, average glucose level, body mass index, and smoking status. The research methodology encompasses data collection, preprocessing, splitting the data into training and testing sets at an 80:20 ratio (3,984 training records and 997 testing records), implementation of both algorithms, and model evaluation using accuracy, precision, recall, and F1-score metrics. The results indicate that the Decision Tree algorithm outperforms Naive Bayes, achieving an accuracy of 91.68%, precision of 91.91%, recall of 91.68%, and F1-score of 91.79%, while Naive Bayes achieved an accuracy of 86.06%, precision of 92.09%, recall of 86.06%, and F1-score of 88.72%. Confusion matrix analysis shows that Decision Tree has a lower misclassification rate, with an accuracy margin of 5.62% over Naive Bayes. This study concludes that decision trees are more effective for stroke risk classification. In practical terms, this model implies the provision of a decision-support tool for medical personnel to perform initial triage of high-risk patients prior to further clinical examination, thereby accelerating intervention and reducing the cost burden of stroke care. Theoretically, the findings enrich the literature on classification algorithm comparisons in the healthcare domain and serve as a baseline for the development of advanced models incorporating class imbalance handling and hyperparameter optimization in future research.

