Principal Component Analysis pada Random Forest untuk Deteksi Dini Dropout Mahasiswa
DOI:
https://doi.org/10.67401/j-csat.v4i2.36Keywords:
Principal Component Analysis, Random Forest, Dropout Mahasiswa, Educational Data Mining, Ekstraksi FiturAbstract
Penelitian ini menguji pengaruh Principal Component Analysis (PCA) terhadap performa Random Forest dalam deteksi dini risiko dropout mahasiswa. Dua skema dibandingkan, yaitu Random Forest tanpa PCA dan dengan PCA, menggunakan Orange Data Mining (Orange3) pada dataset Kaggle Predict Students' Dropout and Academic Success (4.424 data, 35 atribut, tiga kelas target). PCA diterapkan dengan dua komponen utama (explained variance 26%). Kelas Dropout dievaluasi menggunakan validasi silang 10-lipat bertingkat dengan metrik AUC, Akurasi, F1, Presisi, Recall, dan MCC. Random Forest tanpa PCA memperoleh performa lebih tinggi pada seluruh metrik (AUC 0,905; Accuracy 0,856; F1 0,771; Precision 0,790; Recall 0,754; MCC 0,667) dibandingkan Random Forest dengan PCA (AUC 0,834; Accuracy 0,800; F1 0,676; Precision 0,704; Recall 0,650; MCC 0,533), yang diduga disebabkan oleh rendahnya variansi kumulatif yang dipertahankan.
Downloads
References
[1] R. Quincho Apumayta, J. Carrillo Cayllahua, A. Ccencho Pari, V. Inga Choque, J. C. Cárdenas Valverde, and D. Huamán Ataypoma, “University Dropout: A Systematic Review of the Main Determinant Factors (2020-2024),” F1000Research, vol. 13, pp. 1–20, 2024, doi: 10.12688/f1000research.154263.2.
[2] A. Algiffary and T. Sutabri, “Indonesian Journal of Computer Science,” Indones. J. Comput. Sci., vol. 12, no. 2, pp. 284–301, 2023, [Online]. Available: http://ijcs.stmikindonesia.ac.id/ijcs/index.php/ijcs/article/view/3135
[3] Y. Lin, H. Chen, W. Xia, F. Lin, Z. Wang, and Y. Liu, “A Comprehensive Survey on Deep Learning Techniques in Educational Data Mining,” Data Sci. Eng., vol. 10, no. 4, pp. 564–590, 2025, doi: 10.1007/s41019-025-00303-z.
[4] S. G. A. Utami, H. Setiadi, and A. Rohmadi, “Comparative Analysis of Machine Learning Algorithms with RFE-CV for Student Dropout Prediction,” J. Tek. Inform., vol. 6, no. 3, pp. 1319–1338, 2025, doi: 10.52436/1.jutif.2025.6.3.4695.
[5] O. Rute et al., “Jurnal EurekaMatika,” vol. 13, no. 1, pp. 67–80, 2025.
[6] A. M. Fajria, A. Faqih, and G. Dwilestari, “The Impact of Principal Component Analysis Dimensionality Reduction on Sentiment Classification Performance Using Support Vector Machine,” J. Artif. Intell. Eng. Appl., vol. 4, no. 2, pp. 764–770, 2025, doi: 10.59934/jaiea.v4i2.744.
[7] Ridha Afifa, Muhammad Itqan Mazdadi, Triando Hamongan Saragih, Fatma Indriani, and Muliadi, “4015-10989-1-Pb[1],” vol. 13, pp. 1852–1864, 2024.
[8] E. A. F. Elmuna, T. Chamidy, and F. Nugroho, “Optimization of the Random Forest Method Using Principal Component Analysis to Predict House Prices,” Int. J. Adv. Data Inf. Syst., vol. 4, no. 2, pp. 155–166, 2023, doi: 10.25008/ijadis.v4i2.1290.
[9] Z. Dobesova, “Evaluation of Orange data mining software and examples for lecturing machine learning tasks in geoinformatics,” Comput. Appl. Eng. Educ., vol. 32, no. 4, pp. 1–18, 2024, doi: 10.1002/cae.22735.
[10] D. Chicco, N. Tötsch, and G. Jurman, “The matthews correlation coefficient (Mcc) is more reliable than balanced accuracy, bookmaker informedness, and markedness in two-class confusion matrix evaluation,” BioData Min., vol. 14, pp. 1–22, 2021, doi: 10.1186/s13040-021-00244-z.
[11] R. Riaz, M. A. Chatni, T. Rehman, M. Bin Tahir, and S. Khan, “Using Orange Data Mining and Predictive Analytics for Analytical Workflows in Academic Research,” vol. 7, no. 3, pp. 23–30, 2026.
[12] D. Priyanto, H. Hairani, K. Marzuki, and M. Innuddin, “Optimization of Random Forest for Health Data Classification Using PCA and K-Means SMOTE-ENN,” Eng. Technol. Appl. Sci. Res., vol. 15, no. 5, pp. 27646–27652, 2025, doi: 10.48084/etasr.12976.
[13] Z. Abidin, T. Suratno, and M. F. Putri, “Penerapan Random Oversampling dan Principal Component Analysis untuk Meningkatkan Akurasi Prediksi Kebangkrutan Perusahaan di Indonesia dengan Model Machine Learning,” J. Teknol. Inf. dan Ilmu Komput., vol. 12, no. 5, pp. 1209–1220, 2025, [Online]. Available: https://jtiik.ub.ac.id/index.php/jtiik/article/view/9972
[14] L. F. Voges, L. C. Jarren, and S. Seifert, “Exploitation of surrogate variables in random forests for unbiased analysis of mutual impact and importance of features,” Bioinformatics, vol. 39, no. 8, 2023, doi: 10.1093/bioinformatics/btad471.
[15] L. G. R. Putra, D. D. Prasetya, and M. Mayadi, “Student Dropout Prediction Using Random Forest and XGBoost Method,” INTENSIF J. Ilm. Penelit. dan Penerapan Teknol. Sist. Inf., vol. 9, no. 1, pp. 147–157, 2025, doi: 10.29407/intensif.v9i1.21191.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Rully Ghozali, Kevin Ramadhan Oktaviano, Anando Ferdi Setiawan, Muhammad Rio Ramadani, Heru Saputro (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.



