Principal Component Analysis pada Random Forest untuk Deteksi Dini Dropout Mahasiswa

Authors

  • Rully Ghozali Program Studi Sistem Informasi Universitas Islam Nahdlatul Ulama Jepara Author
  • Kevin Ramadhan Oktaviano Program Studi Sistem Informasi Universitas Islam Nahdlatul Ulama Jepara Author
  • Anando Ferdi Setiawan Program Studi Sistem Informasi Universitas Islam Nahdlatul Ulama Jepara Author
  • Muhammad Rio Ramadani Program Studi Sistem Informasi Universitas Islam Nahdlatul Ulama Jepara Author
  • Heru Saputro Program Studi Sistem Informasi Universitas Islam Nahdlatul Ulama Jepara Author

DOI:

https://doi.org/10.67401/j-csat.v4i2.36

Keywords:

Principal Component Analysis, Random Forest, Dropout Mahasiswa, Educational Data Mining, Ekstraksi Fitur

Abstract

Penelitian ini menguji pengaruh Principal Component Analysis (PCA) terhadap performa Random Forest dalam deteksi dini risiko dropout mahasiswa. Dua skema dibandingkan, yaitu Random Forest tanpa PCA dan dengan PCA, menggunakan Orange Data Mining (Orange3) pada dataset Kaggle Predict Students' Dropout and Academic Success (4.424 data, 35 atribut, tiga kelas target). PCA diterapkan dengan dua komponen utama (explained variance 26%). Kelas Dropout dievaluasi menggunakan validasi silang 10-lipat bertingkat dengan metrik AUC, Akurasi, F1, Presisi, Recall, dan MCC. Random Forest tanpa PCA memperoleh performa lebih tinggi pada seluruh metrik (AUC 0,905; Accuracy 0,856; F1 0,771; Precision 0,790; Recall 0,754; MCC 0,667) dibandingkan Random Forest dengan PCA (AUC 0,834; Accuracy 0,800; F1 0,676; Precision 0,704; Recall 0,650; MCC 0,533), yang diduga disebabkan oleh rendahnya variansi kumulatif yang dipertahankan.

Downloads

Download data is not yet available.

References

[1] R. Quincho Apumayta, J. Carrillo Cayllahua, A. Ccencho Pari, V. Inga Choque, J. C. Cárdenas Valverde, and D. Huamán Ataypoma, “University Dropout: A Systematic Review of the Main Determinant Factors (2020-2024),” F1000Research, vol. 13, pp. 1–20, 2024, doi: 10.12688/f1000research.154263.2.

[2] A. Algiffary and T. Sutabri, “Indonesian Journal of Computer Science,” Indones. J. Comput. Sci., vol. 12, no. 2, pp. 284–301, 2023, [Online]. Available: http://ijcs.stmikindonesia.ac.id/ijcs/index.php/ijcs/article/view/3135

[3] Y. Lin, H. Chen, W. Xia, F. Lin, Z. Wang, and Y. Liu, “A Comprehensive Survey on Deep Learning Techniques in Educational Data Mining,” Data Sci. Eng., vol. 10, no. 4, pp. 564–590, 2025, doi: 10.1007/s41019-025-00303-z.

[4] S. G. A. Utami, H. Setiadi, and A. Rohmadi, “Comparative Analysis of Machine Learning Algorithms with RFE-CV for Student Dropout Prediction,” J. Tek. Inform., vol. 6, no. 3, pp. 1319–1338, 2025, doi: 10.52436/1.jutif.2025.6.3.4695.

[5] O. Rute et al., “Jurnal EurekaMatika,” vol. 13, no. 1, pp. 67–80, 2025.

[6] A. M. Fajria, A. Faqih, and G. Dwilestari, “The Impact of Principal Component Analysis Dimensionality Reduction on Sentiment Classification Performance Using Support Vector Machine,” J. Artif. Intell. Eng. Appl., vol. 4, no. 2, pp. 764–770, 2025, doi: 10.59934/jaiea.v4i2.744.

[7] Ridha Afifa, Muhammad Itqan Mazdadi, Triando Hamongan Saragih, Fatma Indriani, and Muliadi, “4015-10989-1-Pb[1],” vol. 13, pp. 1852–1864, 2024.

[8] E. A. F. Elmuna, T. Chamidy, and F. Nugroho, “Optimization of the Random Forest Method Using Principal Component Analysis to Predict House Prices,” Int. J. Adv. Data Inf. Syst., vol. 4, no. 2, pp. 155–166, 2023, doi: 10.25008/ijadis.v4i2.1290.

[9] Z. Dobesova, “Evaluation of Orange data mining software and examples for lecturing machine learning tasks in geoinformatics,” Comput. Appl. Eng. Educ., vol. 32, no. 4, pp. 1–18, 2024, doi: 10.1002/cae.22735.

[10] D. Chicco, N. Tötsch, and G. Jurman, “The matthews correlation coefficient (Mcc) is more reliable than balanced accuracy, bookmaker informedness, and markedness in two-class confusion matrix evaluation,” BioData Min., vol. 14, pp. 1–22, 2021, doi: 10.1186/s13040-021-00244-z.

[11] R. Riaz, M. A. Chatni, T. Rehman, M. Bin Tahir, and S. Khan, “Using Orange Data Mining and Predictive Analytics for Analytical Workflows in Academic Research,” vol. 7, no. 3, pp. 23–30, 2026.

[12] D. Priyanto, H. Hairani, K. Marzuki, and M. Innuddin, “Optimization of Random Forest for Health Data Classification Using PCA and K-Means SMOTE-ENN,” Eng. Technol. Appl. Sci. Res., vol. 15, no. 5, pp. 27646–27652, 2025, doi: 10.48084/etasr.12976.

[13] Z. Abidin, T. Suratno, and M. F. Putri, “Penerapan Random Oversampling dan Principal Component Analysis untuk Meningkatkan Akurasi Prediksi Kebangkrutan Perusahaan di Indonesia dengan Model Machine Learning,” J. Teknol. Inf. dan Ilmu Komput., vol. 12, no. 5, pp. 1209–1220, 2025, [Online]. Available: https://jtiik.ub.ac.id/index.php/jtiik/article/view/9972

[14] L. F. Voges, L. C. Jarren, and S. Seifert, “Exploitation of surrogate variables in random forests for unbiased analysis of mutual impact and importance of features,” Bioinformatics, vol. 39, no. 8, 2023, doi: 10.1093/bioinformatics/btad471.

[15] L. G. R. Putra, D. D. Prasetya, and M. Mayadi, “Student Dropout Prediction Using Random Forest and XGBoost Method,” INTENSIF J. Ilm. Penelit. dan Penerapan Teknol. Sist. Inf., vol. 9, no. 1, pp. 147–157, 2025, doi: 10.29407/intensif.v9i1.21191.

Downloads

Published

2026-08-19