##plugins.themes.academic_pro.article.main##
Prediksi Kegagalan Build Real-Time Pada Pipeline CI/CD Menggunakan Pendekatan Machine Learning Berbasis Data Streaming
Abstract
Kegagalan proses build pada lingkungan Continuous Integration/Continuous Deployment (CI/CD) dapat menghambat rilis perangkat lunak, meningkatkan biaya perbaikan, dan menurunkan efisiensi pengembangan. Hal ini menunjukkan perlunya sistem prediksi yang cepat, adaptif, serta mampu menangani data yang terus mengalir secara real-time. Penelitian ini bertujuan untuk menganalisis dan membandingkan kinerja beberapa algoritma online machine learning dalam memprediksi kegagalan build pada pipeline CI/CD. Dataset yang digunakan berasal dari TravisTorrent yang mencakup aktivitas repositori, perubahan kode, metrik pengujian, serta riwayat build. Tahapan penelitian meliputi preprocessing data, pembentukan skema data streaming, penerapan model Hoeffding Tree, Adaptive Random Forest (ARF), dan Online Logistic Regression, serta evaluasi menggunakan metrik accuracy. Hasil eksperimen menunjukkan model terbaik mencapai accuracy 97,35%, lebih tinggi dibandingkan penelitian sebelumnya sebesar 95,9%. Peningkatan absolut sebesar 1,45 poin persentase atau sekitar 1,51% ini membuktikan bahwa pendekatan yang digunakan lebih efektif dalam mengenali pola kegagalan build pada lingkungan CI/CD berbasis streaming data.
##plugins.themes.academic_pro.article.details##

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
References
A. Barrak, E. E. Eghan, B. Adams, and F. Khomh, “Why Do Builds Fail?—A Conceptual Replication Study,” J. Syst. Softw., 2021, doi: 10.1016/j.jss.2021.110939.
R. Sharma, E. Petrova, and J. O. Connolly, “Machine Learning – Based Failure Prediction in Continuous Integration and Deployment Workflows,” 2025.
M. Hilton, T. Tunnell, K. Huang, D. Marinov, and D. Dig, “Usage, Costs, And Benefits Of Continuous Integration In Open-Source Projects” ACM Int. Conf. Autom. Softw. Eng., pp. 426–437, 2016, doi: 10.1145/2970276.2970358.
P. Domingos and G. Hulten, “Mining High-Speed Data Streams” KDD Conf., pp. 71–80, 2000.
A. Bifet, “Adaptive Learning And Mining For Data Streams And Frequent Patterns,” ACM SIGKDD Explor. Newsl., vol. 11, no. 1, pp. 55–56, 2009, doi: 10.1145/1656274.1656287.
A. Bifet and R. Gavaldà, “Learning From Time-Changing Data With Adaptive Windowing,” Proc. 7th SIAM Int. Conf. Data Min., pp. 443–448, 2007, doi: 10.1137/1.9781611972771.42.
Hirmayanti and E. Utami, “Enhanced Heart Disease Diagnosis Using Machine Learning Algorithms : A Comparison of Feature Selection,” Rekayasa Sist. dan Teknol. Inf., vol. 5, no. 158, pp. 385–392, 2025, doi: https://doi.org/10.29207/resti.v9i2.6175.
A. Fisher, C. Rudin, and F. Dominici, “All Models are Wrong , but Many are Useful : Learning a Variable ’ s Importance by Studying an Entire Class of Prediction Models Simultaneously,” J. Mach. Learn. Res., vol. 20, pp. 1–81, 2019.
I. A. Al-Baltah, N. Al-Shaibany, M. Abdellatief, M. M. Al-Gawda, and S. Y. Al-Sultan, “An Intelligent Multi-Class XGBoost-Based Model for Optimizing DevOps Continuous Integration and Continuous Deployment Failure Prediction,” Inf., vol. 17, no. 2, pp. 1–14, 2026, doi: 10.3390/info17020178.
J. Humble and D. Farley, Continuous Delivery: Reliable Software Releases Through Build, Test, And Deployment Automation. 2011. doi: 10.1201/9781003221579-1.
A. A. Benczúr, L. Kocsis, and R. Pálovics, “Online Machine Learning in Big Data Streams: Overview,” Encycl. Big Data Technol., pp. 1207–1218, 2019, doi: 10.1007/978-3-319-77525-8_326.
H. M. Gomes et al., “Adaptive Random Forests For Evolving Data Stream Classification,” Mach. Learn., vol. 106, no. 9–10, pp. 1469–1495, 2017, doi: 10.1007/s10994-017-5642-8.
J. Gama, I. Zliobaite, A. Bifet, M. Pechenizkiy, and A. Bouchachia, “A Survey on Concept Drift Adaptation,” ACM Comput. Surv., vol. 46, no. 4, 2014, doi: 10.1145/2523813.
J. L. Crowley, Generative Networks: EigenSpace Coding, Auto-Encoders, Variational Autoencoders and Generative Adversarial Networks. 2020. doi: 10.1017/9781009218276.020.
A. Altmann, L. Tolo, O. Sander, and T. Lengauer, “Permutation Importance : a Corrected Feature Importance Measure,” Bioinformatics, vol. 26, no. 10, pp. 1340–1347, 2010, doi: 10.1093/bioinformatics/btq134.