ISSN 2996-8223
International Journal of Computer Science and Technology | Vol. 1, No. 12, December 2020 | pp. 89–96
Research Article
Title: Cross-Project Defect Prediction Using Domain Adaptation and Extreme Gradient Boosting
Names of Authors: Min-Ji Kim¹, Seung-Woo Lee², and Ji-Youn Park³
Authors’ Affiliations:
¹Department of Computer Science and Engineering, Seoul National University, Seoul, South Korea
²Department of Software Engineering, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea
³Department of Information Technology, Yonsei University, Seoul, South Korea
Abstract: Software defect prediction models assist quality assurance teams by identifying bug-prone source code modules, optimizing testing resource allocation. However, building reliable within-project defect predictors requires extensive historical data, which is unavailable for newly initiated or open-source software projects. Cross-Project Defect Prediction (CPDP) addresses this limitation by using historical data from a source project to train a classifier for a target project. The primary challenge in CPDP is the feature distribution mismatch between the source and target domains, which severely degrades the performance of standard machine learning classifiers. This paper presents a novel CPDP framework that integrates Transfer Component Analysis (TCA) for domain adaptation with an optimized Extreme Gradient Boosting (XGBoost) classifier. The TCA module maps software metrics from heterogeneous source and target projects into a low-dimensional latent space where the distance between data distributions is minimized. The XGBoost model, optimized via random search cross-validation, then executes classification within this aligned feature space. Evaluated on 10 open-source software repositories from the PROMISE dataset, the proposed framework improves the F1-measure by 24.6% and the Area Under the Curve (AUC) by 18.3% compared to non-transfer baseline predictors. The results establish that aligning feature distributions across projects mitigates data scarcity and enhances code auditing efficiency in new software systems.
Keywords: Software defect prediction, Cross-project prediction, Domain adaptation, Transfer component analysis, Extreme gradient boosting, Machine learning
Manuscript Timeline: Received: September 10, 2020; Revised: October 18, 2020; Accepted: November 14, 2020; Published: December 1, 2020
Citation: Kim, M. J., Lee, S. W., & Park, J. Y. (2020). Cross-project defect prediction using domain adaptation and extreme gradient boosting. International Journal of Computer Science and Technology, 1(12), 89–96. DOI: 10.46882/2020/IJCST/000012
Subscribe to read the full article: https://internationalscholarsjournals.org/subscribe-to-read