Comparative Study of Artificial Intelligence Accuracy on Project Team Member Assessment
DOI:
https://doi.org/10.22441/jurnal_mix.2026.v16i2.014Keywords:
Artificial Intelligence, Performance Assessment, Prediction Accuracy, Comparative MethodAbstract
Objectives: This study compares five AI platforms (ChatGPT 5.2, Claude Sonnet 4.5, Gemini 3 Pro, GLM 4.7, DeepSeek V3.2) with manual HR assessment in team performance evaluation, examining accuracy and predictive validity by correlating with program success rates.
Methodology: Using a within-subjects design, 32 team members managing 619 locations were evaluated by six methods via standardized prompts. Analysis included comparative tests, accuracy metrics, and correlation between performance scores and graduation rates.
Finding: All AI platforms scored higher than the HR baseline. Correlation analysis revealed substantial differences in predictive validity. Claude Sonnet 4.5 demonstrated the strongest correlation with graduation rates (ρ = 0.70), followed by GLM 4.7 (ρ = 0.67), accounting for approximately 50-55% of the outcome variability. The HR baseline unexpectedly showed a negative correlation with program success (ρ = -0.49), indicating misalignment between the evaluation criteria and actual effectiveness. DeepSeek V3.2 most closely matched HR scores but showed only moderate predictive validity, while ChatGPT 5.2 showed no meaningful relationship with outcomes.
Conclusion: AI platforms differ substantially in predictive validity. Organizations should prioritize outcome-based validation over human agreement when selecting AI for HR decisions and consider recalibrating manual evaluation systems.
References
Adadi, A., & Berrada, M. (2018). Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access, 6, 52138–52160. https://doi.org/10.1109/ACCESS.2018.2870052
Anthropic. (2024). Claude 4.5 model card. https://www.anthropic.com/claude
Barredo A.A, Díaz-Rodríguez A, Del Ser J., Bennetot A., Tabik S., Barbado A., Garcia S. , Gil-Lopez S., Molina D., Benjamins R., Chatila R., Herrera F. (2020). Explainable Artificial Intelligence (XAI): : Concepts, taxonomies, opportunities and challenges toward responsible AI. https://doi.org/10.1016/j.inffus.2019.12.012
Benbya, H., Davenport, T. H., & Pachidi, S. (2021). Special issue editorial: Artificial intelligence in organizations: Current state and future opportunities. MIS Quarterly Executive, 20(4), ix–xxi. https://doi.org/10.17705/2msqe.00054
Brook Fischer. (2024). Candidate evaluation: How to confidently decide who to hire. Homerun. https://www.homerun.co/articles/candidate-evaluations
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
Cascio, W. F., & Aguinis, H. (2019). Applied psychology in talent management (8th ed.). SAGE Publications.
Charlwood, A., & Guenole, N. (2022). Can HR adapt to the paradoxes of artificial intelligence? Human Resource Management Journal, 32(4), 729–742. https://doi.org/10.1111/1748-8583.12433
Chen, T., Fu, M., Liu, R., Xu, X., Zhou, S., & Liu, B. (2019). How do project management competencies change within the project management career model in large Chinese construction companies? International Journal of Project Management, 37(3), 485–500. https://doi.org/10.1016/j.ijproman.2018.12.002
Chien, C.-F., & Chen, L.-F. (2020). Using rough set theory to recruit and retain high-potential talents for semiconductor manufacturing. IEEE Transactions on Semiconductor Manufacturing, 33(4), 529–539.
Chui, M., Hazan, E., Roberts, R., Singla, A., Smaje, K., Sukharevsky, A., Yee, L., & Zemmel, R. (2023). The economic potential of generative AI: The next productivity frontier. McKinsey & Company. https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier
Derry Holt. (2024). Recruitment assessment: In-depth guide for agencies 2024. OneUpSales. https://oneupsales.com/blog/recruitment-assessment
Dwivedi, Y. K., Hughes, L., Ismagilova, E., Aarts, G., Coombs, C., Crick, T., Duan, Y., Dwivedi, R., Edwards, J., Eirug, A., Galanos, V., Ilavarasan, P. V, Janssen, M., Jones, P., Kumar, A., Kar, A., Kizgin, H., Kronemann, B., Lal, B., & Williams, M. D. (2021). Artificial intelligence (AI): Multidisciplinary perspectives on emerging challenges, opportunities, and agenda for research, practice and policy. International Journal of Information Management, 57, 101994. https://doi.org/10.1016/j.ijinfomgt.2019.08.002
Foster, B. (2020). Manajemen sumber daya manusia. Alfabeta.
Intelligence, S. H.-C. A. (2025). Artificial Intelligence Index Report 2025. Stanford University.
Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). Noise: A flaw in human judgment. Little, Brown Spark.
Köchling, J., & Wehner, M. (2020). Discriminated by an algorithm: a systematic review of discrimination and fairness by algorithmic decision-making in the context of HR recruitment and HR development. Business Research, 13(3), 795–848. https://doi.org/10.1007/s40685-020-00134-w
Kirkpatrick, D. L., & Kirkpatrick, J. D. (2016). Kirkpatrick’s Four Levels of Training Evaluation. ATD Press.
Kluyver, T., Ragan-Kelley, B., Pérez, F., Granger, B., Bussonnier, M., Frederic, J., Kelley, K., Hamrick, J., Grout, J., Corlay, S., Ivanov, P., Avila, D., Abdalla, S., & Willing, C. (2016). Jupyter Notebooks—A publishing format for reproducible computational workflows. Positioning and Power in Academic Publishing: Players, Agents and Agendas, 87–90. https://doi.org/10.3233/978-1-61499-649-1-87
Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), Article 33267. https://doi.org/10.1525/collabra.33267
Lambrecht, A., & Tucker, C. (2019). Algorithmic bias? An empirical study of apparent gender-based discrimination in the display of STEM career ads. Management Science, 65(7), 2966–2981. https://doi.org/10.1287/mnsc.2018.3093
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Yian, Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., … Koreeda, Y. (2023). Holistic Evaluation of Language Models. http://arxiv.org/abs/2211.09110
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6). https://doi.org/10.1145/3457607
Özkan, N., & Yücel, İ. (2024). Development and validation of the project manager skills scale (PMSS): An empirical approach. Heliyon, 10(2), e24075. https://doi.org/10.1016/j.heliyon.2024.e24075
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., VanderPlas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. https://jmlr.org/papers/v12/pedregosa11a.html
PPM School of Management. (2025). Management skill: Pengertian, fungsi, jenis dan contoh di dunia kerja. https://www.ppmschool.ac.id/management-skill/
Purnomo V.M, F., Perizade Badia & Syapril Yuliani. (2023). The Influence of Training and Work Experience on Performance with Competence as an Intervening Variable in The Head Of The Technical Implementing Unit Pt. Indonesian Railways. International Journal of Social Service and Research (IJSSR), Vol. 03, No. 08, e-ISSN: 2807-8691. https://doi.org/10.46799/ijssr.v3i8.499
Rosalia, D., Mintarti, S., & Fitriadi. (2018). Pengaruh pelatihan dan pengalaman kerja terhadap produktivitas kerja karyawan Jaya Sakti Sentosa. Jurnal Ekonomi Manajemen Akuntansi, 2(2), 1–15.
Sánchez-Monedero, J., Dencik, L., & Edwards, L. (2020). What does it mean to “solve” the problem of discrimination in hiring? Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 458–468. https://doi.org/10.1145/3351095.3372849
Siddiqui, J., Kosmala, K., Marszałek, A., & Rudawska, E. (2024). Communication in project team management: Identification of research gaps and direction for future research. European Research Studies Journal, 27(Special Issue), 732–751. https://doi.org/10.35808/ersj/3746
Stewart, G. L. (2019). Team member skill and ability levels, personality traits, background and experience in workplace team performance: A meta-analysis. Team Performance Management, 25(1/2), 3–29.
Tambe, P., Cappelli, P., & Yakubovich, V. (2019). Artificial intelligence in human resources management: Challenges and a path forward. California Management Review, 61(4), 15–42. https://doi.org/10.1177/0008125619867910
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., … Scialom, T. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. http://arxiv.org/abs/2307.09288
Vallat, R. (2018). Pingouin: Statistics in Python. Journal of Open Source Software, 3(31), 1026. https://doi.org/10.21105/joss.01026
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., & Larson, E. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17(3), 261–272. https://doi.org/10.1038/s41592-019-0686-2
White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., & Schmidt, D. C. (2023). A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT. http://arxiv.org/abs/2302.11382
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., Tam, W. L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Zhang, P., Dong, Y., & Tang, J. (2023). GLM-130B: An Open Bilingual Pre-trained Model. http://arxiv.org/abs/2210.02414
Zhang, D., Mishra, S., Brynjolfsson, E., Etchemendy, J., Ganguli, D., Grosz, B., & Perrault, R. (2022). The AI Index 2022 Annual Report. Stanford Institute for Human-Centered AI, Stanford University.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 MIX: JURNAL ILMIAH MANAJEMEN

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
The copyright to this article is transferred to Universitas Mercu Buana (UMB) if and when the article is accepted for publication. The undersigned hereby transfers any and all rights in and to the paper including without limitation all copyrights to UMB. The undersigned hereby represents and warrants that the paper is original and that he/she is the author of the paper, except for material that is clearly identified as to its original source, with permission notices from the copyright owners where required. The undersigned represents that he/she has the power and authority to make and execute this assignment.
We declare that this paper has not been published in the same form elsewhere.
Furthermore, I/We hereby transfer the unlimited rights of publication of the above mentioned paper in whole to UMB. The copyright transfer covers the right to reproduce and distribute the article, including reprints, translations, photographic reproductions, microform, electronic form (offline, online) or any other reproductions of similar nature.
The corresponding author signs for and accepts responsibility for releasing this material on behalf of any and all co-authors. This agreement is to be signed by at least one of the authors who have obtained the assent of the co-author(s) where applicable. After submission of this agreement signed by the corresponding author, changes of authorship or in the order of the authors listed will not be accepted.
Retained Rights/Terms and Conditions
Although authors are permitted to re-use all or portions of the Work in other works, this does not include granting third-party requests for reprinting, republishing, or other types of re-use.











