Comparative Study of Artificial Intelligence Accuracy on Project Team Member Assessment

Authors

  • Poltak Evencus Hutajulu Industrial Chemical Technology Polytechnic, Indonesia
  • Andreas Rumata Simanjuntak Industrial Chemical Technology Polytechnic,
  • Toba Sastrawan Manik Industrial Chemical Technology Polytechnic,
  • R Reza El Akbar Siliwangi University,

DOI:

https://doi.org/10.22441/jurnal_mix.2026.v16i2.014

Keywords:

Artificial Intelligence, Performance Assessment, Prediction Accuracy, Comparative Method

Abstract

Objectives: This study compares five AI platforms (ChatGPT 5.2, Claude Sonnet 4.5, Gemini 3 Pro, GLM 4.7, DeepSeek V3.2) with manual HR assessment in team performance evaluation, examining accuracy and predictive validity by correlating with program success rates.
Methodology: Using a within-subjects design, 32 team members managing 619 locations were evaluated by six methods via standardized prompts. Analysis included comparative tests, accuracy metrics, and correlation between performance scores and graduation rates.
Finding: All AI platforms scored higher than the HR baseline. Correlation analysis revealed substantial differences in predictive validity. Claude Sonnet 4.5 demonstrated the strongest correlation with graduation rates (ρ = 0.70), followed by GLM 4.7 (ρ = 0.67), accounting for approximately 50-55% of the outcome variability. The HR baseline unexpectedly showed a negative correlation with program success (ρ = -0.49), indicating misalignment between the evaluation criteria and actual effectiveness. DeepSeek V3.2 most closely matched HR scores but showed only moderate predictive validity, while ChatGPT 5.2 showed no meaningful relationship with outcomes.
Conclusion: AI platforms differ substantially in predictive validity. Organizations should prioritize outcome-based validation over human agreement when selecting AI for HR decisions and consider recalibrating manual evaluation systems.

References

Adadi, A., & Berrada, M. (2018). Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access, 6, 52138–52160. https://doi.org/10.1109/ACCESS.2018.2870052

Anthropic. (2024). Claude 4.5 model card. https://www.anthropic.com/claude

Barredo A.A, Díaz-Rodríguez A, Del Ser J., Bennetot A., Tabik S., Barbado A., Garcia S. , Gil-Lopez S., Molina D., Benjamins R., Chatila R., Herrera F. (2020). Explainable Artificial Intelligence (XAI): : Concepts, taxonomies, opportunities and challenges toward responsible AI. https://doi.org/10.1016/j.inffus.2019.12.012

Benbya, H., Davenport, T. H., & Pachidi, S. (2021). Special issue editorial: Artificial intelligence in organizations: Current state and future opportunities. MIS Quarterly Executive, 20(4), ix–xxi. https://doi.org/10.17705/2msqe.00054

Brook Fischer. (2024). Candidate evaluation: How to confidently decide who to hire. Homerun. https://www.homerun.co/articles/candidate-evaluations

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.

Cascio, W. F., & Aguinis, H. (2019). Applied psychology in talent management (8th ed.). SAGE Publications.

Charlwood, A., & Guenole, N. (2022). Can HR adapt to the paradoxes of artificial intelligence? Human Resource Management Journal, 32(4), 729–742. https://doi.org/10.1111/1748-8583.12433

Chen, T., Fu, M., Liu, R., Xu, X., Zhou, S., & Liu, B. (2019). How do project management competencies change within the project management career model in large Chinese construction companies? International Journal of Project Management, 37(3), 485–500. https://doi.org/10.1016/j.ijproman.2018.12.002

Chien, C.-F., & Chen, L.-F. (2020). Using rough set theory to recruit and retain high-potential talents for semiconductor manufacturing. IEEE Transactions on Semiconductor Manufacturing, 33(4), 529–539.

Chui, M., Hazan, E., Roberts, R., Singla, A., Smaje, K., Sukharevsky, A., Yee, L., & Zemmel, R. (2023). The economic potential of generative AI: The next productivity frontier. McKinsey & Company. https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier

Derry Holt. (2024). Recruitment assessment: In-depth guide for agencies 2024. OneUpSales. https://oneupsales.com/blog/recruitment-assessment

Dwivedi, Y. K., Hughes, L., Ismagilova, E., Aarts, G., Coombs, C., Crick, T., Duan, Y., Dwivedi, R., Edwards, J., Eirug, A., Galanos, V., Ilavarasan, P. V, Janssen, M., Jones, P., Kumar, A., Kar, A., Kizgin, H., Kronemann, B., Lal, B., & Williams, M. D. (2021). Artificial intelligence (AI): Multidisciplinary perspectives on emerging challenges, opportunities, and agenda for research, practice and policy. International Journal of Information Management, 57, 101994. https://doi.org/10.1016/j.ijinfomgt.2019.08.002

Foster, B. (2020). Manajemen sumber daya manusia. Alfabeta.

Intelligence, S. H.-C. A. (2025). Artificial Intelligence Index Report 2025. Stanford University.

Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). Noise: A flaw in human judgment. Little, Brown Spark.

Köchling, J., & Wehner, M. (2020). Discriminated by an algorithm: a systematic review of discrimination and fairness by algorithmic decision-making in the context of HR recruitment and HR development. Business Research, 13(3), 795–848. https://doi.org/10.1007/s40685-020-00134-w

Kirkpatrick, D. L., & Kirkpatrick, J. D. (2016). Kirkpatrick’s Four Levels of Training Evaluation. ATD Press.

Kluyver, T., Ragan-Kelley, B., Pérez, F., Granger, B., Bussonnier, M., Frederic, J., Kelley, K., Hamrick, J., Grout, J., Corlay, S., Ivanov, P., Avila, D., Abdalla, S., & Willing, C. (2016). Jupyter Notebooks—A publishing format for reproducible computational workflows. Positioning and Power in Academic Publishing: Players, Agents and Agendas, 87–90. https://doi.org/10.3233/978-1-61499-649-1-87

Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), Article 33267. https://doi.org/10.1525/collabra.33267

Lambrecht, A., & Tucker, C. (2019). Algorithmic bias? An empirical study of apparent gender-based discrimination in the display of STEM career ads. Management Science, 65(7), 2966–2981. https://doi.org/10.1287/mnsc.2018.3093

Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Yian, Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., … Koreeda, Y. (2023). Holistic Evaluation of Language Models. http://arxiv.org/abs/2211.09110

Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6). https://doi.org/10.1145/3457607

Özkan, N., & Yücel, İ. (2024). Development and validation of the project manager skills scale (PMSS): An empirical approach. Heliyon, 10(2), e24075. https://doi.org/10.1016/j.heliyon.2024.e24075

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., VanderPlas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. https://jmlr.org/papers/v12/pedregosa11a.html

PPM School of Management. (2025). Management skill: Pengertian, fungsi, jenis dan contoh di dunia kerja. https://www.ppmschool.ac.id/management-skill/

Purnomo V.M, F., Perizade Badia & Syapril Yuliani. (2023). The Influence of Training and Work Experience on Performance with Competence as an Intervening Variable in The Head Of The Technical Implementing Unit Pt. Indonesian Railways. International Journal of Social Service and Research (IJSSR), Vol. 03, No. 08, e-ISSN: 2807-8691. https://doi.org/10.46799/ijssr.v3i8.499

Rosalia, D., Mintarti, S., & Fitriadi. (2018). Pengaruh pelatihan dan pengalaman kerja terhadap produktivitas kerja karyawan Jaya Sakti Sentosa. Jurnal Ekonomi Manajemen Akuntansi, 2(2), 1–15.

Sánchez-Monedero, J., Dencik, L., & Edwards, L. (2020). What does it mean to “solve” the problem of discrimination in hiring? Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 458–468. https://doi.org/10.1145/3351095.3372849

Siddiqui, J., Kosmala, K., Marszałek, A., & Rudawska, E. (2024). Communication in project team management: Identification of research gaps and direction for future research. European Research Studies Journal, 27(Special Issue), 732–751. https://doi.org/10.35808/ersj/3746

Stewart, G. L. (2019). Team member skill and ability levels, personality traits, background and experience in workplace team performance: A meta-analysis. Team Performance Management, 25(1/2), 3–29.

Tambe, P., Cappelli, P., & Yakubovich, V. (2019). Artificial intelligence in human resources management: Challenges and a path forward. California Management Review, 61(4), 15–42. https://doi.org/10.1177/0008125619867910

Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., … Scialom, T. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. http://arxiv.org/abs/2307.09288

Vallat, R. (2018). Pingouin: Statistics in Python. Journal of Open Source Software, 3(31), 1026. https://doi.org/10.21105/joss.01026

Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., & Larson, E. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17(3), 261–272. https://doi.org/10.1038/s41592-019-0686-2

White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., & Schmidt, D. C. (2023). A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT. http://arxiv.org/abs/2302.11382

Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., Tam, W. L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Zhang, P., Dong, Y., & Tang, J. (2023). GLM-130B: An Open Bilingual Pre-trained Model. http://arxiv.org/abs/2210.02414

Zhang, D., Mishra, S., Brynjolfsson, E., Etchemendy, J., Ganguli, D., Grosz, B., & Perrault, R. (2022). The AI Index 2022 Annual Report. Stanford Institute for Human-Centered AI, Stanford University.

Downloads

Published

2026-07-26

How to Cite

Hutajulu, P. E., Simanjuntak, A. R., Manik, T. S., & El Akbar, R. R. (2026). Comparative Study of Artificial Intelligence Accuracy on Project Team Member Assessment. MIX: JURNAL ILMIAH MANAJEMEN, 16(2), 644–659. https://doi.org/10.22441/jurnal_mix.2026.v16i2.014

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.