CodeBERT-SENet: Adaptive Syntax-Semantic Fusion via Gated Attention for Python Bug Detection and Localization
DOI:
https://doi.org/10.21512/commit.v20i2.14377Keywords:
CodeBERT-SENet, Syntax Error Detection, Adaptive Gating, Multi-Task LearningAbstract
Syntax error detection is a critical step in software development to ensure code quality and reliability. The research proposes the first end-to-end multitask architecture that jointly detects Python syntax errors and localizes their exact line position by adaptively fusing CodeBERT’s semantic embeddings with handcrafted syntactic features via a lightweight gating mechanism, without relying on Abstract Syntax Trees (AST). The model is trained and evaluated on a structured Python code dataset (testing set: 342 files, 30% of total data) using consistent data splits and standardized evaluation metrics. The results demonstrate that CodeBERT-SENet achieves state-of-the-art performance with 99.71% accuracy, an F1-score of 0.9969, and a Mean Absolute Error (MAE) of 0.1157 in line-level error prediction, outperforming all baselines, including Vanilla CodeBERT, GraphCodeBERT, RoBERTa, Random Forest, and the rule-based approach. The confusion matrix confirms zero false negatives (all 160 buggy files detected) and only one false positive (181 non-bugged files, 180 correctly identified). Training converges stably within five epochs without signs of overfitting, underscoring the effectiveness of the architectural design and optimization strategy. However, limitations remain, including high inference latency (six seconds per file), lack of deep structural code representation (e.g., AST), regression-based line prediction instead of classification, and unverified generalization on real-world production code. Nevertheless, CodeBERTSENet conclusively demonstrates that adaptive fusion of semantic and syntactic features significantly enhances code diagnostic capability. Future research will focus on optimizing inference speed, integrating AST-based representations, and transitioning to per-line classification, transforming CodeBERT-SENet from a mere bug detector into a next-generation, context-aware code diagnostic assistant.
References
[1] N. Aloufi and A. Aljuhani, “Empirical evaluation of prompting strategies for Python syntax error detection with LLMs,” Applied Sciences, vol. 15, no. 16, pp. 1–19, 2025.
[2] H. M. Tran, S. T. Le, S. V. Nguyen, and P. T. Ho, “An analysis of software bug reports using machine learning techniques,” SN Computer Science, vol. 1, 2020.
[3] J. A. Fadhil, K. T. Wei, and K. S. Na, “Artificial intelligence for software engineering: An initial review on software bug detection and prediction,” Journal of Computer Science, vol. 16, no. 12, pp. 1709–1717, 2020.
[4] A. Kukkar, R. Mohana, Y. Kumar, A. Nayyar, M. Bilal, and K. S. Kwak, “Duplicate bug report detection and classification system based on deep learning technique,” IEEE Access, vol. 8, pp. 200 749–200 763, 2020.
[5] N. Ayesha and N. G. Yethiraj, “Review on code examination proficient system in software engineering by using machine learning approach,” in 2018 International Conference on Inventive Research in Computing Applications (ICIRCA). Coimbatore, India: IEEE, July 11–12, 2018, pp. 324–327.
[6] S. Mostafa, S. T. Cynthia, B. Roy, and D. Mondal, “Feature transformation for improved software bug detection and commit classification,” Journal of Systems and Software, vol. 219, pp. 1–13, 2025.
[7] A. D. Septiadi, M. A. W. Prasetyo, and G. E. A. Daffa, “Optimizing function-level source code classification using meta-trained CodeBERT in low-resource settings,” Journal of Applied Data Sciences, vol. 6, no. 3, pp. 2267–2280, 2025.
[8] C. Do Xuan, T. T. Luong, and M. C. Thanh, “Optimising source code vulnerability detection using deep learning and deep graph network,” Connection Science, vol. 37, no. 1, pp. 1–29, 2025.
[9] G. Giray, K. E. Bennin, O¨ . Ko¨ksal, O¨ . Babur, and B. Tekinerdogan, “On the use of deep learning in software defect prediction,” Journal of Systems and Software, vol. 195, pp. 1–26, 2023.
[10] F. Kumeno, “Software engineering challenges for machine learning applications: A literature review,” Intelligent Decision Technologies, vol. 13, no. 4, pp. 463–476, 2019.
[11] S. V¨ansk¨a, K. K. Kemell, T. Mikkonen, and P. Abrahamsson, “Continuous software engineering practices in AI/ML development past the narrow lens of MLOps: Adoption challenges,” EInformatica, vol. 18, no. 1, pp. 1–20, 2024.
[12] B. Imran, E. Wahyudi, S. Riadi, Z. Muahidin, S. Erniwati, and W. A. Wahyuni, “A comparative hybrid approach for Python bug detection using syntactic features, random forest, and neural network,” CommIT (Communication and Information Technology), vol. 19, no. 2, pp. 141–150, 2025.
[13] S. A. Alsaedi, A. Y. Noaman, A. A. A. Gad-Elrab, and F. E. Eassa, “Nature-based prediction model of bug reports based on ensemble machine learning model,” IEEE Access, vol. 11, pp. 63 916–63 931, 2023.
[14] A. Serban, K. Van der Blom, H. Hoos, and J. Visser, “Software engineering practices for machine learning–Adoption, effects, and team assessment,” Journal of Systems and Software, vol. 209, pp. 1–21, 2024.
[15] W. Albattah and M. Alzahrani, “Software defect prediction based on machine learning and deep learning techniques: An empirical approach,” AI, vol. 5, no. 4, pp. 1743–1758, 2024.
[16] S. Amershi et al., “Software engineering for machine learning: A case study,” in 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSESEIP). Montreal, QC, Canada: IEEE, May 25–31, 2019, pp. 291–300.
[17] L. Ma, H. Yang, J. Xu, Z. Yang, Q. Lao, and D. Yuan, “Code analysis with static application security testing for Python program,” Journal of Signal Processing Systems, vol. 94, no. 11, pp. 1169–1182, 2022.
[18] R. G. Hussain, K. C. Yow, and M. Gori, “Leveraging an enhanced CodeBERT-based model for multiclass software defect prediction via defect classification,” IEEE Access, vol. 13, pp. 24 383–24 397, 2025.
[19] Y. Lu, S. Ye, and L. Qi, “Codetranfix: A neural machine translation approach for context-aware Java program repair with CodeBERT,” Applied Sciences, vol. 15, no. 7, pp. 1–14, 2025.
[20] A. T. P. Nguyen and V. D. Hoang, “Development of code evaluation system based on abstract syntax tree,” Journal of Technical Education Science, vol. 19, no. Special Issue 01, pp. 15–24, 2024.
[21] S. D. Immaculate, M. F. Begam, and M. Floramary, “Software bug prediction using supervised machine learning algorithms,” in 2019 International Conference on Data Science and Communication (IconDSC). Bangalore, India: IEEE, March 1–2, 2019, pp. 1–7.
[22] R. Siva, K. S, B. Hariharan, and N. Premkumar, “Automatic software bug prediction using adaptive artificial jelly optimization with long shortterm memory,” Wireless Personal Communications, vol. 132, no. 3, pp. 1975–1998, 2023.
[23] Q. M. ul Haq, F. Arif, K. Aurangzeb, N. ul Ain, J. A. Khan, S. Rubab, and M. S. Anwar, “Identification of software bugs by analyzing natural language-based requirements using optimized deep learning features,” Computers, Materials & Continua, vol. 78, no. 3, pp. 4379–4397, 2024.
[24] S. T. Cynthia, B. Roy, and D. Mondal, “Feature transformation for improved software bug detection models,” in Proceedings of the 15th Innovations in Software Engineering Conference, 2022, pp. 1–10.
[25] J. Hu, L. Shen, and G. Sun, “Squeeze-andexcitation networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7132–7141.
[26] C. Tian et al., “A survey on deep learning fundamentals,” Artificial Intelligence Review, vol. 58, pp. 1–108, 2025.
[27] S. Danish, A. Sadeghi-Niaraki, S. U. Khan, L. M. Dang, L. Tightiz, and H. Moon, “A comprehensive survey of vision-language models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets,” Information Fusion, 2025.
[28] S. Vandenhende, S. Georgoulis, W. Van Gansbeke, M. Proesmans, D. Dai, and L. Van Gool, “Multi-task learning for dense prediction tasks: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, pp. 3614–3633, 2022.
[29] R. Wang and F. Liu, “A cross-attention gating mechanism-based multimodal feature fusion method for software defect prediction,” Applied Sciences, vol. 15, no. 20, pp. 1–28, 2025.
[30] S. Cong and Y. Zhou, “A review of convolutional neural network architectures and their optimizations,” Artificial Intelligence Review, vol. 56, no. 3, pp. 1905–1969, 2023.
[31] Y. Xie, J. Lin, H. Dong, L. Zhang, and Z. Wu, “Survey of code search based on deep learning,” ACM Transactions on Software Engineering and Methodology, vol. 33, no. 2, pp. 1–42, 2023.
[32] W. Ma, S. Liu, M. Zhao, X. Xie, W. Wang, Q. Hu, J. Zhang, and Y. Liu, “Unveiling code pretrained models: Investigating syntax and semantics capacities,” ACM Transactions on Software Engineering and Methodology, vol. 33, no. 7, pp. 1–29, 2024.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Rozali Ilham, Bahtiar Imran, Hasan Basri, Erfan Wahyudi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
a. Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License - Share Alike that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
b. Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
c. Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
Â
USER RIGHTS
All articles published Open Access will be immediately and permanently free for everyone to read and download. We are continuously working with our author communities to select the best choice of license options, currently being defined for this journal as follows: Creative Commons Attribution-Share Alike (CC BY-SA)

















