Execution-Feedback Code Repair Agents with Test-Guided Iterative Refinement
Keywords:
Code Agents, Automated Program Repair, Execution Feedback, Humaneval, Fault Localization, Iterative Refinement, Test-Guided RepairAbstract
Code agents often produce near-correct programs that fail executable checks. This study evaluates an execution-feedback repair agent (EFR) that converts test outcomes into structured evidence, localizes suspicious code, proposes bounded abstract-syntax-tree edits, and re-ranks remaining patches after failed attempts. All 164 HumanEval tasks were used. After canonical solutions passed the official checks, one reversible fault was introduced per task across six edit families. The 1,582 dynamically executed assertions were split into 1,260 feedback and 322 held-out assertions, and 1,470 localized proposals were executed. A task-grouped five-fold logistic ranker supplied a patch prior; coverage, exception compatibility, and test-score changes supplied runtime evidence. EFR repaired 83.54% of programs within five attempts and 87.80% within ten (95% Wilson interval, 81.91%-91.97%), compared with 82.32% for learned fixed-order search, 66.46% for coverage search, and 57.32% for random search. The ten-attempt policy averaged 2.77 attempts and 914.6 lexical tokens. Only one feedback-passing patch failed held-out verification. The results show that a learned prior improves early efficiency, while bounded execution and evidence updates convert local failures into verified repairs.
References
[1] M. Chen et al., "Evaluating Large Language Models Trained on Code," arXiv:2107.03374, 2021.
[2] Y. Li, "Findable then Explainable: Retrieval-Summary Integration for Code Intelligence on a Lightweight CodeSearchNet Subset," J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65-82, Jul. 2024, doi: 10.69987/JACS.2024.40706.
[3] J. Austin et al., "Program Synthesis with Large Language Models," arXiv:2108.07732, 2021.
[4] D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt, "Measuring Coding Challenge Competence With APPS," in Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 12621-12632.
[5] Q. Xin, "Hybrid Cloud Architecture for Efficient and Cost-Effective Large Language Model Deployment," J. Inf. Syst. Informatics, vol. 7, no. 3, pp. 2182-2195, Sep. 2025, doi: 10.51519/journalisi.v7i3.1170.
[6] Y. Li et al., "Competition-Level Code Generation with AlphaCode," Science, vol. 378, no. 6624, pp. 1092-1097, 2022.
[7] E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong, "CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis," in Proc. International Conference on Learning Representations, 2023.
[8] B. Chen, F. Zhang, A. Nguyen, D. Zan, Z. Lin, J.-G. Lou, and W. Chen, "CodeT: Code Generation with Generated Tests," arXiv:2207.10397, 2022.
[9] H. Zhang, "LLM-Driven CI Failure Diagnosis and Automated Repair: From GitHub Actions Logs to Patch Recommendation," J. Technol. Informatics Eng., vol. 4, no. 1, pp. 190-214, Apr. 2025, doi: 10.51903/jtie.v4i1.484.
[10] A. Ni, S. Iyer, D. Radev, V. Stoyanov, W.-T. Yih, S. Wang, and X. V. Lin, "LEVER: Learning to Verify Language-to-Code Generation with Execution," in Proc. 40th International Conference on Machine Learning, PMLR, vol. 202, 2023, pp. 26106-26128.
[11] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, "Intelligent Fault Analysis with AIOps Technology," J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94-100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.
[12] X. Chen, M. Lin, N. Scharli, and D. Zhou, "Teaching Large Language Models to Self-Debug," arXiv:2304.05128, 2023.
[13] J. Nie and D. Zheng, "Ambiguity-Aware HDFS Log Anomaly Detection with Retrieval-Augmented Failure Narratives and Selective Refusal," J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66-80, Jan. 2023, doi: 10.69987/JACS.2023.30105.
[14] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, "Reflexion: Language Agents with Verbal Reinforcement Learning," in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 8634-8652.
[15] Q. Xin, "Explaining OpenStack Failure-Injection Log Anomalies with Retrieved Normal Prototypes," Emerg. Inf. Sci. Technol., vol. 6, no. 2, pp. 125-146, Nov. 2025, doi: 10.18196/eist.v6i2.31232.
[16] A. Madaan et al., "Self-Refine: Iterative Refinement with Self-Feedback," in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 46534-46594.
[17] H. Zhang, "DriftGuard: Multi-Signal Drift Early Warning and Safe Re-Training/Rollback for CTR/CVR Models," J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24-40, Jul. 2023, doi: 10.69987/JACS.2023.30703.
[18] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, "ReAct: Synergizing Reasoning and Acting in Language Models," in Proc. International Conference on Learning Representations, 2023.
[19] J. Nie and D. Zheng, "Noisy-Neighbor-Aware VM Degradation Risk Modeling with Unsupervised Residual Fusion," J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112-123, Apr. 2024, doi: 10.69987/JACS.2024.40409.
[20] T. X. Olausson, J. P. Inala, C. Wang, J. Gao, and A. Solar-Lezama, "Is Self-Repair a Silver Bullet for Code Generation?" arXiv:2306.09896, 2023.
[21] W. Weimer, T. Nguyen, C. Le Goues, and S. Forrest, "Automatically Finding Patches Using Genetic Programming," in Proc. 31st International Conference on Software Engineering, 2009, pp. 364-374.
[22] C. Le Goues, M. Dewey-Vogt, S. Forrest, and W. Weimer, "A Systematic Study of Automated Program Repair: Fixing 55 out of 105 Bugs for $8 Each," in Proc. 34th International Conference on Software Engineering, 2012, pp. 3-13.
[23] H. D. T. Nguyen, D. Qi, A. Roychoudhury, and S. Chandra, "SemFix: Program Repair via Semantic Analysis," in Proc. 35th International Conference on Software Engineering, 2013, pp. 772-781.
[24] Y. Qi, X. Mao, Y. Lei, Z. Dai, and C. Wang, "The Strength of Random Search on Automated Program Repair," in Proc. 36th International Conference on Software Engineering, 2014, pp. 254-265.
[25] F. Long and M. Rinard, "Staged Program Repair with Condition Synthesis," in Proc. 10th Joint Meeting on Foundations of Software Engineering, 2015, pp. 166-178.
[26] F. Long and M. Rinard, "Automatic Patch Generation by Learning Correct Code," in Proc. 43rd ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages, 2016, pp. 298-312.
[27] S. Mechtaev, J. Yi, and A. Roychoudhury, "Angelix: Scalable Multiline Program Patch Synthesis via Symbolic Analysis," in Proc. 38th International Conference on Software Engineering, 2016, pp. 691-701.
[28] Y. Xiong, J. Wang, R. Yan, J. Zhang, S. Han, G. Huang, and L. Zhang, "Precise Condition Synthesis for Program Repair," in Proc. 39th International Conference on Software Engineering, 2017, pp. 416-426.
[29] M. Monperrus, "Automatic Software Repair: A Bibliography," ACM Computing Surveys, vol. 51, no. 1, Art. no. 17, 2018.
[30] M. Papadakis, M. Kintis, J. Zhang, Y. Jia, Y. Le Traon, and M. Harman, "Mutation Testing Advances: An Analysis and Survey," Advances in Computers, vol. 112, pp. 275-378, 2019.
[31] T. Lutellier, H. Pham, L. Pang, Y. Li, M. Wei, and L. Tan, "CoCoNuT: Combining Context-Aware Neural Translation Models Using Ensemble for Program Repair," in Proc. 29th ACM SIGSOFT International Symposium on Software Testing and Analysis, 2020, pp. 101-114.
[32] N. Jiang, T. Lutellier, and L. Tan, "CURE: Code-Aware Neural Machine Translation for Automatic Program Repair," in Proc. 43rd IEEE/ACM International Conference on Software Engineering, 2021, pp. 1161-1173.
[33] C. S. Xia and L. Zhang, "Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-Shot Learning," in Proc. 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2022, pp. 959-971.
[34] C. S. Xia, Y. Wei, and L. Zhang, "Automated Program Repair in the Era of Large Pre-Trained Language Models," in Proc. 45th IEEE/ACM International Conference on Software Engineering, 2023, pp. 1482-1494.
[35] Z. S. Zhong, R. Ma, and H. Zhao, "Human-Uncertainty Distillation for Calibrated Vision Models on CIFAR-10H," J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77-89, Feb. 2023, doi: 10.69987/JACS.2023.30206.
[36] Q. Xin, "Uncertainty-Aware Late Fusion for 3D Perception (Confidence Calibration + Fusion Rule Learning)," J. Technol. Informatics Eng., vol. 4, no. 1, pp. 215-238, Apr. 2025, doi: 10.51903/jtie.v4i1.485.
[37] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, "Cancer Image Classification Based on DenseNet Model," J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, Nov. 2020, doi: 10.1088/1742-6596/1651/1/012143.
[38] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, "Implementation of an AI-Based MRD Evaluation and Prediction Model for Multiple Myeloma," Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127-131, Jan. 2024, doi: 10.54097/zJ4MnbWW.
[39] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, "Autonomous Driving System Driven by Artificial Intelligence Perception Fusion," Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193-198, Feb. 2024, doi: 10.54097/e0b9ak47.
[40] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, "Optimization of Autonomous Driving Image Detection Based on RFAConv and Triplet Attention," Appl. Comput. Eng., vol. 77, no. 1, pp. 210-217, 2024, doi: 10.54254/2755-2721/77/2024MA0067.
[41] R. Ma, L. Zhang, and T. Song, "Computer-Vision-Informed Visual Explanation Cards for Autonomous-Driving Traffic-Sign Alerts: Localization, Classification, and Retrieved Evidence on GTSDB," Int. J. Graph. Des., vol. 3, no. 2, pp. 455-473, Oct. 2025, doi: 10.51903/ijgd.v3i2.3991.
[42] S. Zhou, Z. Li, and E. Wang, "Evidence-Grounded RAG for Tokenized Trade Receivable Disclosure QA under U.S. Capital Market Standards," J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41-57, Jul. 2023, doi: 10.69987/JACS.2023.30704.
[43] Z. Li, S. Zhou, and Z. Zhou, "Financial Risk Dashboard Design for Institutional RWA Investors: Visual Hierarchy, Chart Comprehension, and Explainability in FinChart-Bench," Int. J. Graph. Des., vol. 3, no. 1, pp. 196-210, May 2025, doi: 10.51903/ijgd.v3i1.3715.
[44] S. Zhou, Z. Li, and E. Wang, "Long-Document RAG for Contractual and Insurance Clause Analysis in Receivables RWA Structures," J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88-104, Aug. 2024, doi: 10.69987/JACS.2024.40810.
[45] H. Wang, Y. Ren, and X. Chang, "Layout-Aware Progressive PDF Rendering: AI Prioritization of PDF Slices to Reduce Time-to-Functional-First-Frame on FUNSD," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 425-446, Aug. 2025, doi: 10.51903/jtie.v4i2.523.
[46] B. Zhou, H. Wang, and X. Chang, "Distilling VMAF into an Edge-Deployable Quality Predictor: A Pilot Shot-Level Proxy with LLM-Ready Quality Tokens," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 447-463, Aug. 2025, doi: 10.51903/jtie.v4i2.522.
[47] R. Abreu, P. Zoeteweij, and A. J. C. van Gemund, "An Evaluation of Similarity Coefficients for Software Fault Localization," in Proc. 12th Pacific Rim International Symposium on Dependable Computing, 2006, pp. 39-46.
[48] J. A. Jones and M. J. Harrold, "Empirical Evaluation of the Tarantula Automatic Fault-Localization Technique," in Proc. 20th IEEE/ACM International Conference on Automated Software Engineering, 2005, pp. 273-282.
[49] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, "IoT Traffic Classification and Anomaly Detection Method Based on Deep Autoencoders," Appl. Comput. Eng., vol. 69, no. 1, pp. 64-70, Jul. 2024, doi: 10.54254/2755-2721/69/20241511.
[50] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, "Predictive Optimization of DDoS Attack Mitigation in Distributed Systems Using Machine Learning," Appl. Comput. Eng., vol. 64, no. 1, pp. 89-94, 2024, doi: 10.54254/2755-2721/64/20241350.
[51] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, "Enhancing Bank Credit Risk Management Using the C5.0 Decision Tree Algorithm," J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100-107, Nov. 2024, doi: 10.5281/zenodo.14032041.
[52] H. Zhou and K. Zhang, "News-Based Uncertainty and Macro-Market Fusion for VIX Direction Forecasting: Evidence from 2015-2024 FRED Panel," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 487-501, Aug. 2025, doi: 10.51903/jtie.v4i2.540.
[53] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, "Enhancing Financial Services through Big Data and AI-Driven Customer Insights and Risk Analysis," J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53-62, Jul. 2024, doi: 10.60087/jklst.vol3.n3.p53-62.
[54] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, "Predicting Enterprise Marketing Decision Making with Intelligent Data-Driven Approaches," J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12-19, Jun. 2024, doi: 10.5281/zenodo.11357252.
[55] R. Just, D. Jalali, and M. D. Ernst, "Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs," in Proc. International Symposium on Software Testing and Analysis, 2014, pp. 437-440.
[56] Z. S. Zhong and S. Ling, "Improved Theoretical Guarantee for Rank Aggregation via Spectral Method," Inf. Inference: J. IMA, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.
[57] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, "Intelligent Classification and Personalized Recommendation of E-Commerce Products Based on Machine Learning," Appl. Comput. Eng., vol. 64, no. 1, pp. 143-149, 2024, doi: 10.54254/2755-2721/64/20241365.
[58] H. Zhang, "Counterfactual Learning-to-Rank for Ads: Off-Policy Evaluation on the Open Bandit Dataset," J. Adv. Comput. Syst., vol. 5, no. 12, pp. 1-11, Dec. 2025, doi: 10.69987/JACS.2025.51201.
[59] Z. S. Zhong and S. Ling, "Uncertainty Quantification of Spectral Estimator and MLE for Orthogonal Group Synchronization," arXiv:2408.05944, Aug. 2024, doi: 10.48550/arXiv.2408.05944.
[60] H. Zhang, "Risk-Aware Budget-Constrained Auto-Bidding under First-Price RTB: A Distributional Constrained Deep Reinforcement Learning Framework," J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30-47, Jun. 2024, doi: 10.69987/JACS.2024.40603.
[61] Y. Lu, H. Zhou, and Y. Zhang, "A Constrained, Data-Driven Budgeting Framework Integrating Macro Demand Forecasting and Marketing Response Modeling," J. Technol. Informatics Eng., vol. 4, no. 3, pp. 493-520, Dec. 2025, doi: 10.51903/jtie.v4i3.466.
[62] L. Zhang, R. Ma, and P. Greg, "Digital-Twin Dispatching for Urban Mobility via Spatio-Temporal Transformers and Offline Reinforcement Learning," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 337-363, Aug. 2025, doi: 10.51903/jtie.v4i2.501.
[63] Z. S. Zhong, X. Pan, and Q. Lei, "Bridging Domains with Approximately Shared Features," in Proc. 28th Int. Conf. Artificial Intelligence and Statistics (AISTATS), PMLR, vol. 258, 2025, pp. 559-567.
[64] A. G. S. Raj et al., "Impact of Bilingual CS Education on Student Learning and Engagement in a Data Structures Course," in Proc. 19th Koli Calling Int. Conf. Comput. Educ. Res., 2019, pp. 1-10, doi: 10.1145/3364510.3364518.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Brian Thompson, Fan Zhang, Melissa Adams, Jun Zhao (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Creative Commons Attribution 4.0 International (CC BY 4.0).




This work is licensed under a