Multi-Agent Debate with Disagreement-Aware Arbitration for Mathematical Reasoning
Keywords:
Multi-Agent Reasoning, Debate, Mathematical Reasoning, Disagreement, Arbitration, GSM8K, Self-Consistency, Program InductionAbstract
Multi-agent reasoning promises to improve mathematical problem solving by exposing disagreement before a final answer is selected. This study isolates that coordination problem in a controlled, fully executable setting and evaluates four inference policies on the complete 1,319-problem GSM8K test split: a greedy single chain of thought, lexical self-consistency, three-agent debate, and debate followed by a disagreement-aware arbiter. Executable arithmetic programs were induced from the annotated GSM8K training rationales; 6,999 of 7,473 training problems yielded programs that exactly reconstructed their gold answers. Three retrieval-program agents used raw lexical features, number-normalized structural features, and operation-cue features. The arbiter was a held-out logistic model trained on agreement, support, similarity, trace complexity, and numeric-plausibility features. Initial agents disagreed on 97.12% of test problems. Single-chain and self-consistent inference each solved 27 problems (2.05%). One debate round solved 35 (2.65%), while disagreement-aware arbitration solved 91 (6.90%). The arbiter improved accuracy over the single-chain baseline by 4.85 percentage points; the paired bootstrap 95% interval was 3.56 to 6.29 points and the exact McNemar p-value was 7.37 x 10^-13. Mean token-equivalent costs were 77.6, 725.8, 265.6, and 1,015.0, respectively. Candidate-generation failure accounted for 88.93% of test cases, establishing that arbitration recovers useful minority and extended candidates but does not replace a strong proposer.
References
[1] A. Patel, S. Bhattamishra, and N. Goyal, "Are NLP models really able to solve simple math word problems?" in Proc. NAACL-HLT, pp. 2080-2094, 2021.
[2] W. Ling, D. Yogatama, C. Dyer, and P. Blunsom, "Program induction by rationale generation: Learning to solve and explain algebraic word problems," in Proc. ACL, pp. 158-167, 2017.
[3] R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi, "MAWPS: A math word problem repository," in Proc. NAACL-HLT, pp. 1152-1157, 2016.
[4] S.-Y. Miao, C.-C. Liang, and K.-Y. Su, "A diverse corpus for evaluating and developing English math word problem solvers," in Proc. ACL, pp. 975-984, 2020.
[5] M. J. Hosseini, H. Hajishirzi, O. Etzioni, and N. Kushman, "Learning to solve arithmetic word problems with verb categorization," in Proc. EMNLP, pp. 523-533, 2014.
[6] J. Nie and D. Zheng, "Ambiguity-Aware HDFS Log Anomaly Detection with Retrieval-Augmented Failure Narratives and Selective Refusal," J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66-80, Jan. 2023, doi: 10.69987/JACS.2023.30105.
[7] S. Zhou, Z. Li, and E. Wang, "Evidence-Grounded RAG for Tokenized Trade Receivable Disclosure QA under U.S. Capital Market Standards," J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41-57, Jul. 2023, doi: 10.69987/JACS.2023.30704.
[8] T. B. Brown et al., "Language models are few-shot learners," in Advances in Neural Information Processing Systems, vol. 33, pp. 1877-1901, 2020.
[9] J. Wei et al., "Chain-of-thought prompting elicits reasoning in large language models," in Advances in Neural Information Processing Systems, vol. 35, pp. 24824-24837, 2022.
[10] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, "Large language models are zero-shot reasoners," in Advances in Neural Information Processing Systems, vol. 35, pp. 22199-22213, 2022.
[11] D. Zhou et al., "Least-to-most prompting enables complex reasoning in language models," in Proc. Int. Conf. Learning Representations, 2023.
[12] Z. S. Zhong, R. Ma, and H. Zhao, "Human-Uncertainty Distillation for Calibrated Vision Models on CIFAR-10H," J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77-89, Feb. 2023, doi: 10.69987/JACS.2023.30206.
[13] K. Cobbe et al., "Training verifiers to solve math word problems," arXiv:2110.14168, 2021.
[14] X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, and D. Zhou, "Self-consistency improves chain of thought reasoning in language models," in Proc. Int. Conf. Learning Representations, 2023.
[15] Z. S. Zhong and S. Ling, "Improved theoretical guarantee for rank aggregation via spectral method," Inf. Inference, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.
[16] E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman, "STaR: Bootstrapping reasoning with reasoning," in Advances in Neural Information Processing Systems, vol. 35, pp. 15476-15488, 2022.
[17] L. Gao et al., "PAL: Program-aided language models," arXiv:2211.10435, 2022.
[18] W. Chen, X. Ma, X. Wang, and W. W. Cohen, "Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks," arXiv:2211.12588, 2022.
[19] S. Yao et al., "Tree of thoughts: Deliberate problem solving with large language models," in Advances in Neural Information Processing Systems, vol. 36, pp. 11809-11822, 2023.
[20] A. Madaan et al., "Self-Refine: Iterative refinement with self-feedback," in Advances in Neural Information Processing Systems, vol. 36, pp. 46534-46594, 2023.
[21] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, "Reflexion: Language agents with verbal reinforcement learning," in Advances in Neural Information Processing Systems, vol. 36, pp. 8634-8652, 2023.
[22] J. He-Yueya, G. Poesia, R. E. Wang, and N. D. Goodman, "Solving math word problems by combining language models with symbolic solvers," arXiv:2304.09102, 2023.
[23] Y. Li, "Findable then Explainable: Retrieval-Summary Integration for Code Intelligence on a Lightweight CodeSearchNet Subset," J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65-82, Jul. 2024, doi: 10.69987/JACS.2024.40706.
[24] H. Lightman et al., "Let's verify step by step," arXiv:2305.20050, 2023.
[25] J. Uesato et al., "Solving math word problems with process- and outcome-based feedback," arXiv:2211.14275, 2022.
[26] A. Lewkowycz et al., "Solving quantitative reasoning problems with language models," arXiv:2206.14858, 2022.
[27] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, "Enhancing Bank Credit Risk Management Using the C5.0 Decision Tree Algorithm," J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100-107, Nov. 2024, doi: 10.5281/zenodo.14032041.
[28] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, "Cancer image classification based on DenseNet model," J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, 2020, doi: 10.1088/1742-6596/1651/1/012143.
[29] G. Irving, P. Christiano, and D. Amodei, "AI safety via debate," arXiv:1805.00899, 2018.
[30] J. Brown-Cohen, G. Irving, and G. Piliouras, "Scalable AI safety via doubly-efficient debate," arXiv:2311.14125, 2023.
[31] Y. Du et al., "Improving factuality and reasoning in language models through multiagent debate," arXiv:2305.14325, 2023.
[32] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, "Autonomous Driving System Driven by Artificial Intelligence Perception Fusion," Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193-198, Feb. 2024, doi: 10.54097/e0b9ak47.
[33] J. Huang et al., "Large language models cannot self-correct reasoning yet," arXiv:2310.01798, 2023.
[34] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, "Intelligent Fault Analysis with AIOps Technology," J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94-100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.
[35] S. Zhou, Z. Li, and E. Wang, "Long-Document RAG for Contractual and Insurance Clause Analysis in Receivables RWA Structures," J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88-104, Aug. 2024, doi: 10.69987/JACS.2024.40810.
[36] T. G. Dietterich, "Ensemble methods in machine learning," in Multiple Classifier Systems, LNCS 1857, pp. 1-15, 2000.
[37] L. I. Kuncheva and C. J. Whitaker, "Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy," Mach. Learn., vol. 51, pp. 181-207, 2003.
[38] A. P. Dawid and A. M. Skene, "Maximum likelihood estimation of observer error-rates using the EM algorithm," J. Roy. Stat. Soc. C, vol. 28, no. 1, pp. 20-28, 1979.
[39] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, "IoT traffic classification and anomaly detection method based on deep autoencoders," Appl. Comput. Eng., vol. 69, no. 1, pp. 64-70, Jul. 2024, doi: 10.54254/2755-2721/69/20241511.
[40] J. Nie and D. Zheng, "Noisy-Neighbor-Aware VM Degradation Risk Modeling with Unsupervised Residual Fusion," J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112-123, Apr. 2024, doi: 10.69987/JACS.2024.40409.
[41] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, "Predictive optimization of DDoS attack mitigation in distributed systems using machine learning," Appl. Comput. Eng., vol. 64, no. 1, pp. 94-99, May 2024, doi: 10.54254/2755-2721/64/20241350.
[42] H. Zhang, "Risk-Aware Budget-Constrained Auto-Bidding under First-Price RTB: A Distributional Constrained Deep Reinforcement Learning Framework," J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30-47, Jun. 2024, doi: 10.69987/JACS.2024.40603.
[43] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, "Intelligent classification and personalized recommendation of e-commerce products based on machine learning," in Proc. 6th Int. Conf. Computing and Data Science, 2024, doi: 10.54254/2755-2721/64/20241365.
[44] H. Zhang, "DriftGuard: Multi-Signal Drift Early Warning and Safe Re-Training/Rollback for CTR/CVR Models," J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24-40, Jul. 2023, doi: 10.69987/JACS.2023.30703.
[45] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, "Enhancing Financial Services through Big Data and AI-Driven Customer Insights and Risk Analysis," J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53-62, 2024, doi: 10.60087/jklst.vol3.n3.p53-62.
[46] Z. S. Zhong and S. Ling, "Uncertainty quantification of spectral estimator and MLE for orthogonal group synchronization," arXiv:2408.05944, Aug. 2024.
[47] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, "Optimization of autonomous driving image detection based on RFAConv and triplet attention," Appl. Comput. Eng., vol. 77, no. 1, pp. 210-217, Jul. 2024, doi: 10.54254/2755-2721/77/2024MA0067.
[48] A. G. S. Raj, H. Zhang, V. Abhyankar, S. Mukherjee, E. Zhang, J. Williams, R. Halverson, and J. M. Patel, "Impact of Bilingual CS Education on Student Learning and Engagement in a Data Structures Course," in Proc. 19th Koli Calling Int. Conf. Comput. Educ. Res., 2019, Art. no. 24, pp. 1-10, doi: 10.1145/3364510.3364518.
[49] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, "Predicting Enterprise Marketing Decision Making with Intelligent Data-Driven Approaches," J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12-19, Jun. 2024, doi: 10.5281/zenodo.11357252.
[50] D. Hendrycks et al., "Measuring mathematical problem solving with the MATH dataset," in Proc. NeurIPS Datasets and Benchmarks Track, 2021.
[51] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, "Implementation of an AI-based MRD Evaluation and Prediction Model for Multiple Myeloma," Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127-131, Jan. 2024, doi: 10.54097/zJ4MnbWW.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 Ming Wang, Sarah Collins, Joshua Bell (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Creative Commons Attribution 4.0 International (CC BY 4.0).




This work is licensed under a