Committee-of-Judges: Multi-Agent Consensus for Reliable LLM Evaluation under Difficult Reasoning Tasks
Keywords:
LLM-As-A-Judge, Multi-Agent Evaluation, Pairwise Preference, Consensus, Debate, Calibration, Reasoning Verification, Code Execution, JudgebenchAbstract
Reliable pairwise evaluation of large language model responses is difficult when candidates contain long derivations, executable code, or closely matched conclusions. This study evaluates four protocols on the complete 620-pair JudgeBench corpus: a single judge, three-judge voting, globally weighted aggregation, and a task-aware arbiter activated only on disagreement. Three question-grouped, cross-validated agents modeled semantic contrast, question-response interaction with arithmetic checks, and structural evidence with corroboration and sandboxed code execution. Accuracy was 59.19% for the single judge, 61.61% for voting, 66.94% for weighting, and 67.42% for debate. Debate improved on the single judge by 8.23 points (McNemar p = 0.000218), but its 0.48-point advantage over weighting was not significant. Removing all objective verifiers reduced debate accuracy to 60.65%. Debate cost 3.62 normalized units because 61.77% of pairs entered arbitration, whereas weighting achieved nearly the same accuracy at 3.00 units. Heterogeneous evidence and calibrated aggregation therefore improved difficult-response evaluation, while unconditional extra deliberation offered limited value.
References
[1] T. B. Brown et al., “Language models are few-shot learners,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 1877–1901.
[2] Q. Xin, “Hybrid cloud architecture for efficient and cost-effective large language model deployment,” J. Inf. Syst. Informatics, vol. 7, no. 3, pp. 2182–2195, Sep. 2025, doi: 10.51519/journalisi.v7i3.1170.
[3] OpenAI, “GPT-4 technical report,” arXiv:2303.08774, 2023.
[4] L. Ouyang et al., “Training language models to follow instructions with human feedback,” in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 27730–27744.
[5] Y. Bai et al., “Constitutional AI: Harmlessness from AI feedback,” arXiv:2212.08073, 2022.
[6] N. Stiennon et al., “Learning to summarize with human feedback,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 3008–3021.
[7] P. F. Christiano et al., “Deep reinforcement learning from human preferences,” in Advances in Neural Information Processing Systems, vol. 30, 2017.
[8] A. G. S. Raj et al., “Impact of bilingual CS education on student learning and engagement in a data structures course,” in Proc. 19th Koli Calling Int. Conf. Comput. Educ. Res., Koli, Finland, 2019, pp. 1–10, doi: 10.1145/3364510.3364518.
[9] L. Zheng et al., “Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,” arXiv:2306.05685, 2023.
[10] Z. S. Zhong, R. Ma, and H. Zhao, “Human-uncertainty distillation for calibrated vision models on CIFAR-10H,” J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77–89, Feb. 2023, doi: 10.69987/JACS.2023.30206.
[11] Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu, “G-Eval: NLG evaluation using GPT-4 with better human alignment,” in Proc. 2023 Conf. Empirical Methods in Natural Language Processing, Singapore, 2023, pp. 2511–2522.
[12] J. Fu, S.-K. Ng, Z. Jiang, and P. Liu, “GPTScore: Evaluate as you desire,” arXiv:2302.04166, 2023.
[13] J. Nie and D. Zheng, “Ambiguity-aware HDFS log anomaly detection with retrieval-augmented failure narratives and selective refusal,” J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66–80, Jan. 2023, doi: 10.69987/JACS.2023.30105.
[14] P. Wang et al., “Large language models are not fair evaluators,” arXiv:2305.17926, 2023.
[15] C.-M. Chan, W. Chen, Y. Su, J. Yu, W. Xue, S. Zhang, J. Fu, and Z. Liu, “ChatEval: Towards better LLM-based evaluators through multi-agent debate,” arXiv:2308.07201, 2023.
[16] Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch, “Improving factuality and reasoning in language models through multiagent debate,” arXiv:2305.14325, 2023.
[17] G. Irving, P. Christiano, and D. Amodei, “AI safety via debate,” arXiv:1805.00899, 2018.
[18] Z. S. Zhong and S. Ling, “Improved theoretical guarantee for rank aggregation via spectral method,” Inf. Inference: J. IMA, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.
[19] J. Wei et al., “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 24824–24837.
[20] X. Wang et al., “Self-consistency improves chain of thought reasoning in language models,” in Proc. 11th Int. Conf. Learning Representations, 2023.
[21] K. Cobbe et al., “Training verifiers to solve math word problems,” arXiv:2110.14168, 2021.
[22] H. Lightman et al., “Let’s verify step by step,” arXiv:2305.20050, 2023.
[23] D. Hendrycks et al., “Measuring massive multitask language understanding,” in Proc. 9th Int. Conf. Learning Representations, 2021.
[24] A. Srivastava et al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” arXiv:2206.04615, 2022.
[25] M. Chen et al., “Evaluating large language models trained on code,” arXiv:2107.03374, 2021.
[26] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, “Cancer image classification based on DenseNet model,” J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, Nov. 2020, doi: 10.1088/1742-6596/1651/1/012143.
[27] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, “Implementation of an AI-based MRD evaluation and prediction model for multiple myeloma,” Frontiers Comput. Intell. Syst., vol. 6, no. 3, pp. 127–131, Jan. 2024, doi: 10.54097/zJ4MnbWW.
[28] Y. Li, “Findable then explainable: Retrieval-summary integration for code intelligence on a lightweight CodeSearchNet subset,” J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65–82, Jul. 2024, doi: 10.69987/JACS.2024.40706.
[29] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, “Autonomous driving system driven by artificial intelligence perception fusion,” Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193–198, Feb. 2024, doi: 10.54097/e0b9ak47.
[30] Q. Xin, “Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning),” J. Technol. Informatics Eng., vol. 4, no. 1, pp. 215–238, Apr. 2025, doi: 10.51903/jtie.v4i1.485.
[31] Z. S. Zhong and S. Ling, “Uncertainty quantification of spectral estimator and MLE for orthogonal group synchronization,” arXiv:2408.05944, Aug. 2024.
[32] H. Zhang, “DriftGuard: Multi-signal drift early warning and safe re-training/rollback for CTR/CVR models,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24–40, Jul. 2023, doi: 10.69987/JACS.2023.30703.
[33] S. Zhou, Z. Li, and E. Wang, “Evidence-grounded RAG for tokenized trade receivable disclosure QA under U.S. capital market standards,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41–57, Jul. 2023, doi: 10.69987/JACS.2023.30704.
[34] Q. Xin, “Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes,” Emerg. Inf. Sci. Technol., vol. 6, no. 2, pp. 125–146, Nov. 2025, doi: 10.18196/eist.v6i2.31232.
[35] S. Zhou, Z. Li, and E. Wang, “Long-document RAG for contractual and insurance clause analysis in receivables RWA structures,” J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88–104, Aug. 2024, doi: 10.69987/JACS.2024.40810.
[36] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, “Intelligent fault analysis with AIOps technology,” J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94–100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.
[37] H. Zhang, “LLM-driven CI failure diagnosis and automated repair: From GitHub Actions logs to patch recommendation,” J. Technol. Informatics Eng., vol. 4, no. 1, pp. 190–214, Feb. 2025, doi: 10.51903/jtie.v4i1.484.
[38] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, “Intelligent classification and personalized recommendation of e-commerce products based on machine learning,” Appl. Comput. Eng., vol. 64, no. 1, pp. 143–149, May 2024, doi: 10.54254/2755-2721/64/20241365.
[39] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, “Predicting enterprise marketing decision making with intelligent data-driven approaches,” J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12–19, Jun. 2024, doi: 10.5281/zenodo.11357252.
[40] L. Zhang, R. Ma, and P. Greg, “Digital-twin dispatching for urban mobility via spatio-temporal transformers and offline reinforcement learning,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 337–363, Aug. 2025, doi: 10.51903/jtie.v4i2.501.
[41] H. Zhou and K. Zhang, “News-based uncertainty and macro-market fusion for VIX direction forecasting: Evidence from 2015–2024 FRED panel,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 487–501, Aug. 2025, doi: 10.51903/jtie.v4i2.540.
[42] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, “IoT traffic classification and anomaly detection method based on deep autoencoders,” Appl. Comput. Eng., vol. 69, no. 1, pp. 64–70, Jul. 2024, doi: 10.54254/2755-2721/69/20241511.
[43] J. Nie and D. Zheng, “Noisy-neighbor-aware VM degradation risk modeling with unsupervised residual fusion,” J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112–123, Apr. 2024, doi: 10.69987/JACS.2024.40409.
[44] Z. S. Zhong, X. Pan, and Q. Lei, “Bridging domains with approximately shared features,” in Proc. 28th Int. Conf. Artificial Intelligence and Statistics (AISTATS), PMLR, vol. 258, 2025, pp. 559–567.
[45] T. G. Dietterich, “Ensemble methods in machine learning,” in Multiple Classifier Systems, Berlin, Germany: Springer, 2000, pp. 1–15.
[46] L. I. Kuncheva, Combining Pattern Classifiers: Methods and Algorithms. Hoboken, NJ, USA: Wiley, 2004.
[47] A. P. Dawid and A. M. Skene, “Maximum likelihood estimation of observer error-rates using the EM algorithm,” J. Royal Stat. Soc. Series C, vol. 28, no. 1, pp. 20–28, 1979.
[48] J. Cohen, “A coefficient of agreement for nominal scales,” Educ. Psychol. Meas., vol. 20, no. 1, pp. 37–46, 1960.
[49] J. L. Fleiss, “Measuring nominal scale agreement among many raters,” Psychol. Bull., vol. 76, no. 5, pp. 378–382, 1971.
[50] G. W. Brier, “Verification of forecasts expressed in terms of probability,” Mon. Weather Rev., vol. 78, no. 1, pp. 1–3, 1950.
[51] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proc. 34th Int. Conf. Machine Learning, 2017, pp. 1321–1330.
[52] B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York, NY, USA: Chapman & Hall, 1993.
[53] Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, no. 2, pp. 153–157, 1947.
[54] J. Demšar, “Statistical comparisons of classifiers over multiple data sets,” J. Mach. Learn. Res., vol. 7, pp. 1–30, 2006.
[55] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, “Enhancing bank credit risk management using the C5.0 decision tree algorithm,” J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100–107, Nov. 2024.
[56] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, “Optimization of autonomous driving image detection based on RFAConv and triplet attention,” Appl. Comput. Eng., vol. 77, no. 1, pp. 210–217, Jul. 2024, doi: 10.54254/2755-2721/77/2024MA0067.
[57] H. Zhang, “Risk-aware budget-constrained auto-bidding under first-price RTB: A distributional constrained deep reinforcement learning framework,” J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30–47, Jun. 2024, doi: 10.69987/JACS.2024.40603.
[58] B. Zhou, H. Wang, and X. Chang, “Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 447–463, Aug. 2025, doi: 10.51903/jtie.v4i2.522.
[59] H. Wang, Y. Ren, and X. Chang, “Layout-aware progressive PDF rendering: AI prioritization of PDF slices to reduce time-to-functional-first-frame on FUNSD,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 425–446, Aug. 2025, doi: 10.51903/jtie.v4i2.523.
[60] Y. Lu, H. Zhou, and Y. Zhang, “A constrained, data-driven budgeting framework integrating macro demand forecasting and marketing response modeling,” J. Technol. Informatics Eng., vol. 4, no. 3, pp. 493–520, Dec. 2025, doi: 10.51903/jtie.v4i3.466.
[61] H. Zhang, “Counterfactual learning-to-rank for ads: Off-policy evaluation on the Open Bandit Dataset,” J. Adv. Comput. Syst., vol. 5, no. 12, pp. 1–11, Dec. 2025, doi: 10.69987/JACS.2025.51201.
[62] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, “Enhancing financial services through big data and AI-driven customer insights and risk analysis,” J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53–62, Jul. 2024, doi: 10.60087/jklst.vol3.n3.p53-62.
[63] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, “Predictive optimization of DDoS attack mitigation in distributed systems using machine learning,” Appl. Comput. Eng., vol. 64, no. 1, pp. 94–99, May 2024, doi: 10.54254/2755-2721/64/20241350.
[64] Z. Li, S. Zhou, and Z. Zhou, “Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench,” Int. J. Graph. Des., vol. 3, no. 1, pp. 196–210, May 2025, doi: 10.51903/ijgd.v3i1.3715.
[65] R. Ma, L. Zhang, and T. Song, “Computer-vision-informed visual explanation cards for autonomous-driving traffic-sign alerts: Localization, classification, and retrieved evidence on GTSDB,” Int. J. Graph. Des., vol. 3, no. 2, pp. 455–473, Oct. 2025, doi: 10.51903/ijgd.v3i2.3991.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Emily Wright, Yan Li, Anthony Hall, Bin Zhao (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Creative Commons Attribution 4.0 International (CC BY 4.0).




This work is licensed under a