Verifier-Guided Adaptive Test-Time Scaling for Cost-Accurate Mathematical Reasoning

Authors

Keywords:

Mathematical Reasoning, Test-Time Scaling, Verifier, Adaptive Computation, Best-Of-N, Confidence Calibration, GSM8K, Cost-Aware Inference

Abstract

Inference-time sampling improves mathematical reasoning only when additional candidates are useful and the selector can identify them. This paper presents Verifier-Guided Adaptive Test-Time Scaling (VATS), a cost-aware framework that allocates a nested candidate budget to each problem, ranks candidate answers with a learned verifier, calibrates checkpoint-level correctness, and stops once the selected answer is stable and sufficiently confident. The empirical study uses all 7,473 GSM8K training problems and all 1,319 test problems. A trace-conditioned program sampler retrieves arithmetic programs from the training rationales, transfers them to new number mentions, adds cue-ranked exact-arithmetic programs, and exposes candidate prefixes at budgets of 1, 2, 4, 8, 16, 32, and 64. A 32-feature histogram gradient-boosting verifier raises exact-match accuracy at the largest budget from 4.78% for proposal-score selection and 6.22% for answer voting to 8.19%. The difference over proposal selection is 3.41 percentage points (95% bootstrap interval 1.74 to 5.00; exact McNemar p = 4.53e-5). The development-selected adaptive policy reaches 5.76% accuracy with 6.14 actual candidate evaluations on average, saving 89.59% of the computation used by fixed N = 64. A descriptive high-accuracy Pareto point reaches 8.26% with 45.08 evaluations, preserving fixed-budget accuracy while saving 23.63% of measured candidate scoring. Oracle pass@64 is 43.29%, showing that candidate generation and answer selection remain separate bottlenecks. The results establish a transparent accuracy-compute frontier and show that calibrated stopping reduces unnecessary inference without concealing the limits of the underlying reasoning policy.

References

[1] K. Cobbe et al., "Training Verifiers to Solve Math Word Problems," arXiv:2110.14168, 2021.

[2] J. Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models," in Advances in Neural Information Processing Systems, vol. 35, pp. 24824-24837, 2022.

[3] X. Wang et al., "Self-Consistency Improves Chain of Thought Reasoning in Language Models," in Proc. International Conference on Learning Representations, 2023.

[4] T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, "Large Language Models Are Zero-Shot Reasoners," in Advances in Neural Information Processing Systems, vol. 35, pp. 22199-22213, 2022.

[5] A. Lewkowycz et al., "Solving Quantitative Reasoning Problems with Language Models," in Advances in Neural Information Processing Systems, vol. 35, pp. 3843-3857, 2022.

[6] H. Lightman et al., "Let's Verify Step by Step," arXiv:2305.20050, 2023.

[7] J. Uesato et al., "Solving Math Word Problems with Process- and Outcome-Based Feedback," arXiv:2211.14275, 2022.

[8] E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman, "STaR: Bootstrapping Reasoning with Reasoning," in Advances in Neural Information Processing Systems, vol. 35, pp. 15476-15488, 2022.

[9] M. Nye et al., "Show Your Work: Scratchpads for Intermediate Computation with Language Models," arXiv:2112.00114, 2021.

[10] W. Chen, X. Ma, X. Wang, and W. W. Cohen, "Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks," Transactions on Machine Learning Research, 2023.

[11] L. Gao et al., "PAL: Program-Aided Language Models," in Proc. 40th International Conference on Machine Learning, 2023.

[12] D. Zhou et al., "Least-to-Most Prompting Enables Complex Reasoning in Large Language Models," in Proc. International Conference on Learning Representations, 2023.

[13] S. Yao et al., "Tree of Thoughts: Deliberate Problem Solving with Large Language Models," in Advances in Neural Information Processing Systems, vol. 36, pp. 11809-11822, 2023.

[14] A. Madaan et al., "Self-Refine: Iterative Refinement with Self-Feedback," in Advances in Neural Information Processing Systems, vol. 36, 2023.

[15] N. Shinn et al., "Reflexion: Language Agents with Verbal Reinforcement Learning," in Advances in Neural Information Processing Systems, vol. 36, 2023.

[16] Z. Zhang, A. Zhang, M. Li, and A. Smola, "Automatic Chain of Thought Prompting in Large Language Models," in Proc. International Conference on Learning Representations, 2023.

[17] T. Khot et al., "Decomposed Prompting: A Modular Approach for Solving Complex Tasks," in Proc. International Conference on Learning Representations, 2023.

[18] D. Hendrycks et al., "Measuring Mathematical Problem Solving with the MATH Dataset," in Advances in Neural Information Processing Systems, vol. 34, 2021.

[19] T. B. Brown et al., "Language Models Are Few-Shot Learners," in Advances in Neural Information Processing Systems, vol. 33, pp. 1877-1901, 2020.

[20] A. Vaswani et al., "Attention Is All You Need," in Advances in Neural Information Processing Systems, vol. 30, 2017.

[21] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, "On Calibration of Modern Neural Networks," in Proc. 34th International Conference on Machine Learning, vol. 70, pp. 1321-1330, 2017.

[22] A. Niculescu-Mizil and R. Caruana, "Predicting Good Probabilities with Supervised Learning," in Proc. 22nd International Conference on Machine Learning, pp. 625-632, 2005.

[23] A. Graves, "Adaptive Computation Time for Recurrent Neural Networks," arXiv:1603.08983, 2016.

[24] A. Banino, J. Balaguer, and C. Blundell, "PonderNet: Learning to Ponder," arXiv:2107.05407, 2021.

[25] A. Wald, Sequential Analysis. New York, NY, USA: Wiley, 1947.

[26] J. C. Platt, "Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods," in Advances in Large Margin Classifiers, A. J. Smola, P. L. Bartlett, B. Schoelkopf, and D. Schuurmans, Eds. Cambridge, MA, USA: MIT Press, pp. 61-74, 1999.

[27] Z. S. Zhong, R. Ma, and H. Zhao, “Human-Uncertainty Distillation for Calibrated Vision Models on CIFAR-10H,” J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77–89, Feb. 2023, doi: 10.69987/JACS.2023.30206.

[28] H. Zhang, “Risk-Aware Budget-Constrained Auto-Bidding under First-Price RTB: A Distributional Constrained Deep Reinforcement Learning Framework,” J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30–47, Jun. 2024, doi: 10.69987/JACS.2024.40603.

[29] S. Zhou, Z. Li, and E. Wang, “Evidence-Grounded RAG for Tokenized Trade Receivable Disclosure QA under U.S. Capital Market Standards,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41–57, Jul. 2023, doi: 10.69987/JACS.2023.30704.

[30] Y. Li, “Findable then Explainable: Retrieval–Summary Integration for Code Intelligence on a Lightweight CodeSearchNet Subset,” J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65–82, Jul. 2024, doi: 10.69987/JACS.2024.40706.

[31] S. Zhou, Z. Li, and E. Wang, “Long-Document RAG for Contractual and Insurance Clause Analysis in Receivables RWA Structures,” J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88–104, Aug. 2024, doi: 10.69987/JACS.2024.40810.

[32] Z. S. Zhong and S. Ling, “Improved Theoretical Guarantee for Rank Aggregation via Spectral Method,” Inf. Inference: J. IMA, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.

[33] H. Zhang, “DriftGuard: Multi-Signal Drift Early Warning and Safe Re-Training/Rollback for CTR/CVR Models,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24–40, Jul. 2023, doi: 10.69987/JACS.2023.30703.

[34] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, “Intelligent Classification and Personalized Recommendation of E-Commerce Products Based on Machine Learning,” in Proc. 6th Int. Conf. Computing and Data Science, 2024, pp. 143–149, doi: 10.54254/2755-2721/64/20241365.

[35] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, “Predicting Enterprise Marketing Decision Making with Intelligent Data-Driven Approaches,” J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12–19, Jun. 2024, doi: 10.5281/zenodo.11357252.

[36] J. Nie and D. Zheng, “Ambiguity-Aware HDFS Log Anomaly Detection with Retrieval-Augmented Failure Narratives and Selective Refusal,” J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66–80, Jan. 2023, doi: 10.69987/JACS.2023.30105.

[37] Z. S. Zhong and S. Ling, “Uncertainty Quantification of Spectral Estimator and MLE for Orthogonal Group Synchronization,” arXiv:2408.05944, Aug. 2024.

[38] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, “Enhancing Financial Services through Big Data and AI-Driven Customer Insights and Risk Analysis,” J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53–62, Jul. 2024, doi: 10.60087/jklst.vol3.n3.p53-62.

[39] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, “Enhancing Bank Credit Risk Management Using the C5.0 Decision Tree Algorithm,” J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100–107, Nov. 2024, doi: 10.5281/zenodo.14032041.

[40] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, “Intelligent Fault Analysis with AIOps Technology,” J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94–100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.

[41] J. Nie and D. Zheng, “Noisy-Neighbor-Aware VM Degradation Risk Modeling with Unsupervised Residual Fusion,” J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112–123, Apr. 2024, doi: 10.69987/JACS.2024.40409.

[42] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, “Predictive Optimization of DDoS Attack Mitigation in Distributed Systems Using Machine Learning,” in Proc. 6th Int. Conf. Computing and Data Science, 2024, pp. 89–94, doi: 10.54254/2755-2721/64/20241350.

[43] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, “IoT Traffic Classification and Anomaly Detection Method Based on Deep Autoencoders,” in Proc. 6th Int. Conf. Computing and Data Science, 2024, pp. 64–70, doi: 10.54254/2755-2721/69/20241511.

[44] A. G. S. Raj et al., “Impact of Bilingual CS Education on Student Learning and Engagement in a Data Structures Course,” in Proc. 19th Koli Calling Int. Conf. Computing Education Research, 2019, pp. 1–10, doi: 10.1145/3364510.3364518.

[45] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, “Cancer Image Classification Based on DenseNet Model,” J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, Nov. 2020, doi: 10.1088/1742-6596/1651/1/012143.

[46] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, “Implementation of an AI-Based MRD Evaluation and Prediction Model for Multiple Myeloma,” Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127–131, Jan. 2024, doi: 10.54097/zJ4MnbWW.

[47] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, “Autonomous Driving System Driven by Artificial Intelligence Perception Fusion,” Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193–198, Feb. 2024, doi: 10.54097/e0b9ak47.

[48] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, “Optimization of Autonomous Driving Image Detection Based on RFAConv and Triplet Attention,” Appl. Comput. Eng., vol. 77, no. 1, pp. 210–217, Jul. 2024, doi: 10.54254/2755-2721/77/2024MA0067.

Published

2024-12-31

How to Cite

Scott, J., Chen, T., Young, B., Zhao, L., & Reed, E. (2024). Verifier-Guided Adaptive Test-Time Scaling for Cost-Accurate Mathematical Reasoning. Journal of Information Technology and Informatics Engineering, 1(1), 01-11. https://journal.jci.co.id/jitie/article/view/816