Schema-Aware Retrieval-Augmented Tool-Use Agents for API Selection, Planning, and Calling: An Empirical Study on API-Bank

Authors

Keywords:

Tool-Use Agents, Function Calling, API Retrieval, Agent Planning, Schema Validation, API-Bank, Retrieval-Augmented Generation

Abstract

Tool-use agents must identify the correct application programming interface, construct a valid call, and preserve execution state across multi-step dialogue. This study evaluates a retrieval-augmented architecture that separates those responsibilities into global retrieval, trajectory-local candidate formation, schema-aware planning, typed argument binding, and execution validation. The experiments use 1,012 API-Bank records drawn from 261 dialogue files: 534 API-call targets and 478 response targets spanning 50 tools, 47 parameter names, and three difficulty levels. All selection models are evaluated with five-fold StratifiedGroupKFold splits by dialogue file, which prevents calls from the same trajectory from appearing in both training and test folds. Global BM25 and latent-semantic embedding retrieval achieve 41.57% and 41.01% Top-1 accuracy. Restricting retrieval to tools exposed by the active trajectory raises Top-1 accuracy to 64.98%-68.16%. The proposed SchemaPlan-RAC pairwise planner reaches 95.32% Top-1, 99.44% Top-3, and 0.973 mean reciprocal rank. Coupled with a deterministic schema-aware argument agent, it obtains 92.26% parameter-name F1, 74.97% argument-value accuracy, 64.42% strict call exact match, and 88.20% schema-executable accuracy. The planner improvement over the strongest manifest embedding baseline is significant under an exact McNemar test (p < 10^-38). Ablations show that trajectory candidates, schema constraints, and transition state contribute sequential gains. Error analysis identifies free-text values and temporal normalization as the dominant remaining failures.

References

[1] M. Li et al., “API-Bank: A comprehensive benchmark for tool-augmented LLMs,” in Proc. EMNLP, 2023, pp. 3102–3116.

[2] J. Nie and D. Zheng, “Ambiguity-aware HDFS log anomaly detection with retrieval-augmented failure narratives and selective refusal,” J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66–80, Jan. 2023, doi: 10.69987/JACS.2023.30105.

[3] S. Zhou, Z. Li, and E. Wang, “Evidence-grounded RAG for tokenized trade receivable disclosure QA under U.S. capital market standards,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41–57, Jul. 2023, doi: 10.69987/JACS.2023.30704.

[4] T. Schick et al., “Toolformer: Language models can teach themselves to use tools,” in Adv. Neural Inf. Process. Syst., vol. 36, 2023.

[5] S. Yao et al., “ReAct: Synergizing reasoning and acting in language models,” in Proc. ICLR, 2023.

[6] S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez, “Gorilla: Large language model connected with massive APIs,” arXiv:2305.15334, 2023.

[7] Y. Qin et al., “ToolLLM: Facilitating large language models to master 16000+ real-world APIs,” arXiv:2307.16789, 2023.

[8] Y. Song et al., “RestGPT: Connecting large language models with real-world RESTful APIs,” arXiv:2306.06624, 2023.

[9] Y. Shen et al., “HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face,” in Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 38154–38180.

[10] B. Paranjape et al., “ART: Automatic multi-step reasoning and tool-use for large language models,” arXiv:2303.09014, 2023.

[11] E. Karpas et al., “MRKL systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning,” arXiv:2205.00445, 2022.

[12] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Adv. Neural Inf. Process. Syst., vol. 33, 2020, pp. 9459–9474.

[13] K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang, “Retrieval augmented language model pre-training,” in Proc. ICML, vol. 119, 2020, pp. 3929–3938.

[14] V. Karpukhin et al., “Dense passage retrieval for open-domain question answering,” in Proc. EMNLP, 2020, pp. 6769–6781.

[15] Y. Li, “Findable then explainable: Retrieval-summary integration for code intelligence on a lightweight CodeSearchNet subset,” J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65–82, Jul. 2024, doi: 10.69987/JACS.2024.40706.

[16] S. Zhou, Z. Li, and E. Wang, “Long-document RAG for contractual and insurance clause analysis in receivables RWA structures,” J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88–104, Aug. 2024, doi: 10.69987/JACS.2024.40810.

[17] O. Khattab and M. Zaharia, “ColBERT: Efficient and effective passage search via contextualized late interaction over BERT,” in Proc. ACM SIGIR, 2020, pp. 39–48.

[18] S. Robertson and H. Zaragoza, “The probabilistic relevance framework: BM25 and beyond,” Found. Trends Inf. Retr., vol. 3, no. 4, pp. 333–389, 2009.

[19] S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by latent semantic analysis,” J. Amer. Soc. Inf. Sci., vol. 41, no. 6, pp. 391–407, 1990.

[20] G. Salton and C. Buckley, “Term-weighting approaches in automatic text retrieval,” Inf. Process. Manage., vol. 24, no. 5, pp. 513–523, 1988.

[21] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proc. EMNLP-IJCNLP, 2019, pp. 3982–3992.

[22] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186.

[23] C. Raffel et al., “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res., vol. 21, no. 140, pp. 1–67, 2020.

[24] Z. S. Zhong, R. Ma, and H. Zhao, “Human-uncertainty distillation for calibrated vision models on CIFAR-10H,” J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77–89, Feb. 2023, doi: 10.69987/JACS.2023.30206.

[25] Q. Xin, “Uncertainty-aware late fusion for 3D perception: Confidence calibration and fusion rule learning,” J. Technol. Informatics Eng., vol. 4, no. 1, pp. 215–238, Feb. 2025, doi: 10.51903/jtie.v4i1.485.

[26] J. Wei et al., “Chain-of-thought prompting elicits reasoning in large language models,” in Adv. Neural Inf. Process. Syst., vol. 35, 2022, pp. 24824–24837.

[27] X. Wang et al., “Self-consistency improves chain of thought reasoning in language models,” in Proc. ICLR, 2023.

[28] L. Gao et al., “PAL: Program-aided language models,” in Proc. ICML, vol. 202, 2023, pp. 10764–10799.

[29] M. Ahn et al., “Do as I can, not as I say: Grounding language in robotic affordances,” in Proc. CoRL, vol. 205, 2023, pp. 287–318.

[30] Z. S. Zhong and S. Ling, “Improved theoretical guarantee for rank aggregation via spectral method,” Inf. Inference, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.

[31] G. V. Cormack, C. L. A. Clarke, and S. Buettcher, “Reciprocal rank fusion outperforms Condorcet and individual rank learning methods,” in Proc. ACM SIGIR, 2009, pp. 758–759.

[32] S. Borgeaud et al., “Improving language models by retrieving from trillions of tokens,” in Proc. ICML, vol. 162, 2022, pp. 2206–2240.

[33] U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis, “Generalization through memorization: Nearest neighbor language models,” in Proc. ICLR, 2020.

[34] Z. S. Zhong and S. Ling, “Uncertainty quantification of spectral estimator and MLE for orthogonal group synchronization,” arXiv:2408.05944, Aug. 2024.

[35] F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.

[36] Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, no. 2, pp. 153–157, 1947.

[37] H. Zhang, “DriftGuard: Multi-signal drift early warning and safe re-training/rollback for CTR/CVR models,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24–40, 2023, doi: 10.69987/JACS.2023.30703.

[38] J. Nie and D. Zheng, “Noisy-neighbor-aware VM degradation risk modeling with unsupervised residual fusion,” J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112–123, Apr. 2024, doi: 10.69987/JACS.2024.40409.

[39] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, “Intelligent fault analysis with AIOps technology,” J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94–100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.

[40] Q. Xin, “Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes,” Emerg. Inf. Sci. Technol., vol. 6, no. 2, pp. 125–146, Nov. 2025, doi: 10.18196/eist.v6i2.31232.

[41] H. Zhang, “LLM-driven CI failure diagnosis and automated repair: From GitHub Actions logs to patch recommendation,” J. Technol. Informatics Eng., vol. 4, no. 1, pp. 190–214, Feb. 2025.

[42] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, “Enhancing bank credit risk management using the C5.0 decision tree algorithm,” J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100–107, Nov. 2024, doi: 10.5281/zenodo.14032041.

[43] H. Zhang, “Risk-aware budget-constrained auto-bidding under first-price RTB: A distributional constrained deep reinforcement learning framework,” J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30–47, Jun. 2024, doi: 10.69987/JACS.2024.40603.

[44] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, “Enhancing financial services through big data and AI-driven customer insights and risk analysis,” J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53–62, Jul. 2024, doi: 10.60087/jklst.vol3.n3.p53-62.

[45] H. Zhang, “Counterfactual learning-to-rank for ads: Off-policy evaluation on the Open Bandit Dataset,” J. Adv. Comput. Syst., vol. 5, no. 12, pp. 1–11, 2025.

[46] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, “Predicting enterprise marketing decision making with intelligent data-driven approaches,” J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12–19, Jun. 2024, doi: 10.5281/zenodo.11357252.

[47] Y. Lu, H. Zhou, and Y. Zhang, “A constrained, data-driven budgeting framework integrating macro demand forecasting and marketing response modeling,” J. Technol. Informatics Eng., vol. 4, no. 3, Dec. 2025, doi: 10.51903/jtie.v4i3.466.

[48] H. Zhou and K. Zhang, “News-based uncertainty and macro-market fusion for VIX direction forecasting: Evidence from 2015–2024 FRED panel,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 487–501, Aug. 2025, doi: 10.51903/jtie.v4i2.540.

[49] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, “Intelligent classification and personalized recommendation of e-commerce products based on machine learning,” Appl. Comput. Eng., vol. 64, no. 1, pp. 143–149, May 2024, doi: 10.54254/2755-2721/64/20241365.

[50] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, “Autonomous driving system driven by artificial intelligence perception fusion,” Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193–198, Feb. 2024, doi: 10.54097/e0b9ak47.

[51] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, “Optimization of autonomous driving image detection based on RFAConv and triplet attention,” Appl. Comput. Eng., vol. 77, no. 1, pp. 210–217, 2024, doi: 10.54254/2755-2721/77/2024MA0067.

[52] L. Zhang, R. Ma, and P. Greg, “Digital-twin dispatching for urban mobility via spatio-temporal transformers and offline reinforcement learning,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 337–363, Aug. 2025, doi: 10.51903/jtie.v4i2.501.

[53] R. Ma, L. Zhang, and T. Song, “Computer-vision-informed visual explanation cards for autonomous-driving traffic-sign alerts: Localization, classification, and retrieved evidence on GTSDB,” Int. J. Graph. Des., vol. 3, no. 2, pp. 455–473, Oct. 2025, doi: 10.51903/ijgd.v3i2.3991.

[54] H. Wang, Y. Ren, and X. Chang, “Layout-aware progressive PDF rendering: AI prioritization of PDF slices to reduce time-to-functional-first-frame on FUNSD,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 425–446, Aug. 2025, doi: 10.51903/jtie.v4i2.523.

[55] B. Zhou, H. Wang, and X. Chang, “Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 447–463, Aug. 2025, doi: 10.51903/jtie.v4i2.522.

[56] Z. Li, S. Zhou, and Z. Zhou, “Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench,” Int. J. Graph. Des., vol. 3, no. 1, pp. 196–210, May 2025, doi: 10.51903/ijgd.v3i1.3715.

[57] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, “Implementation of an AI-based MRD evaluation and prediction model for multiple myeloma,” Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127–131, Jan. 2024, doi: 10.54097/zJ4MnbWW.

[58] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, “Cancer image classification based on DenseNet model,” J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, Nov. 2020, doi: 10.1088/1742-6596/1651/1/012143.

[59] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, “IoT traffic classification and anomaly detection method based on deep autoencoders,” Appl. Comput. Eng., vol. 69, no. 1, pp. 64–70, Jul. 2024, doi: 10.54254/2755-2721/69/20241511.

[60] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, “Predictive optimization of DDoS attack mitigation in distributed systems using machine learning,” Appl. Comput. Eng., vol. 64, no. 1, pp. 89–94, May 2024, doi: 10.54254/2755-2721/64/20241350.

[61] A. G. S. Raj et al., “Impact of bilingual CS education on student learning and engagement in a data structures course,” in Proc. 19th Koli Calling Int. Conf. Comput. Educ. Res., 2019, pp. 1–10, doi: 10.1145/3364510.3364518.

[62] Z. S. Zhong, X. Pan, and Q. Lei, “Bridging domains with approximately shared features,” in Proc. AISTATS, PMLR, vol. 258, 2025, pp. 559–567.

[63] Q. Xin, “Hybrid cloud architecture for efficient and cost-effective large language model deployment,” J. Inf. Syst. Informatics, vol. 7, no. 3, pp. 2182–2195, Sep. 2025, doi: 10.51519/journalisi.v7i3.1170.

Published

2025-08-31

How to Cite

Ran, H., Cooper, R., Bennett, M., & Wang, J. (2025). Schema-Aware Retrieval-Augmented Tool-Use Agents for API Selection, Planning, and Calling: An Empirical Study on API-Bank. Journal of Information Systems and Business Technology, 1(2), 84-95. https://journal.jci.co.id/jisbt/article/view/809

Most read articles by the same author(s)