Out-of-Scope-Aware Tool Routing Agents with Selective Invocation and Calibrated Rejection

Authors

Keywords:

Tool Routing, Intent Classification, Out-Of-Scope Detection, Selective Prediction, Confidence Calibration, Abstention, CLINC150

Abstract

Tool-using agents need a router that selects the correct capability while refusing requests outside the available catalog. This study formulates routing on CLINC150 as joint 150-way tool classification, out-of-scope (OOS) detection, and selective invocation. The proposed Calibrated Selective Invocation Router (CSIR) combines word and character TF-IDF features, a linear support vector machine route head, temperature-scaled route probabilities, a 151-class OOS specialist, class-prototype distance, and a calibrated rejection gate. The official partitions provide 15,000/3,000/4,500 in-scope training/validation/test utterances and 100/100/1,000 OOS utterances. Disjoint validation subsets calibrate probabilities and select thresholds before frozen evaluation on 5,500 test queries. The hybrid router achieves 92.73% top-1, 97.36% top-3, and 98.11% top-5 accuracy. Temperature scaling reduces expected calibration error from 0.8705 to 0.0212. OOS fusion reaches 0.9585 AUROC, 0.8593 AUPR-OOS, and 0.1682 FPR at 95% OOS recall. Under moderate costs, CSIR accepts 75.11% of queries at 94.55% precision, rejects 90.90% of OOS inputs, and yields 0.5046 utility per query, exceeding 0.4796 for a maximum-probability gate and -0.3287 for always invoking. Paired bootstrap intervals support the routing, OOS, and utility gains, while warmed batch inference requires 0.0529 ms per query.

References

[1] E. Karpas et al., "MRKL systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning," arXiv:2205.00445, 2022.

[2] Q. Xin, "Hybrid cloud architecture for efficient and cost-effective large language model deployment," J. Inf. Syst. Informatics, vol. 7, no. 3, pp. 2182-2195, Sep. 2025, doi: 10.51519/journalisi.v7i3.1170.

[3] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao, "ReAct: Synergizing reasoning and acting in language models," in Proc. ICLR, 2023.

[4] H. Zhang, "LLM-driven CI failure diagnosis and automated repair: From GitHub Actions logs to patch recommendation," J. Technol. Informatics Eng., vol. 4, no. 1, pp. 190-214, Apr. 2025, doi: 10.51903/jtie.v4i1.484.

[5] T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom, "Toolformer: Language models can teach themselves to use tools," in Adv. Neural Inf. Process. Syst., vol. 36, pp. 68539-68551, 2023.

[6] Y. Li, "Findable then explainable: Retrieval-summary integration for code intelligence on a lightweight CodeSearchNet subset," J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65-82, Jul. 2024, doi: 10.69987/JACS.2024.40706.

[7] S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez, "Gorilla: Large language model connected with massive APIs," arXiv:2305.15334, 2023.

[8] S. Zhou, Z. Li, and E. Wang, "Evidence-grounded RAG for tokenized trade receivable disclosure QA under U.S. capital market standards," J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41-57, Jul. 2023, doi: 10.69987/JACS.2023.30704.

[9] G. Mialon et al., "Augmented language models: A survey," Trans. Mach. Learn. Res., 2023.

[10] S. Zhou, Z. Li, and E. Wang, "Long-document RAG for contractual and insurance clause analysis in receivables RWA structures," J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88-104, Aug. 2024, doi: 10.69987/JACS.2024.40810.

[11] Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, "HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face," in Adv. Neural Inf. Process. Syst., vol. 36, pp. 38154-38180, 2023.

[12] Q. Xin, "Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes," Emerg. Inf. Sci. Technol., vol. 6, no. 2, pp. 125-146, Nov. 2025, doi: 10.18196/eist.v6i2.31232.

[13] Y. Qin et al., "ToolLLM: Facilitating large language models to master 16000+ real-world APIs," arXiv:2307.16789, 2023.

[14] J. Nie and D. Zheng, "Ambiguity-aware HDFS log anomaly detection with retrieval-augmented failure narratives and selective refusal," J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66-80, Jan. 2023, doi: 10.69987/JACS.2023.30105.

[15] S. Larson, A. Mahendran, J. J. Peper, C. Clarke, A. Lee, P. Hill, J. K. Kummerfeld, K. Leach, M. A. Laurenzano, L. Tang, and J. Mars, "An evaluation dataset for intent classification and out-of-scope prediction," in Proc. EMNLP-IJCNLP, 2019, pp. 1311-1316.

[16] A. G. S. Raj et al., "Impact of bilingual CS education on student learning and engagement in a data structures course," in Proc. 19th Koli Calling Int. Conf. Comput. Educ. Res., 2019, pp. 1-10, doi: 10.1145/3364510.3364518.

[17] A. Coucke et al., "Snips voice platform: An embedded spoken language understanding system for private-by-design voice interfaces," arXiv:1805.10190, 2018.

[18] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, "Intelligent classification and personalized recommendation of e-commerce products based on machine learning," Appl. Comput. Eng., vol. 64, no. 1, pp. 143-149, May 2024, doi: 10.54254/2755-2721/64/20241365.

[19] I. Casanueva, T. Temcinas, D. Gerz, M. Henderson, and I. Vulic, "Efficient intent detection with dual sentence encoders," in Proc. 2nd Workshop NLP for ConvAI, 2020, pp. 38-45.

[20] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, "Cancer image classification based on DenseNet model," J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, Nov. 2020, doi: 10.1088/1742-6596/1651/1/012143.

[21] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, "Implementation of an AI-based MRD evaluation and prediction model for multiple myeloma," Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127-131, 2024, doi: 10.54097/zJ4MnbWW.

[22] D. Hendrycks and K. Gimpel, "A baseline for detecting misclassified and out-of-distribution examples in neural networks," in Proc. ICLR, 2017.

[23] Z. S. Zhong, R. Ma, and H. Zhao, "Human-uncertainty distillation for calibrated vision models on CIFAR-10H," J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77-89, Feb. 2023, doi: 10.69987/JACS.2023.30206.

[24] K. Lee, K. Lee, H. Lee, and J. Shin, "A simple unified framework for detecting out-of-distribution samples and adversarial attacks," in Adv. Neural Inf. Process. Syst., vol. 31, 2018.

[25] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, "IoT traffic classification and anomaly detection method based on deep autoencoders," Appl. Comput. Eng., vol. 69, pp. 64-70, Jul. 2024, doi: 10.54254/2755-2721/69/20241511.

[26] D. Hendrycks, M. Mazeika, and T. Dietterich, "Deep anomaly detection with outlier exposure," in Proc. ICLR, 2019.

[27] J. Nie and D. Zheng, "Noisy-neighbor-aware VM degradation risk modeling with unsupervised residual fusion," J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112-123, Apr. 2024, doi: 10.69987/JACS.2024.40409.

[28] W. Liu, X. Wang, J. D. Owens, and Y. Li, "Energy-based out-of-distribution detection," in Adv. Neural Inf. Process. Syst., vol. 33, pp. 21464-21475, 2020.

[29] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, "Predictive optimization of DDoS attack mitigation in distributed systems using machine learning," Appl. Comput. Eng., vol. 64, no. 1, pp. 94-99, May 2024, doi: 10.54254/2755-2721/64/20241350.

[30] S. Vaze, K. Han, A. Vedaldi, and A. Zisserman, "Open-set recognition: A good closed-set classifier is all you need?" in Proc. ICLR, 2022.

[31] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, "Autonomous driving system driven by artificial intelligence perception fusion," Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193-198, Feb. 2024, doi: 10.54097/e0b9ak47.

[32] T.-E. Lin and H. Xu, "Deep unknown intent detection with margin loss," in Proc. 57th Annu. Meeting Assoc. Comput. Linguistics, 2019, pp. 5491-5496.

[33] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, "Optimization of autonomous driving image detection based on RFAConv and triplet attention," in Proc. 2nd Int. Conf. Software Eng. Mach. Learn. (SEML), 2024; arXiv:2407.09530.

[34] H. Zhang, H. Xu, and T.-E. Lin, "Deep open intent classification with adaptive decision boundary," Proc. AAAI Conf. Artif. Intell., vol. 35, no. 16, pp. 14374-14382, 2021.

[35] Q. Xin, "Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning)," J. Technol. Informatics Eng., vol. 4, no. 1, pp. 215-238, Apr. 2025, doi: 10.51903/jtie.v4i1.485.

[36] L.-M. Zhan, H. Liang, B. Liu, L. Fan, X.-M. Wu, and A. Y. S. Lam, "Out-of-scope intent detection with self-supervision and discriminative training," in Proc. 59th Annu. Meeting Assoc. Comput. Linguistics, 2021, pp. 3521-3532.

[37] Z. S. Zhong, X. Pan, and Q. Lei, "Bridging domains with approximately shared features," in Proc. 28th Int. Conf. Artif. Intell. Statist. (AISTATS), PMLR, vol. 258, pp. 559-567, 2025.

[38] Y. Ouyang, J. Ye, Y. Chen, X. Dai, S. Huang, and J. Chen, "Energy-based unknown intent detection with data manipulation," in Findings ACL-IJCNLP, 2021, pp. 2852-2861.

[39] H. Zhang, "DriftGuard: Multi-signal drift early warning and safe re-training/rollback for CTR/CVR models," J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24-40, Jul. 2023, doi: 10.69987/JACS.2023.30703.

[40] H. Lang, Y. Zheng, J. Sun, F. Huang, L. Si, and Y. Li, "Estimating soft labels for out-of-domain intent detection," in Proc. EMNLP, 2022, pp. 261-276.

[41] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, "Intelligent fault analysis with AIOps technology," J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94-100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.

[42] J. C. Platt, "Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods," in Advances in Large Margin Classifiers, A. J. Smola, P. L. Bartlett, B. Scholkopf, and D. Schuurmans, Eds. Cambridge, MA, USA: MIT Press, 1999, pp. 61-74.

[43] Z. S. Zhong and S. Ling, "Improved theoretical guarantee for rank aggregation via spectral method," Inf. Inference: J. IMA, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.

[44] A. Niculescu-Mizil and R. Caruana, "Predicting good probabilities with supervised learning," in Proc. 22nd Int. Conf. Mach. Learn., 2005, pp. 625-632.

[45] Z. S. Zhong and S. Ling, "Uncertainty quantification of spectral estimator and MLE for orthogonal group synchronization," arXiv:2408.05944, Aug. 2024.

[46] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, "On calibration of modern neural networks," in Proc. 34th Int. Conf. Mach. Learn., 2017, pp. 1321-1330.

[47] H. Zhou and K. Zhang, "News-based uncertainty and macro-market fusion for VIX direction forecasting: Evidence from 2015-2024 FRED panel," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 487-501, Aug. 2025, doi: 10.51903/jtie.v4i2.540.

[48] R. El-Yaniv and Y. Wiener, "On the foundations of noise-free selective classification," J. Mach. Learn. Res., vol. 11, pp. 1605-1641, 2010.

[49] H. Zhang, "Risk-aware budget-constrained auto-bidding under first-price RTB: A distributional constrained deep reinforcement learning framework," J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30-47, Jun. 2024, doi: 10.69987/JACS.2024.40603.

[50] Y. Geifman and R. El-Yaniv, "SelectiveNet: A deep neural network with an integrated reject option," in Proc. 36th Int. Conf. Mach. Learn., 2019, pp. 2151-2159.

[51] H. Zhang, "Counterfactual learning-to-rank for ads: Off-policy evaluation on the Open Bandit Dataset," J. Adv. Comput. Syst., vol. 5, no. 12, pp. 1-11, Dec. 2025, doi: 10.69987/JACS.2025.51201.

[52] L. Zhang, R. Ma, and P. Greg, "Digital-twin dispatching for urban mobility via spatio-temporal transformers and offline reinforcement learning," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 337-363, Aug. 2025, doi: 10.51903/jtie.v4i2.501.

[53] G. Salton and C. Buckley, "Term-weighting approaches in automatic text retrieval," Inf. Process. Manage., vol. 24, no. 5, pp. 513-523, 1988.

[54] C. Cortes and V. Vapnik, "Support-vector networks," Mach. Learn., vol. 20, pp. 273-297, 1995.

[55] R.-E. Fan, K.-W. Chang, C.-J. Hsieh, X.-R. Wang, and C.-J. Lin, "LIBLINEAR: A library for large linear classification," J. Mach. Learn. Res., vol. 9, pp. 1871-1874, 2008.

[56] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, "Enhancing bank credit risk management using the C5.0 decision tree algorithm," J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100-107, Nov. 2024, doi: 10.5281/zenodo.14032041.

[57] Z. Li, S. Zhou, and Z. Zhou, "Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench," Int. J. Graphic Des., vol. 3, no. 1, pp. 196-210, May 2025, doi: 10.51903/ijgd.v3i1.3715.

[58] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, "Predicting enterprise marketing decision making with intelligent data-driven approaches," J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12-19, Jun. 2024, doi: 10.5281/zenodo.11357252.

[59] R. Ma, L. Zhang, and T. Song, "Computer-vision-informed visual explanation cards for autonomous-driving traffic-sign alerts: Localization, classification, and retrieved evidence on GTSDB," Int. J. Graphic Des., vol. 3, no. 2, pp. 455-473, Oct. 2025, doi: 10.51903/ijgd.v3i2.3991.

[60] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, "Enhancing financial services through big data and AI-driven customer insights and risk analysis," J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53-62, Jul. 2024, doi: 10.60087/jklst.vol3.n3.p53-62.

[61] B. Zhou, H. Wang, and X. Chang, "Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 447-463, Aug. 2025, doi: 10.51903/jtie.v4i2.522.

[62] Y. Lu, H. Zhou, and Y. Zhang, "A constrained, data-driven budgeting framework integrating macro demand forecasting and marketing response modeling," J. Technol. Informatics Eng., vol. 4, no. 3, pp. 493-520, Dec. 2025, doi: 10.51903/jtie.v4i3.466.

[63] H. Wang, Y. Ren, and X. Chang, "Layout-aware progressive PDF rendering: AI prioritization of PDF slices to reduce time-to-functional-first-frame on FUNSD," J. Technol. Informatics Eng., vol. 4, no. 2, pp. 425-446, Aug. 2025, doi: 10.51903/jtie.v4i2.523.

Downloads

Published

2025-10-31

How to Cite

Taylor, M., Wu, M., & Harris, M. (2025). Out-of-Scope-Aware Tool Routing Agents with Selective Invocation and Calibrated Rejection. Journal of Information Technology and Informatics Engineering, 1(2), 43-51. https://journal.jci.co.id/jitie/article/view/822