Uncertainty-Gated Adaptive RAG for Scientific Claim Verification with Evidence-Calibrated Abstention
Keywords:
Scientific Claim Verification, Retrieval-Augmented Generation, Adaptive RAG, Uncertainty Calibration, Selective Prediction, Abstention, Scifact, Trustworthy AIAbstract
Scientific claim verification must retrieve relevant evidence, assign SUPPORT, CONTRADICT, or NEI, and avoid confident decisions when evidence or the verifier is unreliable. We evaluate an uncertainty-gated adaptive RAG pipeline on 5,183 SciFact abstracts and 1,109 labeled claims through five official cross-validation folds. Four systems share one evidence selector and verifier: BM25 RAG, dense LSA RAG, always-on sparse-first hybrid RAG, and an adaptive system that begins with one BM25 document and two sentences, escalates uncertain claims to five hybrid documents and five sentences, arbitrates between fast and full posteriors, and applies calibrated abstention. The adaptive system answered 95.7% of claims with 54.4% selective accuracy and 47.9% macro-F1, versus 52.2% accuracy and 47.3% macro-F1 for BM25. Forced adaptive accuracy was 53.4%, while average context fell from 241.9 to 139.6 tokens, a 42.3% reduction. Abstention captured errors in 68.8% of rejected cases and reduced wrong answers by 8.7% relative to BM25. BM25 retained the strongest evidence recall, indicating that the adaptive benefit came mainly from calibrated routing and refusal rather than unconditional fusion. Contradiction reasoning remained the dominant error source.
References
[1] D. Wadden, S. Lin, K. Lo, L. L. Wang, M. van Zuylen, A. Cohan, and H. Hajishirzi, “Fact or Fiction: Verifying Scientific Claims,” in Proc. 2020 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 7534–7550, doi: 10.18653/v1/2020.emnlp-main.609.
[2] J. Thorne, A. Vlachos, O. Cocarascu, C. Christodoulopoulos, and A. Mittal, “FEVER: a Large-Scale Dataset for Fact Extraction and VERification,” in Proc. 2018 Conf. North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 1, 2018, pp. 809–819, doi: 10.18653/v1/N18-1074.
[3] N. Kotonya and F. Toni, “Explainable Automated Fact-Checking for Public Health Claims,” in Proc. 2020 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 7740–7754, doi: 10.18653/v1/2020.emnlp-main.623.
[4] A. Saakyan, T. Chakrabarty, and S. Muresan, “COVID-Fact: Fact Extraction and Verification of Real-World Claims on COVID-19 Pandemic,” in Proc. 59th Annual Meeting of the Association for Computational Linguistics and 11th Int. Joint Conf. Natural Language Processing, vol. 1, 2021, pp. 2116–2129, doi: 10.18653/v1/2021.acl-long.165.
[5] M. Sarrouti, A. Ben Abacha, Y. Mrabet, and D. Demner-Fushman, “Evidence-based Fact-Checking of Health-related Claims,” in Findings of the Association for Computational Linguistics: EMNLP 2021, 2021, pp. 3499–3512, doi: 10.18653/v1/2021.findings-emnlp.297.
[6] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, “Implementation of an AI-based MRD evaluation and prediction model for multiple myeloma,” Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127-131, Jan. 2024, doi: 10.54097/zJ4MnbWW.
[7] D. Wadden, K. Lo, B. Kuehl, A. Cohan, I. Beltagy, L. L. Wang, and H. Hajishirzi, “SciFact-Open: Towards Open-Domain Scientific Claim Verification,” in Findings of the Association for Computational Linguistics: EMNLP 2022, 2022, pp. 4719–4734, doi: 10.18653/v1/2022.findings-emnlp.347.
[8] J. Nie and D. Zheng, “Ambiguity-aware HDFS log anomaly detection with retrieval-augmented failure narratives and selective refusal,” J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66-80, Jan. 2023, doi: 10.69987/JACS.2023.30105.
[9] R. Pradeep, X. Ma, R. Nogueira, and J. Lin, “Scientific Claim Verification with VerT5erini,” in Proc. 12th Int. Workshop on Health Text Mining and Information Analysis, 2021, pp. 94–103, doi: 10.18653/v1/2021.louhi-1.11.
[10] S. Zhou, Z. Li, and E. Wang, “Evidence-grounded RAG for tokenized trade receivable disclosure QA under U.S. capital market standards,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41-57, Jul. 2023, doi: 10.69987/JACS.2023.30704.
[11] J. Vladika and F. Matthes, “Scientific Fact-Checking: A Survey of Resources and Approaches,” in Findings of the Association for Computational Linguistics: ACL 2023, 2023, pp. 6215–6230, doi: 10.18653/v1/2023.findings-acl.387.
[12] P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 9459–9474.
[13] S. Zhou, Z. Li, and E. Wang, “Long-document RAG for contractual and insurance clause analysis in receivables RWA structures,” J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88-104, Aug. 2024, doi: 10.69987/JACS.2024.40810.
[14] Y. Li, “Findable then explainable: Retrieval-summary integration for code intelligence on a lightweight CodeSearchNet subset,” J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65-82, Jul. 2024, doi: 10.69987/JACS.2024.40706.
[15] Q. Xin, “Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes,” Emerg. Inf. Sci. Technol., vol. 6, no. 2, pp. 125-146, Nov. 2025, doi: 10.18196/eist.v6i2.31232.
[16] H. Zhang, “LLM-driven CI failure diagnosis and automated repair: From GitHub Actions logs to patch recommendation,” J. Technol. Informatics Eng., vol. 4, no. 1, pp. 190-214, Apr. 2025, doi: 10.51903/jtie.v4i1.484.
[17] S. Robertson and H. Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Foundations and Trends in Information Retrieval, vol. 3, no. 4, pp. 333–389, 2009, doi: 10.1561/1500000019.
[18] G. Salton, A. Wong, and C. S. Yang, “A Vector Space Model for Automatic Indexing,” Communications of the ACM, vol. 18, no. 11, pp. 613–620, 1975, doi: 10.1145/361219.361220.
[19] S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by Latent Semantic Analysis,” Journal of the American Society for Information Science, vol. 41, no. 6, pp. 391–407, 1990, doi: 10.1002/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9.
[20] G. V. Cormack, C. L. A. Clarke, and S. Büttcher, “Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods,” in Proc. 32nd Int. ACM SIGIR Conf. Research and Development in Information Retrieval, 2009, pp. 758–759, doi: 10.1145/1571941.1572114.
[21] Z. S. Zhong and S. Ling, “Improved theoretical guarantee for rank aggregation via spectral method,” Inf. Inference: J. IMA, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.
[22] V. Karpukhin et al., “Dense Passage Retrieval for Open-Domain Question Answering,” in Proc. 2020 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 6769–6781, doi: 10.18653/v1/2020.emnlp-main.550.
[23] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. 2019 Conf. North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 1, 2019, pp. 4171–4186, doi: 10.18653/v1/N19-1423.
[24] I. Beltagy, K. Lo, and A. Cohan, “SciBERT: A Pretrained Language Model for Scientific Text,” in Proc. 2019 Conf. Empirical Methods in Natural Language Processing and 9th Int. Joint Conf. Natural Language Processing, 2019, pp. 3615–3620, doi: 10.18653/v1/D19-1371.
[25] Z. S. Zhong, X. Pan, and Q. Lei, “Bridging domains with approximately shared features,” in Proc. 28th Int. Conf. Artificial Intelligence and Statistics (AISTATS), vol. 258, 2025, pp. 559-567.
[26] A. Cohan, S. Feldman, I. Beltagy, D. Downey, and D. S. Weld, “SPECTER: Document-level Representation Learning using Citation-informed Transformers,” in Proc. 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 2270–2282, doi: 10.18653/v1/2020.acl-main.207.
[27] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in Proc. 2019 Conf. Empirical Methods in Natural Language Processing and 9th Int. Joint Conf. Natural Language Processing, 2019, pp. 3982–3992, doi: 10.18653/v1/D19-1410.
[28] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On Calibration of Modern Neural Networks,” in Proc. 34th Int. Conf. Machine Learning, vol. 70, 2017, pp. 1321–1330.
[29] Z. S. Zhong, R. Ma, and H. Zhao, “Human-uncertainty distillation for calibrated vision models on CIFAR-10H,” J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77-89, Feb. 2023, doi: 10.69987/JACS.2023.30206.
[30] S. Desai and G. Durrett, “Calibration of Pre-trained Transformers,” in Proc. 2020 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 295–302, doi: 10.18653/v1/2020.emnlp-main.21.
[31] Z. S. Zhong and S. Ling, “Uncertainty quantification of spectral estimator and MLE for orthogonal group synchronization,” arXiv:2408.05944, Aug. 2024.
[32] M. Minderer, J. Djolonga, R. Romijnders, F. Hubis, X. Zhai, N. Houlsby, D. Tran, and M. Lucic, “Revisiting the Calibration of Modern Neural Networks,” in Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 15682–15694.
[33] Q. Xin, “Uncertainty-aware late fusion for 3D perception: Confidence calibration and fusion rule learning,” J. Technol. Informatics Eng., vol. 4, no. 1, pp. 215-238, Apr. 2025, doi: 10.51903/jtie.v4i1.485.
[34] Y. Geifman and R. El-Yaniv, “Selective Classification for Deep Neural Networks,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 4885–4894.
[35] H. Zhou and K. Zhang, “News-based uncertainty and macro-market fusion for VIX direction forecasting: Evidence from 2015-2024 FRED panel,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 487-501, Aug. 2025, doi: 10.51903/jtie.v4i2.540.
[36] Y. Geifman and R. El-Yaniv, “SelectiveNet: A Deep Neural Network with an Integrated Reject Option,” in Proc. 36th Int. Conf. Machine Learning, vol. 97, 2019, pp. 2151–2159.
[37] A. Kamath, R. Jia, and P. Liang, “Selective Question Answering under Domain Shift,” in Proc. 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 5684–5696, doi: 10.18653/v1/2020.acl-main.503.
[38] D. Hendrycks and K. Gimpel, “A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks,” in Proc. Int. Conf. Learning Representations, 2017.
[39] B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 6402–6413.
[40] Y. Gal and Z. Ghahramani, “Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning,” in Proc. 33rd Int. Conf. Machine Learning, vol. 48, 2016, pp. 1050–1059.
[41] F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
[42] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, “Optimization of autonomous driving image detection based on RFAConv and triplet attention,” in Proc. 2nd Int. Conf. Software Engineering and Machine Learning (SEML), 2024, pp. 210-217, doi: 10.54254/2755-2721/77/2024MA0067.
[43] R. Nogueira and K. Cho, “Passage Re-ranking with BERT,” arXiv:1901.04085, 2019.
[44] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, “Autonomous driving system driven by artificial intelligence perception fusion,” Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193-198, Feb. 2024, doi: 10.54097/e0b9ak47.
[45] B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York, NY, USA: Chapman & Hall, 1993.
[46] Q. Xin, “Hybrid cloud architecture for efficient and cost-effective large language model deployment,” J. Inf. Syst. Informatics, vol. 7, no. 3, pp. 2182-2195, Sep. 2025, doi: 10.51519/journalisi.v7i3.1170.
[47] B. Zhou, H. Wang, and X. Chang, “Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 447-463, Aug. 2025, doi: 10.51903/jtie.v4i2.522.
[48] L. Zhang, R. Ma, and P. Greg, “Digital-twin dispatching for urban mobility via spatio-temporal transformers and offline reinforcement learning,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 337-363, Aug. 2025, doi: 10.51903/jtie.v4i2.501.
[49] H. Wang, Y. Ren, and X. Chang, “Layout-aware progressive PDF rendering: AI prioritization of PDF slices to reduce time-to-functional-first-frame on FUNSD,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 425-446, Aug. 2025, doi: 10.51903/jtie.v4i2.523.
[50] H. Zhang, “DriftGuard: Multi-signal drift early warning and safe re-training/rollback for CTR/CVR models,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24-40, Jul. 2023, doi: 10.69987/JACS.2023.30703.
[51] J. Nie and D. Zheng, “Noisy-neighbor-aware VM degradation risk modeling with unsupervised residual fusion,” J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112-123, Apr. 2024, doi: 10.69987/JACS.2024.40409.
[52] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, “Intelligent fault analysis with AIOps technology,” J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94-100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.
[53] H. Zhang, “Risk-aware budget-constrained auto-bidding under first-price RTB: A distributional constrained deep reinforcement learning framework,” J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30-47, Jun. 2024, doi: 10.69987/JACS.2024.40603.
[54] Y. Lu, H. Zhou, and Y. Zhang, “A constrained, data-driven budgeting framework integrating macro demand forecasting and marketing response modeling,” J. Technol. Informatics Eng., vol. 4, no. 3, pp. 493-520, Dec. 2025, doi: 10.51903/jtie.v4i3.466.
[55] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, “IoT traffic classification and anomaly detection method based on deep autoencoders,” in Proc. 6th Int. Conf. Computing and Data Science (CDS), 2024, pp. 64-70, doi: 10.54254/2755-2721/69/20241511.
[56] Z. Li, S. Zhou, and Z. Zhou, “Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench,” Int. J. Graphic Des., vol. 3, no. 1, pp. 196-210, May 2025, doi: 10.51903/ijgd.v3i1.3715.
[57] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, “Enhancing bank credit risk management using the C5.0 decision tree algorithm,” J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100-107, Nov. 2024, doi: 10.5281/zenodo.14032041.
[58] R. Ma, L. Zhang, and T. Song, “Computer-vision-informed visual explanation cards for autonomous-driving traffic-sign alerts: Localization, classification, and retrieved evidence on GTSDB,” Int. J. Graphic Des., vol. 3, no. 2, pp. 455-473, Oct. 2025, doi: 10.51903/ijgd.v3i2.3991.
[59] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, “Predictive optimization of DDoS attack mitigation in distributed systems using machine learning,” in Proc. 6th Int. Conf. Computing and Data Science (CDS), 2024, pp. 89-94, doi: 10.54254/2755-2721/64/20241350.
[60] A. G. S. Raj et al., “Impact of bilingual CS education on student learning and engagement in a data structures course,” in Proc. 19th Koli Calling Int. Conf. Comput. Educ. Res., 2019, pp. 1-10, doi: 10.1145/3364510.3364518.
[61] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, “Enhancing financial services through big data and AI-driven customer insights and risk analysis,” J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53-62, Jul. 2024, doi: 10.60087/jklst.vol3.n3.p53-62.
[62] H. Zhang, “Counterfactual learning-to-rank for ads: Off-policy evaluation on the Open Bandit Dataset,” J. Adv. Comput. Syst., vol. 5, no. 12, pp. 1-11, Dec. 2025, doi: 10.69987/JACS.2025.51201.
[63] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, “Intelligent classification and personalized recommendation of e-commerce products based on machine learning,” in Proc. 6th Int. Conf. Computing and Data Science (ICCDS), 2024, pp. 143-149, doi: 10.54254/2755-2721/64/20241365.
[64] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, “Cancer image classification based on DenseNet model,” J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, Nov. 2020, doi: 10.1088/1742-6596/1651/1/012143.
[65] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, “Predicting enterprise marketing decision making with intelligent data-driven approaches,” J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12-19, Jun. 2024, doi: 10.5281/zenodo.11357252.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Lin Yang, Kevin Parker, David Huang (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Creative Commons Attribution 4.0 International (CC BY 4.0).




This work is licensed under a