GraphRAG for Scientific Evidence Verification through Claim–Document–Entity Graph Reasoning

Authors

Keywords:

Graphrag, Scientific Fact Checking, Evidence Retrieval, Knowledge Graph, Retrieval-Augmented Generation, Scifact, Claim Verification

Abstract

Scientific claim verification requires a system to retrieve relevant papers, isolate sentence-level evidence, and assign a SUPPORT, CONTRADICT, or not-enough-information label. This study evaluates a corpus-derived GraphRAG design on SciFact using the five predefined cross-validation folds. The system represents 5,183 papers, 45,952 abstract sentences, and 21,025 scientific entity phrases as a heterogeneous graph with 459,461 typed edges. A 256-dimensional latent-semantic Vector RAG baseline is compared with graph propagation and a training-fold-selected hybrid fusion. Across 1,109 out-of-fold claims, the Hybrid system achieved document recall@5 of 0.7519, compared with 0.6421 for Vector RAG and 0.6654 for GraphRAG. Complete evidence-set recall@5 reached 0.4645, an absolute gain of 0.1082 over Vector RAG. GraphRAG produced the strongest retrieval-augmented verification accuracy (0.4598) and strict grounded accuracy (0.3462), while the claim-only classifier retained the highest label accuracy (0.4860). Paired analysis confirmed that Hybrid improved evidence-set success over Vector RAG, but its grounded-label gain was not statistically significant. The results establish that graph structure reliably broadens evidence coverage and improves resilience to query token loss, while downstream verification still depends on relation-sensitive reasoning that cannot be replaced by retrieval depth alone.

References

Conf. Empirical Methods in Natural Language Processing, 2020, pp. 7534–7550, doi: 10.18653/v1/2020.emnlp-main.609.

[2] J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal, “FEVER: A Large-Scale Dataset for Fact Extraction and VERification,” in Proc. NAACL-HLT, 2018, pp. 809–819, doi: 10.18653/v1/N18-1074.

[3] N. Kotonya and F. Toni, “Explainable Automated Fact-Checking for Public Health Claims,” in Proc. 2020 Conf. Empirical Methods in Natural Language Processing, 2020, pp. 7740–7754, doi: 10.18653/v1/2020.emnlp-main.623.

[4] A. Saakyan, T. Chakrabarty, and S. Muresan, “COVID-Fact: Fact Extraction and Verification of Real-World Claims on COVID-19,” in Proc. 59th Annu. Meeting Assoc. Comput. Linguistics, 2021, pp. 2116–2129, doi: 10.18653/v1/2021.acl-long.165.

[5] T. Diggelmann, J. Boyd-Graber, J. Bulian, M. Ciaramita, and M. Leippold, “CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims,” arXiv:2012.00614, 2020.

[6] J. Vladika and F. Matthes, “Scientific Fact-Checking: A Survey of Resources and Approaches,” in Findings of the Association for Computational Linguistics: ACL 2023, 2023, pp. 6215–6230, doi: 10.18653/v1/2023.findings-acl.387.

[7] R. Pradeep, X. Ma, R. Nogueira, and J. Lin, “Scientific Claim Verification with VerT5erini,” in Proc. 12th Int. Workshop Health Text Mining and Information Analysis, 2021, pp. 94–103, doi: 10.18653/v1/2021.louhi-1.11.

[8] Z. Zhang, J. Li, F. Fukumoto, and Y. Ye, “Abstract, Rationale, Stance: A Joint Model for Scientific Claim Verification,” in Proc. 2021 Conf. Empirical Methods in Natural Language Processing, 2021, pp. 3580–3586, doi: 10.18653/v1/2021.emnlp-main.290.

[9] D. Wadden, K. Lo, L. L. Wang, S. Lin, M. van Zuylen, A. Cohan, and I. Beltagy, “MultiVerS: Improving Scientific Claim Verification with Weak Supervision and Full-Document Context,” in Findings of NAACL 2022, 2022, pp. 61–76, doi: 10.18653/v1/2022.findings-naacl.6.

[10] P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 9459–9474.

[11] V. Karpukhin et al., “Dense Passage Retrieval for Open-Domain Question Answering,” in Proc. 2020 Conf. Empirical Methods in Natural Language Processing, 2020, pp. 6769–6781, doi: 10.18653/v1/2020.emnlp-main.550.

[12] K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang, “REALM: Retrieval-Augmented Language Model Pre-Training,” in Proc. 37th Int. Conf. Machine Learning, 2020, pp. 3929–3938.

[13] G. Izacard and E. Grave, “Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering,” in Proc. 16th Conf. European Chapter Assoc. Comput. Linguistics, 2021, pp. 874–880, doi: 10.18653/v1/2021.eacl-main.74.

[14] G. Izacard et al., “Atlas: Few-Shot Learning with Retrieval Augmented Language Models,” Journal of Machine Learning Research, vol. 24, no. 251, pp. 1–43, 2023.

[15] F. Petroni et al., “KILT: A Benchmark for Knowledge Intensive Language Tasks,” in Proc. NAACL-HLT, 2021, pp. 2523–2544, doi: 10.18653/v1/2021.naacl-main.200.

[16] I. Beltagy, K. Lo, and A. Cohan, “SciBERT: A Pretrained Language Model for Scientific Text,” in Proc. EMNLP-IJCNLP, 2019, pp. 3615–3620, doi: 10.18653/v1/D19-1371.

[17] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks,” in Proc. EMNLP-IJCNLP, 2019, pp. 3982–3992, doi: 10.18653/v1/D19-1410.

[18] H. Sun, B. Dhingra, M. Zaheer, K. Mazaitis, R. Salakhutdinov, and W. Cohen, “Open Domain Question Answering Using Early Fusion of Knowledge Bases and Text,” in Proc. 2018 Conf. Empirical Methods in Natural Language Processing, 2018, pp. 4231–4242, doi: 10.18653/v1/D18-1455.

[19] H. Sun, T. Bedrax-Weiss, and W. Cohen, “PullNet: Open Domain Question Answering with Iterative Retrieval on Knowledge Bases and Text,” in Proc. EMNLP-IJCNLP, 2019, pp. 2380–2390, doi: 10.18653/v1/D19-1242.

[20] M. Yasunaga, H. Ren, A. Bosselut, P. Liang, and J. Leskovec, “QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering,” in Proc. NAACL-HLT, 2021, pp. 535–546, doi: 10.18653/v1/2021.naacl-main.45.

[21] M. Yasunaga, A. Bosselut, H. Ren, X. Zhang, C. D. Manning, P. Liang, and J. Leskovec, “Deep Bidirectional Language-Knowledge Graph Pretraining,” in Proc. Int. Conf. Learning Representations, 2022.

[22] D. Yu, C. Zhu, Y. Fang, W. Yu, S. Wang, Y. Xu, X. Ren, Y. Yang, and M. Zeng, “KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question Answering,” in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics, 2022, pp. 4961–4974, doi: 10.18653/v1/2022.acl-long.340.

[23] T. N. Kipf and M. Welling, “Semi-Supervised Classification with Graph Convolutional Networks,” in Proc. Int. Conf. Learning Representations, 2017.

[24] P. W. Battaglia et al., “Relational Inductive Biases, Deep Learning, and Graph Networks,” arXiv:1806.01261, 2018.

[25] R. Mihalcea and P. Tarau, “TextRank: Bringing Order into Text,” in Proc. 2004 Conf. Empirical Methods in Natural Language Processing, 2004, pp. 404–411.

[26] S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman, “Indexing by Latent Semantic Analysis,” Journal of the American Society for Information Science, vol. 41, no. 6, pp. 391–407, 1990.

[27] G. Salton and C. Buckley, “Term-Weighting Approaches in Automatic Text Retrieval,” Information Processing & Management, vol. 24, no. 5, pp. 513–523, 1988.

[28] L. Page, S. Brin, R. Motwani, and T. Winograd, “The PageRank Citation Ranking: Bringing Order to the Web,” Stanford InfoLab, Technical Report 1999-66, 1999.

[29] Y. Li, “Findable then Explainable: Retrieval–Summary Integration for Code Intelligence on a Lightweight CodeSearchNet Subset,” J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65–82, Jul. 2024, doi: 10.69987/JACS.2024.40706.

[30] S. Zhou, Z. Li, and E. Wang, “Long-Document RAG for Contractual and Insurance Clause Analysis in Receivables RWA Structures,” J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88–104, Aug. 2024, doi: 10.69987/JACS.2024.40810.

[31] J. Nie and D. Zheng, “Ambiguity-Aware HDFS Log Anomaly Detection with Retrieval-Augmented Failure Narratives and Selective Refusal,” J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66–80, Jan. 2023, doi: 10.69987/JACS.2023.30105.

[32] S. Zhou, Z. Li, and E. Wang, “Evidence-Grounded RAG for Tokenized Trade Receivable Disclosure QA under U.S. Capital Market Standards,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41–57, Jul. 2023, doi: 10.69987/JACS.2023.30704.

[33] Z. S. Zhong, R. Ma, and H. Zhao, “Human-Uncertainty Distillation for Calibrated Vision Models on CIFAR-10H,” J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77–89, Feb. 2023, doi: 10.69987/JACS.2023.30206.

[34] Z. S. Zhong and S. Ling, “Improved Theoretical Guarantee for Rank Aggregation via Spectral Method,” Inf. Inference: J. IMA, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.

[35] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, “Autonomous Driving System Driven by Artificial Intelligence Perception Fusion,” Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193–198, Feb. 2024, doi: 10.54097/e0b9ak47.

[36] H. Zhang, “DriftGuard: Multi-Signal Drift Early Warning and Safe Re-Training/Rollback for CTR/CVR Models,” J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24–40, Jul. 2023, doi: 10.69987/JACS.2023.30703.

[37] Q. Xin, “Uncertainty-Aware Late Fusion for 3D Perception (Confidence Calibration + Fusion Rule Learning),” J. Technol. Informatics Eng., vol. 4, no. 1, pp. 215–238, 2025, doi: 10.51903/jtie.v4i1.485.

[38] Z. S. Zhong and S. Ling, “Uncertainty Quantification of Spectral Estimator and MLE for Orthogonal Group Synchronization,” arXiv:2408.05944, Aug. 2024.

[39] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, “IoT Traffic Classification and Anomaly Detection Method Based on Deep Autoencoders,” Appl. Comput. Eng., vol. 69, pp. 64–70, 2024, doi: 10.54254/2755-2721/69/20241511.

[40] J. Nie and D. Zheng, “Noisy-Neighbor-Aware VM Degradation Risk Modeling with Unsupervised Residual Fusion,” J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112–123, Apr. 2024, doi: 10.69987/JACS.2024.40409.

[41] Q. Xin, “Explaining OpenStack Failure-Injection Log Anomalies with Retrieved Normal Prototypes,” Emerg. Inf. Sci. Technol., vol. 6, no. 2, pp. 125–146, Nov. 2025, doi: 10.18196/eist.v6i2.31232.

[42] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, “Intelligent Fault Analysis with AIOps Technology,” J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94–100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.

[43] H. Zhang, “LLM-Driven CI Failure Diagnosis and Automated Repair: From GitHub Actions Logs to Patch Recommendation,” J. Technol. Informatics Eng., vol. 4, no. 1, pp. 190–214, 2025, doi: 10.51903/jtie.v4i1.484.

[44] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, “Optimization of Autonomous Driving Image Detection Based on RFAConv and Triplet Attention,” Appl. Comput. Eng., vol. 77, pp. 210–217, 2024, doi: 10.54254/2755-2721/77/2024MA0067.

[45] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, “Cancer Image Classification Based on DenseNet Model,” J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, Nov. 2020, doi: 10.1088/1742-6596/1651/1/012143.

[46] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, “Implementation of an AI-Based MRD Evaluation and Prediction Model for Multiple Myeloma,” Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127–131, Jan. 2024, doi: 10.54097/zJ4MnbWW.

[47] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, “Enhancing Bank Credit Risk Management Using the C5.0 Decision Tree Algorithm,” J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100–107, Nov. 2024, doi: 10.5281/zenodo.14032041.

[48] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, “Enhancing Financial Services through Big Data and AI-Driven Customer Insights and Risk Analysis,” J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53–62, Jul. 2024, doi: 10.60087/jklst.vol3.n3.p53-62.

[49] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, “Intelligent Classification and Personalized Recommendation of E-Commerce Products Based on Machine Learning,” Appl. Comput. Eng., vol. 64, pp. 147–153, 2024, doi: 10.54254/2755-2721/64/20241365.

[50] H. Zhang, “Risk-Aware Budget-Constrained Auto-Bidding under First-Price RTB: A Distributional Constrained Deep Reinforcement Learning Framework,” J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30–47, Jun. 2024, doi: 10.69987/JACS.2024.40603.

[51] Y. Lu, H. Zhou, and Y. Zhang, “A Constrained, Data-Driven Budgeting Framework Integrating Macro Demand Forecasting and Marketing Response Modeling,” J. Technol. Informatics Eng., vol. 4, no. 3, pp. 493–520, Dec. 2025, doi: 10.51903/jtie.v4i3.466.

[52] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, “Predicting Enterprise Marketing Decision Making with Intelligent Data-Driven Approaches,” J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12–19, Jun. 2024, doi: 10.5281/zenodo.11357252.

[53] H. Zhang, “Counterfactual Learning-to-Rank for Ads: Off-Policy Evaluation on the Open Bandit Dataset,” J. Adv. Comput. Syst., vol. 5, no. 12, pp. 1–11, 2025, doi: 10.69987/JACS.2025.51201.

[54] Q. Xin, “Hybrid Cloud Architecture for Efficient and Cost-Effective Large Language Model Deployment,” J. Inf. Syst. Informatics, vol. 7, no. 3, pp. 2182–2195, Sep. 2025, doi: 10.51519/journalisi.v7i3.1170.

[55] B. Zhou, H. Wang, and X. Chang, “Distilling VMAF into an Edge-Deployable Quality Predictor: A Pilot Shot-Level Proxy with LLM-Ready Quality Tokens,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 447–463, Aug. 2025, doi: 10.51903/jtie.v4i2.522.

[56] H. Wang, Y. Ren, and X. Chang, “Layout-Aware Progressive PDF Rendering: AI Prioritization of PDF Slices to Reduce Time-to-Functional-First-Frame on FUNSD,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 425–446, Aug. 2025, doi: 10.51903/jtie.v4i2.523.

[57] L. Zhang, R. Ma, and P. Greg, “Digital-Twin Dispatching for Urban Mobility via Spatio-Temporal Transformers and Offline Reinforcement Learning,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 337–363, Aug. 2025, doi: 10.51903/jtie.v4i2.501.

[58] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, “Predictive Optimization of DDoS Attack Mitigation in Distributed Systems Using Machine Learning,” Appl. Comput. Eng., vol. 64, pp. 94–99, 2024, doi: 10.54254/2755-2721/64/20241350.

[59] H. Zhou and K. Zhang, “News-Based Uncertainty and Macro-Market Fusion for VIX Direction Forecasting: Evidence from 2015–2024 FRED Panel,” J. Technol. Informatics Eng., vol. 4, no. 2, pp. 487–501, Aug. 2025, doi: 10.51903/jtie.v4i2.540.

[60] Z. S. Zhong, X. Pan, and Q. Lei, “Bridging Domains with Approximately Shared Features,” in Proc. 28th Int. Conf. Artificial Intelligence and Statistics (AISTATS), PMLR, vol. 258, 2025, pp. 559–567.

[61] A. G. S. Raj et al., “Impact of Bilingual CS Education on Student Learning and Engagement in a Data Structures Course,” in Proc. 19th Koli Calling Int. Conf. Comput. Educ. Res., 2019, Art. no. 24, pp. 1–10, doi: 10.1145/3364510.3364518.

[62] Z. Li, S. Zhou, and Z. Zhou, “Financial Risk Dashboard Design for Institutional RWA Investors: Visual Hierarchy, Chart Comprehension, and Explainability in FinChart-Bench,” Int. J. Graphic Design, vol. 3, no. 1, pp. 196–210, May 2025, doi: 10.51903/ijgd.v3i1.3715.

[63] R. Ma, L. Zhang, and T. Song, “Computer-Vision-Informed Visual Explanation Cards for Autonomous-Driving Traffic-Sign Alerts: Localization, Classification, and Retrieved Evidence on GTSDB,” Int. J. Graphic Design, vol. 3, no. 2, pp. 455–473, Oct. 2025, doi: 10.51903/ijgd.v3i2.3991.

Downloads

Published

2025-10-31

How to Cite

Miller, S., Chen, B., Peterson, J., Wang, Y., & Clark, D. (2025). GraphRAG for Scientific Evidence Verification through Claim–Document–Entity Graph Reasoning. Journal of Information Technology and Informatics Engineering, 1(2), 62-74. https://journal.jci.co.id/jitie/article/view/824