Jailbreak-Resilient LLM Agents through Instruction Hierarchy, Input Gating, and Self-Verification
Keywords:
LLM Agents, Jailbreak Defense, Prompt Injection, Guardrails, Instruction Hierarchy, Input Gating, Self-Verification, Advbench, CLINC150Abstract
Language-model agents can convert untrusted text into tool calls, data access, and other consequential actions, so a jailbreak succeeds before a harmful completion is written whenever the request crosses the execution boundary. This study evaluates a four-stage boundary consisting of an unguarded baseline, a calibrated input classifier, an explicit instruction hierarchy, and an independent self-verifier. The experiments use all 520 AdvBench harmful goals and 22,500 in-scope CLINC150 utterances. AdvBench is evaluated with five-fold out-of-fold prediction; CLINC150 retains its official 15,000/3,000/4,500 train, validation, and test partitions. The input gate reduced raw attack success rate (ASR) from 100.00% to 0.38% at 0.25% benign over-refusal. Adding the instruction hierarchy blocked both remaining raw attacks, producing 0.00% raw ASR and 0.32% over-refusal. The full self-verifying pipeline preserved 0.00% raw ASR and reduced ASR under lexical obfuscation from 10.19% for the gate and 9.62% for the hierarchy to 2.12%, while benign over-refusal increased to 0.41%. Across raw, authority-override, role-play, and obfuscated conditions, the full pipeline achieved 0.53% ASR. Ablation, threshold, category, domain, latency, and paired-significance analyses show that the classifier supplies most raw-request discrimination, hierarchy rules close sparse policy misses and neutralize priority conflicts, and normalized self-verification contributes mainly under surface-form attacks. The best operating point depends on threat exposure: gate plus hierarchy maximizes measured utility on unmodified requests, whereas the full stack provides stronger obfuscation resilience at a small refusal and latency cost.
References
[1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, "Attention is all you need," in Advances in Neural Information Processing Systems, vol. 30, 2017.
[2] T. B. Brown et al., "Language models are few-shot learners," in Advances in Neural Information Processing Systems, vol. 33, pp. 1877-1901, 2020.
[3] R. Bommasani et al., "On the opportunities and risks of foundation models," arXiv:2108.07258, 2021.
[4] Z. Zhong, M. Zheng, H. Mai, J. Zhao, and X. Liu, "Cancer image classification based on DenseNet model," J. Phys.: Conf. Ser., vol. 1651, no. 1, Art. no. 012143, Nov. 2020, doi: 10.1088/1742-6596/1651/1/012143.
[5] L. Ouyang et al., "Training language models to follow instructions with human feedback," in Advances in Neural Information Processing Systems, vol. 35, pp. 27730-27744, 2022.
[6] Y. Bai et al., "Constitutional AI: Harmlessness from AI feedback," arXiv:2212.08073, 2022.
[7] OpenAI, "GPT-4 technical report," arXiv:2303.08774, 2023.
[8] L. Weidinger et al., "Taxonomy of risks posed by language models," in Proc. ACM Conf. Fairness, Accountability, and Transparency, 2022, pp. 214-229.
[9] J. Chen, J. Xiong, Y. Wang, Q. Xin, and H. Zhou, "Implementation of an AI-based MRD evaluation and prediction model for multiple myeloma," Front. Comput. Intell. Syst., vol. 6, no. 3, pp. 127-131, Jan. 2024, doi: 10.54097/zJ4MnbWW.
[10] A. Wei, N. Haghtalab, and J. Steinhardt, "Jailbroken: How does LLM safety training fail?" in Advances in Neural Information Processing Systems, vol. 36, 2023.
[11] A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, "Universal and transferable adversarial attacks on aligned language models," arXiv:2307.15043, 2023.
[12] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, "Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection," arXiv:2302.12173, 2023.
[13] F. Perez and I. Ribeiro, "Ignore previous prompt: Attack techniques for language models," arXiv:2211.09527, 2022.
[14] Q. Xin, Z. Xu, L. Guo, F. Zhao, and B. Wu, "IoT traffic classification and anomaly detection method based on deep autoencoders," in Proc. 6th Int. Conf. Comput. Data Sci. (CDS), 2024, pp. 64-70, doi: 10.54254/2755-2721/69/20241511.
[15] N. Jain et al., "Baseline defenses for adversarial attacks against aligned language models," arXiv:2309.00614, 2023.
[16] A. Robey, E. Wong, H. Hassani, and G. J. Pappas, "SmoothLLM: Defending large language models against jailbreaking attacks," arXiv:2310.03684, 2023.
[17] Z. Zhang, J. Yang, P. Ke, S. Cui, C. Zheng, H. Wang, and M. Huang, "Defending large language models against jailbreaking attacks through goal prioritization," arXiv:2311.09096, 2023.
[18] H. Inan et al., "Llama Guard: LLM-based input-output safeguard for human-AI conversations," arXiv:2312.06674, 2023.
[19] Z. Wang, F. Yang, L. Wang, P. Zhao, H. Wang, L. Chen, Q. Lin, and K.-F. Wong, "Self-Guard: Empower the LLM to safeguard itself," arXiv:2310.15851, 2023.
[20] J. Nie and D. Zheng, "Ambiguity-Aware HDFS Log Anomaly Detection with Retrieval-Augmented Failure Narratives and Selective Refusal," J. Adv. Comput. Syst., vol. 3, no. 1, pp. 66-80, Jan. 2023, doi: 10.69987/JACS.2023.30105.
[21] S. Zhou, Z. Li, and E. Wang, "Evidence-Grounded RAG for Tokenized Trade Receivable Disclosure QA under U.S. Capital Market Standards," J. Adv. Comput. Syst., vol. 3, no. 7, pp. 41-57, Jul. 2023, doi: 10.69987/JACS.2023.30704.
[22] S. Larson et al., "An evaluation dataset for intent classification and out-of-scope prediction," in Proc. Conf. Empirical Methods in Natural Language Processing and Int. Joint Conf. Natural Language Processing, 2019, pp. 1311-1316.
[23] Y. Wang, S. Du, Q. Xin, Y. He, and W. Qian, "Autonomous Driving System Driven by Artificial Intelligence Perception Fusion," Acad. J. Sci. Technol., vol. 9, no. 2, pp. 193-198, Feb. 2024, doi: 10.54097/e0b9ak47.
[24] J. Nie and D. Zheng, "Noisy-Neighbor-Aware VM Degradation Risk Modeling with Unsupervised Residual Fusion," J. Adv. Comput. Syst., vol. 4, no. 4, pp. 112-123, Apr. 2024, doi: 10.69987/JACS.2024.40409.
[25] A. G. S. Raj, H. Zhang, V. Abhyankar, S. Mukerjee, E. Zhang, J. Williams, R. Halverson, and J. M. Patel, "Impact of bilingual CS education on student learning and engagement in a data structures course," in Proc. 19th Koli Calling Int. Conf. Comput. Educ. Res., Koli, Finland, 2019, pp. 1-10, doi: 10.1145/3364510.3364518.
[26] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, "On calibration of modern neural networks," in Proc. 34th Int. Conf. Machine Learning, 2017, pp. 1321-1330.
[27] Z. S. Zhong, R. Ma, and H. Zhao, "Human-Uncertainty Distillation for Calibrated Vision Models on CIFAR-10H," J. Adv. Comput. Syst., vol. 3, no. 2, pp. 77-89, Feb. 2023, doi: 10.69987/JACS.2023.30206.
[28] M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh, "Beyond accuracy: Behavioral testing of NLP models with CheckList," in Proc. 58th Annu. Meeting Assoc. Comput. Linguistics, 2020, pp. 4902-4912.
[29] J. X. Morris, E. Lifland, J. Y. Yoo, J. Grigsby, D. Jin, and Y. Qi, "TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP," in Proc. Conf. Empirical Methods in Natural Language Processing: System Demonstrations, 2020, pp. 119-126.
[30] E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh, "Universal adversarial triggers for attacking and analyzing NLP," in Proc. Conf. Empirical Methods in Natural Language Processing and Int. Joint Conf. Natural Language Processing, 2019, pp. 2153-2162.
[31] B. Wang, Y. He, Z. Shui, Q. Xin, and H. Lei, "Predictive optimization of DDoS attack mitigation in distributed systems using machine learning," in Proc. 6th Int. Conf. Comput. Data Sci. (CDS), 2024, pp. 89-94, doi: 10.54254/2755-2721/64/20241350.
[32] E. Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. Gaithersburg, MD, USA: National Institute of Standards and Technology, 2023.
[33] D. Sculley et al., "Hidden technical debt in machine learning systems," in Advances in Neural Information Processing Systems, vol. 28, 2015.
[34] S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, "RealToxicityPrompts: Evaluating neural toxic degeneration in language models," in Findings of the Association for Computational Linguistics: EMNLP, 2020, pp. 3356-3369.
[35] P. Rottger, B. Vidgen, D. Nguyen, Z. Waseem, H. Margetts, and J. Pierrehumbert, "HateCheck: Functional tests for hate speech detection models," in Proc. 59th Annu. Meeting Assoc. Comput. Linguistics and 11th Int. Joint Conf. Natural Language Processing, 2021, pp. 41-58.
[36] S. Lin, J. Hilton, and O. Evans, "TruthfulQA: Measuring how models mimic human falsehoods," in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics, 2022, pp. 3214-3252.
[37] D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt, "Aligning AI with shared human values," in Proc. Int. Conf. Learning Representations, 2021.
[38] Z. Ling, Q. Xin, Y. Lin, G. Su, and Z. Shui, "Optimization of autonomous driving image detection based on RFAConv and triplet attention," in Proc. 2nd Int. Conf. Softw. Eng. Mach. Learn. (SEML), 2024, pp. 210-217, doi: 10.54254/2755-2721/77/2024MA0067.
[39] L. Dixon, J. Li, J. Sorensen, N. Thain, and L. Vasserman, "Measuring and mitigating unintended bias in text classification," in Proc. AAAI/ACM Conf. AI, Ethics, and Society, 2018, pp. 67-73.
[40] K. Xu, H. Zhou, H. Zheng, M. Zhu, and Q. Xin, "Intelligent classification and personalized recommendation of e-commerce products based on machine learning," in Proc. 6th Int. Conf. Comput. Data Sci. (ICCDS), 2024, doi: 10.54254/2755-2721/64/20241365.
[41] Y. Li, "Findable then Explainable: Retrieval-Summary Integration for Code Intelligence on a Lightweight CodeSearchNet Subset," J. Adv. Comput. Syst., vol. 4, no. 7, pp. 65-82, Jul. 2024, doi: 10.69987/JACS.2024.40706.
[42] Q. Xin, R. Song, Z. Wang, Z. Xu, and F. Zhao, "Enhancing Bank Credit Risk Management Using the C5.0 Decision Tree Algorithm," J. Comput. Technol. Appl. Math., vol. 1, no. 4, pp. 100-107, Nov. 2024, doi: 10.5281/zenodo.14032041.
[43] Z. S. Zhong and S. Ling, "Improved theoretical guarantee for rank aggregation via spectral method," Inf. Inference: J. IMA, vol. 13, no. 3, Art. no. iaae020, Sep. 2024, doi: 10.1093/imaiai/iaae020.
[44] H. Zhang, "DriftGuard: Multi-Signal Drift Early Warning and Safe Re-Training/Rollback for CTR/CVR Models," J. Adv. Comput. Syst., vol. 3, no. 7, pp. 24-40, Jul. 2023, doi: 10.69987/JACS.2023.30703.
[45] Z. S. Zhong and S. Ling, "Uncertainty Quantification of Spectral Estimator and MLE for Orthogonal Group Synchronization," arXiv:2408.05944, Aug. 2024.
[46] T. Yang, Q. Xin, X. Zhan, S. Zhuang, and H. Li, "Enhancing Financial Services through Big Data and AI-Driven Customer Insights and Risk Analysis," J. Knowl. Learn. Sci. Technol., vol. 3, no. 3, pp. 53-62, Jul. 2024, doi: 10.60087/jklst.vol3.n3.p53-62.
[47] H. Zhang, "Risk-Aware Budget-Constrained Auto-Bidding under First-Price RTB: A Distributional Constrained Deep Reinforcement Learning Framework," J. Adv. Comput. Syst., vol. 4, no. 6, pp. 30-47, Jun. 2024, doi: 10.69987/JACS.2024.40603.
[48] E. Perez et al., "Red teaming language models with language models," in Proc. Conf. Empirical Methods in Natural Language Processing, 2022, pp. 3419-3448.
[49] Y. He, Y. Pan, Y. Wang, S. Du, and Q. Xin, "Intelligent Fault Analysis with AIOps Technology," J. Theory Pract. Eng. Sci., vol. 4, no. 1, pp. 94-100, Feb. 2024, doi: 10.53469/jtpes.2024.04(01).13.
[50] S. Zhou, Z. Li, and E. Wang, "Long-Document RAG for Contractual and Insurance Clause Analysis in Receivables RWA Structures," J. Adv. Comput. Syst., vol. 4, no. 8, pp. 88-104, Aug. 2024, doi: 10.69987/JACS.2024.40810.
[51] J. Wang, Q. Xin, Y. Liu, J. Wang, and T. Yang, "Predicting Enterprise Marketing Decision Making with Intelligent Data-Driven Approaches," J. Ind. Eng. Appl. Sci., vol. 2, no. 3, pp. 12-19, Jun. 2024, doi: 10.5281/zenodo.11357252.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 Robert Garcia, Qiang Li, Amy Baker, Fan Wu (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Creative Commons Attribution 4.0 International (CC BY 4.0).




This work is licensed under a