Engineering Approaches to Improving Security and Reliability in Generative Artificial Intelligence Systems
Main Article Content
Abstract
Downloads
Article Details
Data Availability Statement
Section

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
References
Bin, W., Jiazheng, Q., Yu, X., Hansen, H., Hao, Y., Gao, A., Wan, Z., Li, H., & Tsang, I. (2026). Safety Sidecar: Reflection-Driven Runtime Control for Safer Agents. Findings of the Association for Computational Linguistics: ACL 2026, pp. 30842–30856. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.1542
Chang, Z., Li, M., Huang, Y., Jiang, Z., Jia, X., Xiong, Q., Wang, J., Li, Z., & Wang, Q. (2026). Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning. Findings of the Association for Computational Linguistics: ACL 2026, pp. 3540–3561. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.findings-acl.173
Wang, X., Jian, S., Li, S., Li, X., Li, Z., Ji, B., Wang, B., & Yu, J. (2026). JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, pp. 7661–7674. Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.acl-long.348
Chen, Y., Li, H., Zheng, Z., Wu, D., Song, Y., & Hooi, B. (2025). Defense Against Prompt Injection Attack by Leveraging Attack Techniques. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, pp. 18331–18347. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl- long.897
Li, H., Liu, X., Zhang, N., & Xiao, C. (2025). PIGuard: Prompt Injection Guardrail via Mitigating Overdefense for Free. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, pp. 30420–30437. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.1468
Hackett, W., Birch, L., Trawicki, S., Suri, N., & Garraghan, P. (2025). Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems. Proceedings of the First Workshop on LLM Security (LLMSEC), pp. 101–114. Association for Computational Linguistics.
Wang, S., Wang, X., Mei, J., Xie, Y., Chen, S.-Q., & Xiong, W. (2025). Developing a Reliable, Fast, General-Purpose Hallucination Detection and Mitigation Service. Proceedings of NAACL 2025: Industry Track, pp. 971–978. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-industry.72
Xu, Z., Liu, Y., Deng, G., Li, Y., & Picek, S. (2024). A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024, pp. 7432–7449. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.443
Xie, Y., Fang, M., Pi, R., & Gong, N. (2024). GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pp. 507–518. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.30
Zhao, W., Li, Z., Li, Y., Zhang, Y., & Sun, J. (2024). Defending Large Language Models Against Jailbreak Attacks via Layer-Specific Editing. Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 5094–5109. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.293