CySecBench: A Cybersecurity Prompt Dataset for Benchmarking LLM Jailbreaks
A publicly released benchmark of 12,662 close-ended prompts, organized into 10 cyber-attack-type categories, for measuring how well large language models resist jailbreaking specifically in the cybersecurity domain. Unlike broad, open-ended jailbreak datasets, CySecBench's domain-specific, close-ended design yields more consistent and accurate assessment of attack effectiveness, and its documented generative-AI generation-and-filtration pipeline transfers to other domains. Demonstrating the benchmark, the authors introduce a prompt-obfuscation jailbreak method with attack success rates of 65% (ChatGPT), 88% (Gemini), and 17% (the more resilient Claude), plus 78.5% on AdvBench, exceeding prior methods and illustrating the value of domain-specific evaluation datasets for LLM security.
CySecBench: Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large
A domain-specific benchmark for evaluating how well large language models resist jailbreaking in the cybersecurity domain. Existing jailbreak datasets are broad and open-ended, making it hard to asse…