E-08Adversarial Evals
Mix attack tasks into the benchmark and score refusals as wins.
Further reading
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents Debenedetti et al., NeurIPS 2024
- Jailbroken: How Does LLM Safety Training Fail? Wei et al., NeurIPS 2023