S-03Direct Prompt Injection
The user themselves is the attacker: the prompt tries to override system rules and misuse the agent’s tools.
Further reading
- Jailbroken: How Does LLM Safety Training Fail? Wei et al., NeurIPS 2023
- Universal and Transferable Adversarial Attacks on Aligned Language Models Zou et al., 2023