Pandipedia entry
Which cyberattack technique relies on adversarial text prompts to manipulate AI
The cyberattack technique that uses adversarial text prompts to manipulate large language models is known as prompt injection, which is often paired with LLM jailbreaking techniques [1].
In these attacks, actors insert crafted prompts designed to bypass model restrictions and force models to reveal hidden step-by-step internal reasoning, system instructions, or proprietary functionalities [2]. Threat actors have used these prompt manipulation techniques to extract reasoning capabilities and training data from frontier AI models [3].
Save this answer
Create your account to keep this answer and continue from it later.
Sorry, Pandi could not find an answer.
Let's look at alternatives:
- Modify the query.
- Start a new thread.
- Remove sources (if manually added).