Entrada de Pandipedia
Actualizado el 16 Sept 20261 fuenteExplorar Pandipedia
Joan

Which cyberattack technique relies on adversarial text prompts to manipulate AI

The cyberattack technique that uses adversarial text prompts to manipulate large language models is known as prompt injection, which is often paired with LLM jailbreaking techniques [1].

In these attacks, actors insert crafted prompts designed to bypass model restrictions and force models to reveal hidden step-by-step internal reasoning, system instructions, or proprietary functionalities [2]. Threat actors have used these prompt manipulation techniques to extract reasoning capabilities and training data from frontier AI models [3].

Sigue explorando
Ver todo