How should I handle prompt injection and other generative AI-specific threats?
Contrasting
Assume prompts can subvert system instructions. Do not rely on secret prompt structure or vendor resilience alone — add filtering, logging, human oversight, and continuous re-testing as models change.
The Playbook describes an architectural inability to distinguish user prompts from system instructions. AI Insights says most vendor solutions are quite resilient to related vulnerabilities, while still requiring continuous vigilance and non-secret defences.
How to navigate this: Treat vendor resilience as helpful, not sufficient. Retain Playbook assumptions about prompt subversion.
From the guidance
Primary (how) AI Insights: Prompt Risks
Section: Prompt injection
Primary (how) AI Insights: Prompt Risks
Section: Vigilance
Contrasting AI Playbook for the UK Government
Section: Prompt injection
Read this in AI Playbook for the UK Government (opens in new tab)
Contrasting AI Insights: Prompt Risks
Section: Prompt injection