Most organisations do not deploy a bare language model. They deploy an application around one: a system prompt, retrieved documents, connections to internal tools, and code that acts on what the model says. That surrounding application is where most real security failures happen.
This guide walks through how LLM applications are built, where attackers get in, which controls help, and how to scope a security test that produces evidence rather than screenshots of a chatbot saying something rude. For the ranked list of risks, see OWASP Top 10 for LLM applications 2026. For governance, see the LLM security and governance checklist.
Key points
- Treat the model as an untrusted component. Anything it reads can influence it, and anything it outputs can be wrong or hostile.
- The biggest risks usually come from what the model can reach: documents, tools, APIs and other users' data.
- Prompt injection cannot be fully prevented by filters. Limit what a successful injection can do.
- Retrieval must enforce the same permissions as the rest of your application.
- Testing should follow your architecture, use realistic roles, and report validated impact.
How an LLM application is put together
| Component | What it does | What can go wrong |
|---|---|---|
| System prompt and hidden context | Sets behaviour and holds instructions the user should not see | Leaks configuration, secrets or business rules |
| User input | Questions, uploads, chat history | Direct prompt injection and abuse |
| Retrieval (RAG) | Pulls documents or records into the prompt | Indirect injection through content; data shown to the wrong user |
| Tools and agents | Lets the model call APIs, send email, query databases or run code | Actions taken with more privilege than the user has |
| Output handling | Displays or executes what the model returns | Cross-site scripting, injection into downstream systems, unsafe automation |
| Model and supply chain | Hosted or self-hosted models, libraries, plugins, fine-tuning data | Compromised components or poisoned data |
| Usage and cost controls | Rate limits, quotas, monitoring | Runaway cost, denial of service, unnoticed abuse |
Draw this map for your own system before anything else. It is the basis for both your controls and your test scope.
Where attackers get in
Direct and indirect prompt injection
Direct injection is a user typing instructions that override the system prompt. Indirect injection is more dangerous: instructions hidden in a web page, email, PDF or support ticket that the model later reads. If the model can act on what it reads, a single poisoned document can trigger actions on behalf of whoever is using the assistant.
Data exposure through retrieval
If the retrieval layer searches everything and relies on the prompt to hide what a user should not see, it will eventually show it. Permissions must be enforced before documents reach the model, using the same identity and access rules as the rest of the application.
Excessive agency
An assistant that can send emails, change records or call admin APIs with a shared service account turns every injection into a potential incident. The question is not whether the model might be tricked. It is what happens when it is.
Unsafe output handling
Model output is user-controlled input by another route. If it is rendered as HTML, inserted into SQL, or passed to a shell or workflow engine without validation, classic web vulnerabilities return.
Cost, abuse and reliability
Long prompts, recursive tool calls and automated abuse can drive up cost or degrade service. Separately, confident but wrong answers become a business risk when staff or customers act on them without checking.
Mapping to the OWASP LLM Top 10 (2026)
| OWASP ID | Category | Where it shows up above |
|---|---|---|
| LLM01 | Prompt Injection | Direct and indirect injection |
| LLM02 | Sensitive Information Disclosure | Retrieval exposure, leaked context |
| LLM03 | Excessive Agency | Tools and agents |
| LLM04 | Supply Chain | Models, libraries, plugins |
| LLM05 | Data and Model Poisoning | Fine-tuning data, poisoned retrieval content |
| LLM06 | Unbounded Consumption | Cost and denial of service |
| LLM07 | Misinformation | Wrong answers acted on without checks |
| LLM08 | Hidden Context Exposure | System prompt and hidden instructions |
| LLM09 | Vector and Embedding Weaknesses | Retrieval stores and embeddings |
| LLM10 | Improper Output Handling | Output passed to browsers or downstream systems |
Earlier versions of this page used the 2023 list. Categories such as Insecure Plugin Design, Overreliance and Model Theft no longer appear under those names; their concerns are now covered mainly by Excessive Agency, Misinformation and Unbounded Consumption.
Controls that reduce real risk
- Least privilege for tools: give each tool the narrowest permission it needs, and act with the end user's identity rather than a shared super-account.
- Human approval for high-impact actions: payments, deletions, external emails and permission changes should need confirmation.
- Permission-aware retrieval: filter documents by the requesting user's access before they reach the prompt.
- Treat output as untrusted: encode it for the context where it is displayed, and validate it before any system acts on it.
- Keep secrets out of prompts: assume the system prompt can be extracted.
- Limits and monitoring: rate limits, token and cost quotas, and logs that let you investigate what the model was asked and what it did.
- Evaluate before release: run adversarial test cases whenever prompts, models, tools or data sources change.
The UK government's Code of Practice for the Cyber Security of AI (January 2025) sets out similar baseline principles for developers and operators, and is a useful reference for governance discussions.
How to scope a useful LLM security test
- Start from the architecture map, not from a list of prompts. Name every data source, tool and output consumer.
- Provide realistic roles: at least two users with different permissions, so data boundaries can be tested.
- Include indirect paths: documents, emails or web content the assistant ingests should be part of the test.
- Test the tools: confirm what each tool can do when the model is manipulated, and whether user identity is enforced.
- Cover the surrounding application: authentication, APIs and output rendering still need conventional web and API testing.
- Agree reporting: findings should show the path from input to impact, with reproduction steps that account for non-deterministic responses.
Our LLM security assessment follows this approach and can be combined with web application and API testing of the same product.
Limits of LLM testing
- Model behaviour is non-deterministic. A test shows what was achievable during the engagement, not every response the model could ever give.
- Provider model updates can change behaviour without any change to your code. Retest after significant model or prompt changes.
- Filters and guardrails reduce noise but do not remove the need to limit what a manipulated model can do.
- Attack Vector identifies weaknesses and evidences impact. Remediation stays with your engineering and ML owners.
For an indicative estimate, use the instant quote calculator or email [email protected]. Engagements start from £1,500 for a small, well-defined job.
FAQ
Is prompt injection the same as jailbreaking?
They overlap. Jailbreaking usually means persuading a model to ignore its safety rules. Prompt injection is broader: any input, direct or hidden in content, that changes what the application does. In business systems, injection that triggers tool use or data access is the bigger concern.
Do we need to test if we use a major provider's model?
Yes. Providers secure their models and platforms. Your system prompt, retrieval sources, tool permissions and output handling are yours, and that is where most application-level weaknesses sit.
Can a guardrail product solve prompt injection?
It can reduce it. OWASP's own guidance is clear that prevention is unreliable, so design for containment: least privilege, approvals and permission-aware retrieval.
Should we test in production or a copy?
Usually a staging environment with the same prompts, tools and representative data, plus limited production checks where configuration differs. See staging or production for how to decide.
How often should LLM features be retested?
After significant changes to prompts, models, tools or data sources, and at least annually for customer-facing features.
Suggested Resources
Ready to strengthen your security?
Talk to our consultants about your penetration testing requirements, or get a fast, transparent quote.