Home/Blog/AI Security
AI Security

LLM Security Guide: How AI Applications Get Attacked and How to Test Them

Most organisations do not deploy a bare language model. They deploy an application around one: a system prompt, retrieved documents, connections to internal tools, and code that acts on what the model says. That surrounding application is where most real security failures happen.

This guide walks through how LLM applications are built, where attackers get in, which controls help, and how to scope a security test that produces evidence rather than screenshots of a chatbot saying something rude. For the ranked list of risks, see OWASP Top 10 for LLM applications 2026. For governance, see the LLM security and governance checklist.

Key points

  • Treat the model as an untrusted component. Anything it reads can influence it, and anything it outputs can be wrong or hostile.
  • The biggest risks usually come from what the model can reach: documents, tools, APIs and other users' data.
  • Prompt injection cannot be fully prevented by filters. Limit what a successful injection can do.
  • Retrieval must enforce the same permissions as the rest of your application.
  • Testing should follow your architecture, use realistic roles, and report validated impact.

How an LLM application is put together

ComponentWhat it doesWhat can go wrong
System prompt and hidden contextSets behaviour and holds instructions the user should not seeLeaks configuration, secrets or business rules
User inputQuestions, uploads, chat historyDirect prompt injection and abuse
Retrieval (RAG)Pulls documents or records into the promptIndirect injection through content; data shown to the wrong user
Tools and agentsLets the model call APIs, send email, query databases or run codeActions taken with more privilege than the user has
Output handlingDisplays or executes what the model returnsCross-site scripting, injection into downstream systems, unsafe automation
Model and supply chainHosted or self-hosted models, libraries, plugins, fine-tuning dataCompromised components or poisoned data
Usage and cost controlsRate limits, quotas, monitoringRunaway cost, denial of service, unnoticed abuse

Draw this map for your own system before anything else. It is the basis for both your controls and your test scope.

Where attackers get in

Direct and indirect prompt injection

Direct injection is a user typing instructions that override the system prompt. Indirect injection is more dangerous: instructions hidden in a web page, email, PDF or support ticket that the model later reads. If the model can act on what it reads, a single poisoned document can trigger actions on behalf of whoever is using the assistant.

Data exposure through retrieval

If the retrieval layer searches everything and relies on the prompt to hide what a user should not see, it will eventually show it. Permissions must be enforced before documents reach the model, using the same identity and access rules as the rest of the application.

Excessive agency

An assistant that can send emails, change records or call admin APIs with a shared service account turns every injection into a potential incident. The question is not whether the model might be tricked. It is what happens when it is.

Unsafe output handling

Model output is user-controlled input by another route. If it is rendered as HTML, inserted into SQL, or passed to a shell or workflow engine without validation, classic web vulnerabilities return.

Cost, abuse and reliability

Long prompts, recursive tool calls and automated abuse can drive up cost or degrade service. Separately, confident but wrong answers become a business risk when staff or customers act on them without checking.

Mapping to the OWASP LLM Top 10 (2026)

OWASP IDCategoryWhere it shows up above
LLM01Prompt InjectionDirect and indirect injection
LLM02Sensitive Information DisclosureRetrieval exposure, leaked context
LLM03Excessive AgencyTools and agents
LLM04Supply ChainModels, libraries, plugins
LLM05Data and Model PoisoningFine-tuning data, poisoned retrieval content
LLM06Unbounded ConsumptionCost and denial of service
LLM07MisinformationWrong answers acted on without checks
LLM08Hidden Context ExposureSystem prompt and hidden instructions
LLM09Vector and Embedding WeaknessesRetrieval stores and embeddings
LLM10Improper Output HandlingOutput passed to browsers or downstream systems

Earlier versions of this page used the 2023 list. Categories such as Insecure Plugin Design, Overreliance and Model Theft no longer appear under those names; their concerns are now covered mainly by Excessive Agency, Misinformation and Unbounded Consumption.

Controls that reduce real risk

  1. Least privilege for tools: give each tool the narrowest permission it needs, and act with the end user's identity rather than a shared super-account.
  2. Human approval for high-impact actions: payments, deletions, external emails and permission changes should need confirmation.
  3. Permission-aware retrieval: filter documents by the requesting user's access before they reach the prompt.
  4. Treat output as untrusted: encode it for the context where it is displayed, and validate it before any system acts on it.
  5. Keep secrets out of prompts: assume the system prompt can be extracted.
  6. Limits and monitoring: rate limits, token and cost quotas, and logs that let you investigate what the model was asked and what it did.
  7. Evaluate before release: run adversarial test cases whenever prompts, models, tools or data sources change.

The UK government's Code of Practice for the Cyber Security of AI (January 2025) sets out similar baseline principles for developers and operators, and is a useful reference for governance discussions.

How to scope a useful LLM security test

  • Start from the architecture map, not from a list of prompts. Name every data source, tool and output consumer.
  • Provide realistic roles: at least two users with different permissions, so data boundaries can be tested.
  • Include indirect paths: documents, emails or web content the assistant ingests should be part of the test.
  • Test the tools: confirm what each tool can do when the model is manipulated, and whether user identity is enforced.
  • Cover the surrounding application: authentication, APIs and output rendering still need conventional web and API testing.
  • Agree reporting: findings should show the path from input to impact, with reproduction steps that account for non-deterministic responses.

Our LLM security assessment follows this approach and can be combined with web application and API testing of the same product.

Limits of LLM testing

  • Model behaviour is non-deterministic. A test shows what was achievable during the engagement, not every response the model could ever give.
  • Provider model updates can change behaviour without any change to your code. Retest after significant model or prompt changes.
  • Filters and guardrails reduce noise but do not remove the need to limit what a manipulated model can do.
  • Attack Vector identifies weaknesses and evidences impact. Remediation stays with your engineering and ML owners.

For an indicative estimate, use the instant quote calculator or email [email protected]. Engagements start from £1,500 for a small, well-defined job.

FAQ

Is prompt injection the same as jailbreaking?

They overlap. Jailbreaking usually means persuading a model to ignore its safety rules. Prompt injection is broader: any input, direct or hidden in content, that changes what the application does. In business systems, injection that triggers tool use or data access is the bigger concern.

Do we need to test if we use a major provider's model?

Yes. Providers secure their models and platforms. Your system prompt, retrieval sources, tool permissions and output handling are yours, and that is where most application-level weaknesses sit.

Can a guardrail product solve prompt injection?

It can reduce it. OWASP's own guidance is clear that prevention is unreliable, so design for containment: least privilege, approvals and permission-aware retrieval.

Should we test in production or a copy?

Usually a staging environment with the same prompts, tools and representative data, plus limited production checks where configuration differs. See staging or production for how to decide.

How often should LLM features be retested?

After significant changes to prompts, models, tools or data sources, and at least annually for customer-facing features.

Suggested Resources

#llm#ai-security#owasp

Ready to strengthen your security?

Talk to our consultants about your penetration testing requirements, or get a fast, transparent quote.