Large language models do not separate “instructions” from “data” the way a parameterised SQL query does. Everything arrives as tokens: system text, user chat, retrieved documents, tool replies, memory. That is why LLM applications fail in ways the classic web Top 10 only partly describes. If assurance only covers XSS and missing host patches, you are measuring the wrong surface for an agent that can read tickets and call privileged APIs.
The OWASP GenAI Security Project maintains the Top 10 for LLM Applications as the shared risk vocabulary. The current edition is 2026 (published August 2026).
This is a risk-taxonomy explainer. For leadership actions, see the governance checklist. For adversary tactics, see MITRE ATLAS. Keep them separate.
At a glance
- Official 2026 edition: OWASP GenAI LLM Top 10 2026.
- 2026 keeps Prompt Injection and Sensitive Information Disclosure at #1 and #2.
- Excessive Agency rises to #3; Unbounded Consumption rises; Improper Output Handling falls to #10.
- System Prompt Leakage is renamed and widened to Hidden Context Exposure.
- Ranking used community input plus incident evidence. It remains an awareness list, not a UK legal mandate.
Context: why a separate Top 10 exists
The web application Top 10 (2025 edition explained here) still applies to the HTTP APIs, auth and cloud around an LLM feature. The LLM list covers model-specific and orchestration-specific failures: injection through content the model reads, over-permissioned tools, retrieval leakage, poisoned corpora, and cost amplification.
Attack Vector is UK-based. LLM assessments identify weaknesses and evidence impact. Remediation stays with your engineering and ML owners. Lead credentials include CSTL and ChCSP. Testing starts from £1,500 for tightly defined scopes.
What changed from 2025 to 2026
Eight of ten entries moved. Categories were not replaced wholesale. The table compares the 2026 ranking with 2025:
| 2026 | Risk | vs 2025 |
|---|---|---|
| LLM01 | Prompt Injection | Unchanged |
| LLM02 | Sensitive Information Disclosure | Unchanged |
| LLM03 | Excessive Agency | Up from #6 |
| LLM04 | Supply Chain | Down from #3 |
| LLM05 | Data and Model Poisoning | Down from #4 |
| LLM06 | Unbounded Consumption | Up from #10 |
| LLM07 | Misinformation | Up from #9 |
| LLM08 | Hidden Context Exposure | Renamed/widened from System Prompt Leakage; down one |
| LLM09 | Vector and Embedding Weaknesses | Down from #8 |
| LLM10 | Improper Output Handling | Down from #5 |
The practical message from the 2026 emphasis: stop pretending you can permanently “filter” the model into obedience. Bound what a fooled model can reach.
The 2026 risks, in plain English
LLM01 Prompt Injection
Untrusted content alters model behaviour. Direct chat abuse is only one path. Indirect injection plants instructions in documents, tickets, emails or tool output that the application later feeds to the model. Defence is mostly architectural: reduce privileges, require confirmation for risky actions, and treat retrieved text as hostile.
LLM02 Sensitive Information Disclosure
The system reveals secrets, personal data or proprietary material through answers, tool arguments, logs, traces or retrieval. Sometimes the bug is memorisation. More often the retrieval index simply contains documents the user should never have searched.
LLM03 Excessive Agency
The model can act: send mail, change tickets, run queries, call payment APIs. Excessive functionality, excessive permissions or excessive autonomy turns a manipulated completion into a real-world incident. This is why agency rose in 2026.
LLM04 Supply Chain
Third-party models, datasets, adapters, conversion tools and serving stacks. Unsafe model serialisation, poisoned packages (including names suggested by coding assistants), and compromised registries all sit here.
LLM05 Data and Model Poisoning
Training, fine-tuning, embedding or RAG corpora corrupted so harmful behaviour is baked in. Triggers may stay dormant until a phrase appears. Fixes often mean retrain or replace, not a one-line patch.
LLM06 Unbounded Consumption
Attackers spend little to force expensive compute, long agent loops or huge outputs. Think Denial-of-Wallet as well as classic availability loss. Rate limits alone may fail when one request fans out into many tool calls.
LLM07 Misinformation
Confident wrong output that people or downstream automations act on. Higher stakes when answers drive tools or other agents without a human gate.
LLM08 Hidden Context Exposure
Formerly framed mainly as system-prompt leakage. Now covers non-user-facing context more broadly: hidden instructions, tool schemas, retrieved policy text, workflow rules. Assume discovery is possible; design so disclosure does limited harm (no secrets in the prompt).
LLM09 Vector and Embedding Weaknesses
Similarity search and embedding stores become trust boundaries. Cross-tenant inference, poisoned embeddings and access-control gaps around vector DBs are in scope even when returned text looks “clean.”
LLM10 Improper Output Handling
Model output reaches a dangerous sink without validation: HTML, SQL, shell, Markdown image fetches, terminal escape sequences. Familiar web fixes (encoding, parameterisation, allow-lists) apply. It fell from fifth to tenth in the 2026 ranking, but the risk has not gone away, and the entry now also covers insecure code generated by AI assistants.
How to assess LLM applications properly
A useful LLM application assessment usually:
- Inventories context sources and tool identities.
- Tests indirect injection before chasing clever chat jailbreaks.
- Follows outputs into every sink.
- Probes retrieval tenancy and access control.
- Measures cost/compute amplification under abuse.
- Reports validated impact, not only screenshots of rude replies.
Pair technical testing with the governance checklist so ownership exists for inventory, policy and incident response. Map interesting behaviours to MITRE ATLAS when stakeholders want ATT&CK-style IDs.
Why filters and tools are not enough
Two common buyer mistakes:
- Filter myth: believing a second model or keyword list makes prompt injection “solved.” OWASP’s own prevention narrative stresses unreliable prevention and the need to bound blast radius.
- Tool myth: shipping an agent with a powerful service account “just for the demo,” then forgetting to revoke it.
A third mistake is commissioning a generic web test that never touches retrieval, tools or multi-tenant indexes, then claiming LLM Top 10 coverage.
What the LLM Top 10 is and is not
- The OWASP LLM Top 10 is an awareness document, not a UK statute or a certificate.
- Edition years matter. Prefer 2026 wording in new briefs; treat older 2025-only materials as historical.
- Attack Vector identifies weaknesses. We do not remediate your model stack inside the standard test fee.
- No assessment guarantees a model cannot be manipulated.
For a scoped assessment of an LLM-enabled product, see LLM security assessment, try the instant quote calculator, or email [email protected]. Broader services: penetration testing services.
FAQ
Which edition should we cite in 2026 RFPs?
Cite the OWASP Top 10 for LLM Applications 2026 unless a customer contract freezes an older edition. Link the 2026 resource page in your methodology appendix.
Is prompt injection actually preventable?
Not reliably as a pure model problem. Reduce impact by limiting tools, data access and outbound channels so a successful injection cannot do much.
Why did Excessive Agency rise?
Because agentic systems made tool misuse concrete. A manipulated answer that can modify records or send messages is a different risk class from a chat-only toy.
How does this relate to the web OWASP Top 10?
Use both. Web Top 10 for the surrounding application and infrastructure; LLM Top 10 for model orchestration, retrieval and tool risks.
Do we need machine-learning researchers to test this?
Often no. Many findings are trust-boundary and authorisation failures familiar to application testers, plus LLM-specific abuse cases. Deep model research helps some specialised threats; it is not the whole engagement.
Will Attack Vector fix prompt injection for us?
No. We evidence weaknesses and provide remediation guidance. Engineering owns control changes; retest can verify closures when scoped.
Suggested Resources
Ready to strengthen your security?
Talk to our consultants about your penetration testing requirements, or get a fast, transparent quote.