TL;DR
- Prompt injection matters because it can be the route into much bigger problems.
- The real risk comes from what the model can read, what it can reach, and what the application decides to trust.
- Most serious LLM failures are not single neat issues. They are combinations of data exposure, weak permissions, poor output handling, unsafe RAG, and over-trust.
- Treat the LLM as an untrusted component. Limit data, limit tools, validate output, and keep humans in the loop where judgement matters.
Read, reach and trust
Prompt injection has gone mainstream. We have all seen the screenshots: someone tells a chatbot to ignore its previous instructions, it says something daft, and the internet has a good laugh. That does not make prompt injection trivial. It just means the easy examples sometimes hide why it matters.
From a security testing point of view, prompt injection is often the starting point rather than the whole finding. A chatbot saying something embarrassing is one thing. A chatbot with access to customer records, internal documents, source code, admin functions, or live booking systems is a very different problem.
The question I care about on an engagement is not only “can I make the model ignore its instructions?” It is: “if I can influence the model, what will that let me do?”
The new August 2026 version of the OWASP Top 10 for Large Language Model Applications is a useful way of thinking about this, but I would not treat the entries as ten tidy, isolated bugs. In the real world they overlap. That is when things get messy.
What the model can read
LLM02: Sensitive Information Disclosure
This is the obvious one, but it still gets missed. If the model has been trained on, fine-tuned with, or given access to sensitive data, assume there is a route for that data to come back out again.
A simple example: a retailer builds an internal customer service assistant and fine-tunes it on old support conversations. Those conversations contain names, addresses, phone numbers, order numbers and complaints. Nobody sanitises the data first, because the team is trying to move quickly and the data is “internal”. The model now has personal data baked into the system. At some point, directly or indirectly, that data may leak into a response.
Putting “never reveal personal data” in the system prompt is not a control I would want to bet the business on. This control relies on the agent doing what you have asked it, which is the very thing prompt injection undermines. It might help in some cases, but it is not a boundary. If the model does not need sensitive data, do not give it sensitive data. If it does need access, scope it tightly and monitor what comes back out.
LLM08: Hidden context exposure
LLM agents are given context prior to communicating with users. This will include a system prompt which tells the agent its name and what it can and can’t do, it may also include other documents specific to the particular context of the query such as the user’s details or relevant factual documents provided as part of a RAG system.
Leaking a system prompt does not automatically mean leaking a secret. However we have found it can reveal the instructions used to steer the model’s behaviour. That can help an attacker understand the application, but the security impact depends on what the context contains and which controls sit behind it.
If credentials or sensitive business logic have been put in the system prompt or operational context, they may be extracted and exposed. But that is a design failure. It is important to put all secrets and sensitive logic tucked away in proper secret storage. Enforce authorisation outside the model. Use output filtering where needed. The system prompt should describe behaviour, not carry the crown jewels.
LLM09: Vector and embedding weaknesses
Retrieval Augmented Generation (RAG) is used to increase the accuracy of LLM output by altering the user’s prompt and adding more factual information to it from a trusted data source, before then sending it to the model for processing. Typically, the model receives a selection of matching passages rather than the entire document collection, but those passages can still contain sensitive information or malicious instructions if the retrieval pipeline is not properly controlled.
If that document store contains sensitive information the user should not see, you have a disclosure problem. If the document store contains malicious instructions, you have an integrity problem. Both are easy to underestimate because the data source is often treated as trusted.
Imagine an attacker manages to get a document indexed by the RAG pipeline. The document looks like a harmless guide about bees, but hidden in the text is an instruction telling the model to lie whenever it answers questions on that topic. Now when the user asks about bees, the application retrieves the poisoned document and kindly injects the attack into its own prompt.
That is why LLM security must be about more than just prompt hardening. With RAG, it is about the pipeline. Who can add content? What gets indexed? Is content sanitised? Are permissions preserved? Can the model retrieve documents the user would not normally be allowed to read?
Search quality matters, but access control matters more. If the retrieval layer gets permissions wrong, the model becomes a very polite data leakage interface.
Keep the corpus clean, keep permissions intact, and treat externally supplied content as hostile until proven otherwise.
What the model can reach
LLM10: Improper output handling
This is where LLM security starts looking a lot like ordinary application security. If an application takes model output and passes it straight into a database, shell, browser, API call or workflow engine, the model has become part of the input chain.
For example, a travel company has a support bot that can look up booking information. An attacker asks for a booking under a name containing SQL control characters. The model dutifully passes that value into a backend lookup without filtering the output, causing SQL injection. Here the model acts as a proxy, passing malicious content onto backend systems and passing sensitive business data back to the attacker.
The fix is not “make the model smarter”. The fix is boring, proven engineering: parameterised queries, strict validation, allow-listed tool calls, encoding, escaping, and sensible trust boundaries.
Treat the model like any untrusted user. It can be helpful, but it can also be manipulated, confused, or simply wrong.
LLM03: excessive agency
Agentic AI makes demos look brilliant and threat models look worse. The more tools an AI system can use, the more carefully those tools need to be controlled.
Take an internal coding assistant. The team wants it to run code, query databases and interact with Git repositories. Someone gives it terminal access and a set of service credentials because that solves several problems at once. It can now do SQL, Git, scripts, package installs, file edits and probably a lot more.
That might be convenient, but it is also a huge blast radius. If a malicious user can steer the agent, or if the model simply makes a bad call, that unbounded agency can lead to unbounded destruction.
The answer is the same as it is for humans and service accounts: least privilege. If the model only needs to read one database table, give it read access to that table and nothing else. If it needs to execute code, do that in a sandbox. If it wants to make a material change, require approval.
Do not give an AI agent broad production access just because the demo worked nicely. Demos rarely include the attacker.
What the application trusts
LLM07: Misinformation
LLMs can be confidently wrong. That is not new, but it becomes serious when people or systems start treating the output as authoritative.
A well-known example is the US legal case where lawyers submitted court filings containing fake case references after using AI for legal research. The tool produced convincing-looking citations, the output was trusted, and the consequences were very real.
The enterprise version of this risk is everywhere. A sales assistant invents a product capability. A compliance bot gives an answer that sounds official but is out of date. A SOC assistant mislabels noise as an incident, or worse, an incident as noise.
RAG and fine-tuning can improve accuracy, but they do not remove the need for judgement. For high-impact decisions, the model should support the human process, not replace it. Make uncertainty visible, cite sources where possible, and do not let a fluent answer bypass review.
What has been introduced into the system
LLM04: Supply chain
LLM applications have supply chains too. The model, datasets, embeddings, plug-ins, orchestration framework, tools, prompts and hosting environment can all introduce risk.
With normal software, we worry about third-party packages and updates. With AI systems, we still worry about those things, but we also need to worry about the model and the data that shaped it.
If you build on someone else’s model, you inherit some of its behaviour. If you pull in third-party training data, you inherit some of its assumptions and flaws. If you connect a plug-in or tool, you inherit its attack surface.
That does not mean “never use third-party AI components”. That would be unrealistic. It means know what you are depending on, track versions, understand where data comes from, and have a plan for updates when new issues are found.
AI supply-chain review should sit alongside existing software supply-chain controls, not in a separate novelty bucket.
LLM05: Data and model poisoning
Poisoning is what happens when bad content gets into the system on purpose. That might be poisoned training data, poisoned fine-tuning data, a malicious model, or a hostile document that later gets pulled into a RAG answer.
The awkward bit is that poisoning does not have to look dramatic. It might just be a document in a shared folder, a support article in a knowledge base, or a dataset that nobody has properly reviewed. Once it is trusted by the pipeline, the model may repeat or act on it.
Good controls here are mostly about provenance and review. Know where training and retrieval content came from. Restrict who can add to it. Monitor for odd outputs. Re-index carefully. Do not let “it came from our documents” become a blanket trust decision.
The important point is this: once poisoned content becomes part of the model’s world, it can influence answers long after the original input has been forgotten.
How the system can be abused at scale
LLM06: Unbounded consumption
Finally, there is the simple problem of scale. LLMs can be expensive to run, slow to respond under load, and vulnerable to abuse if there are no sensible limits.
An attacker can send very long prompts, trigger costly operations, automate repeated requests, or force the system into expensive tool use. If the service is billed per token or per API call, that quickly becomes a denial-of-wallet problem as well as a performance problem.
There is also a confidentiality angle. If an attacker can query a model enough, they may be able to infer behaviour, extract useful outputs, or train a copycat system that approximates what the original does.
Rate limits, token limits, timeouts, monitoring and cost alerts are not exciting, but they are essential. If you expose an LLM-backed feature to users, assume someone will eventually try to make it do too much, too often, for too long.
So what should we actually do?
Start by accepting that an LLM is not a trusted security boundary. It is a component in a wider application. Secure the application.
- Limit what the model can read.
- Limit what the model can do.
- Validate anything the model outputs before another system acts on it.
- Preserve user permissions through RAG and tool calls.
- Keep secrets out of prompts.
- Monitor use, cost, failures and weird outputs.
- Put human approval in front of high-impact actions.
Prompt injection absolutely needs testing. It is one of the clearest ways an attacker can try to steer an AI system away from what its designers intended. But it should not be tested in isolation. The real question is what happens when prompt injection succeeds. Can the model reach sensitive data? Can it call privileged tools? Can it make changes? Can it mislead a user into doing something risky?
That is where AI security becomes more interesting. Prompt injection can be the entry point, but the impact comes from the data, tools and trust sitting behind the model.