Key Takeaways
- Eleven attack patterns account for the overwhelming majority of real LLM incidents in production, and every one of them happens at runtime—after your build-time controls have already passed.
- The highest-frequency vector is not a user typing something clever. It is indirect prompt injection: instructions hidden in content your AI retrieves.
- Each attack has a specific, observable runtime signal. Detection is a telemetry problem before it is a modeling problem.
- Agent tool calls are the highest-severity surface, because that is where a language problem becomes a systems problem.
- If you log only prompts and responses, you cannot detect seven of the eleven. Tool-call telemetry is the gap most programs have.
Most enterprise AI security programs can tell you which models are approved. Very few can tell you what happened during the 4.2 million inference calls their organization made last month. That asymmetry is the whole problem—attacks against AI systems do not arrive as vulnerabilities in an artifact. They arrive as content, in traffic, at runtime.
Below are the eleven patterns we see most often against enterprise LLM and agent deployments, what each looks like in practice, the runtime signal that reveals it, and the control that stops it. They map closely to the OWASP Top 10 for LLM Applications, extended for agent and MCP behavior.
1. Direct Prompt Injection
What it is: A user supplies input designed to override the system prompt—role reassignment, instruction negation, encoded payloads, or a long context designed to push guardrail instructions out of attention.
Runtime signal: Input containing imperative instruction patterns aimed at the model rather than the task; sudden divergence between the request’s stated intent and the tools the model then attempts to use.
Control: Input classification before inference, plus action authorization behind it—so that even a successful injection cannot reach a privileged tool.
2. Indirect Prompt Injection
What it is: The instruction is not typed by the user. It is planted in a document, web page, email, ticket, or code comment that your AI retrieves and treats as context. This is the most common real-world vector and the one teams most often fail to instrument.
Runtime signal: Instruction-shaped text appearing in retrieved content rather than user input; a tool call whose parameters trace to retrieved text rather than to the user’s request.
Control: Classify retrieved context with the same rigor as user input. Maintain provenance so every tool argument can be traced to its source.
3. Tool Poisoning
What it is: A malicious or compromised tool—often an MCP server—carries a description or schema containing instructions to the model. The agent reads the tool definition and behaves accordingly, before any user has done anything.
Runtime signal: Newly registered tools appearing in call chains; tool descriptions containing imperative language; behavior changing after a tool catalog refresh with no code deploy.
Control: Inventory and pin every registered tool. Validate schemas. Alert on any change to a tool definition. The NSA’s 2026 guidance on Model Context Protocol security calls out exactly this class of risk.
4. Excessive Agency Exploitation
What it is: The agent has more capability than its task requires—write access when it needs read, a broad API scope when it needs one endpoint—and an attacker steers it into the surplus.
Runtime signal: Tool invocations outside the agent’s historical pattern; use of a permission that has never been exercised in production; parameter values touching data classifications the workflow does not need.
Control: Per-agent action policy evaluated at call time, not just at credential-issuance time. Deny by default and expand from observed need.
5. System Prompt Leakage
What it is: The model discloses its own configuration—instructions, guardrail logic, tool inventory, sometimes embedded credentials. Attackers use it as reconnaissance for a more precise second attempt.
Runtime signal: Outputs containing substrings matching the system prompt; a burst of meta-questions about the assistant’s own rules from one principal.
Control: Output inspection against known configuration text, and the structural fix—never place secrets in a system prompt.
6. Sensitive Information Disclosure
What it is: The model returns regulated or confidential data the requester is not entitled to—pulled from context, retrieval, or memorized training data.
Runtime signal: Outputs matching sensitive data classifiers; entitlement mismatch between the requesting principal and the classification of retrieved documents.
Control: Output filtering plus retrieval-time entitlement enforcement. The correct fix is upstream: the model should never see what the user cannot.
7. RAG Index Poisoning
What it is: An attacker plants content in a source your pipeline ingests—a wiki page, a shared drive, a public site—so that it is retrieved later and shapes answers or triggers actions.
Runtime signal: New or recently modified documents appearing disproportionately in retrieval results; answers citing sources with no legitimate authorship trail.
Control: Provenance and integrity checks on ingestion, plus retrieval monitoring that treats a sudden shift in source distribution as an event.
8. Vector and Embedding Attacks
What it is: Manipulation of the embedding layer—adversarial content crafted to dominate similarity search, or inversion attacks that reconstruct source text from stored vectors.
Runtime signal: Queries producing anomalously high similarity scores against a narrow document set; bulk or programmatic access patterns against the vector store.
Control: Access controls and rate limits on the vector store itself, and treatment of embeddings as sensitive data—because they are.
9. Improper Output Handling
What it is: Model output flows unsanitized into something that executes it—a browser, a shell, a SQL query, another agent. The classic web vulnerabilities return, with the model as the injection point.
Runtime signal: Outputs containing markup, SQL fragments, shell syntax, or URLs bound for internal addresses, in responses destined for an executing consumer.
Control: Treat every model output as untrusted input downstream. Encode, validate, and constrain at the consumer.
10. Unbounded Consumption
What it is: Resource exhaustion—prompt storms, recursive agent loops, deliberately expensive queries. Denial of service and denial of wallet at once.
Runtime signal: Call volume or token spend far outside a principal’s baseline; agent loops exceeding expected recursion depth; cost per session with a long tail.
Control: Per-principal rate, token, and cost budgets enforced at the gateway, plus hard recursion limits on agent chains.
11. Agent Identity Abuse and the Confused Deputy
What it is: An agent holds a broad service credential and acts on behalf of many users. An attacker gets it to perform a privileged action for the wrong principal. Tokens are replayed, sessions are reused, and the audit log shows only the agent.
Runtime signal: Actions attributed to a service identity with no traceable human principal; the same token appearing across sessions or source contexts; entitlement checks that pass for the agent but would fail for the requester.
Control: Propagate end-user identity through the agent call chain and authorize on the effective principal. Short-lived, narrowly scoped credentials per task.
The Telemetry Gap Underneath All Eleven
Read back through the detection signals and a pattern appears: four of the eleven are visible in prompt-and-response logs. The other seven require tool-call telemetry—invocation, arguments, source provenance, and the identity behind it.
Most enterprise AI logging today captures the first and not the second. That is why so many organizations can describe their AI policy in detail and cannot answer a basic question about what their agents did last Tuesday.
Where Cranium Fits
Detection is only useful inside a loop that also finds the systems, sets the policy, tests the model, and records the outcome. The AI Trust Loop runs Discover, Observe, Govern, Secure, and Prove:
- Discover CodeSensor™ and AgentSensor™ map the models, agent tools and MCP servers present in your code, so detection coverage has a real denominator.
- Observe Guardian captures interaction-level telemetry between your AI application and the LLM: full prompt and response, signal results on both, and a Tool Inspector that records which tools an agent called during the event.
- Govern Profiles decide what happens on each signal — block, modify or pass — on input and on output, with Listen mode available while you tune against real traffic.
- Secure Guardian’s documented security signals map directly onto several patterns above: jailbreaking, instruction override, prompt leaking, role impersonation, direct command injection, self-referential injection and goal hijacking. Arena™ additionally probes for prompt injection, data leakage and jailbreak susceptibility before release.
- Prove Events, signals and actions are retained as the record an investigator or examiner will ask for, and can be streamed to your SIEM.
You already know how to run detection engineering. What you have been missing is a signal source for the AI layer.
Frequently Asked Questions
Which of these should we address first?
Indirect prompt injection and excessive agency. The first is the most frequent; the second turns any successful injection into a real-world action.
Can a WAF or API gateway detect these?
No. Gateways reason about protocol and volume. Every attack here is semantically valid, well-formed traffic—the payload is meaning, not malformation.
Does using a major commercial model make us safe from these?
Frontier labs harden against several patterns, and that helps. But excessive agency, tool poisoning, output handling, and identity abuse are properties of your architecture, not of the model you called.
How do we test our detection coverage?
Red team against this list specifically, with tool-call telemetry enabled, and measure how many you detected rather than how many succeeded. Those are different numbers, and the gap between them is your real posture.
See Cranium in Action
See the signals Cranium detects in live AI traffic—schedule a personalized demo: cranium.ai/get-a-demo/
