AI & LLM Security¶
LLM security is a distinct attack surface: the model is a component in a larger system, and the attacks exploit the gap between what the developer intended the model to do and what an adversary can cause it to do by controlling its inputs.
Start with the three hub docs to orient, then work through the OWASP LLM Top 10 in order, then the agent-specific attack classes, then the protocol deep dives.
Hubs¶
| Doc | Purpose |
|---|---|
| Web LLM Attacks | Entry point — maps the attack surface of LLMs integrated into web apps |
| MCP Protocol Security | Model Context Protocol threat model |
| Agentic AI Threats | Threat taxonomy for autonomous agent systems |
OWASP LLM Top 10¶
| # | Doc |
|---|---|
| LLM01a | Direct Prompt Injection |
| LLM01b | Indirect Prompt Injection |
| LLM02 | Sensitive Information Disclosure |
| LLM03 | LLM Supply Chain |
| LLM04 | Data & Model Poisoning |
| LLM05 | Improper Output Handling |
| LLM06 | Excessive Agency |
| LLM07 | System Prompt Leakage |
| LLM08 | Vector & Embedding Weaknesses |
| LLM09 | Misinformation & Hallucination |
| LLM10 | Unbounded Consumption |
Agent-specific attacks¶
| Doc | Attack class |
|---|---|
| Memory Poisoning | Persistent injection via long-term memory |
| Plan & Goal Hijacking | Loop-level objective rewrite |
| Cascading Hallucination | Cross-agent privilege laundering |
| HITL Bypass | Approval fatigue, spoofed UI |
| Cross-Agent Trust & A2A Injection | Unauthenticated semantic content |
| Tool-Schema Confusion | Typed-argument violations |
| Credential Passthrough | Over-broad scopes, RFC 8707 violations |
| Sandbox Escape via Composition | Composition-level escape |
| MCP Cross-Server Shadowing | Tool-description hijack |
| Rug Pull & Tool-Definition Drift | Metadata-plane supply chain |
| Orchestrator Prompt Injection | Unescaped template variables |
Protocols & architecture¶
MCP Deep Dive · A2A Protocol · Function-Calling Protocols · RAG Architecture · Vector Stores · Model Serving · Guardrail Systems · Model File Formats