Skip to content

AI & LLM Security

LLM security is a distinct attack surface: the model is a component in a larger system, and the attacks exploit the gap between what the developer intended the model to do and what an adversary can cause it to do by controlling its inputs.

Start with the three hub docs to orient, then work through the OWASP LLM Top 10 in order, then the agent-specific attack classes, then the protocol deep dives.

Hubs

Doc Purpose
Web LLM Attacks Entry point — maps the attack surface of LLMs integrated into web apps
MCP Protocol Security Model Context Protocol threat model
Agentic AI Threats Threat taxonomy for autonomous agent systems

OWASP LLM Top 10

# Doc
LLM01a Direct Prompt Injection
LLM01b Indirect Prompt Injection
LLM02 Sensitive Information Disclosure
LLM03 LLM Supply Chain
LLM04 Data & Model Poisoning
LLM05 Improper Output Handling
LLM06 Excessive Agency
LLM07 System Prompt Leakage
LLM08 Vector & Embedding Weaknesses
LLM09 Misinformation & Hallucination
LLM10 Unbounded Consumption

Agent-specific attacks

Doc Attack class
Memory Poisoning Persistent injection via long-term memory
Plan & Goal Hijacking Loop-level objective rewrite
Cascading Hallucination Cross-agent privilege laundering
HITL Bypass Approval fatigue, spoofed UI
Cross-Agent Trust & A2A Injection Unauthenticated semantic content
Tool-Schema Confusion Typed-argument violations
Credential Passthrough Over-broad scopes, RFC 8707 violations
Sandbox Escape via Composition Composition-level escape
MCP Cross-Server Shadowing Tool-description hijack
Rug Pull & Tool-Definition Drift Metadata-plane supply chain
Orchestrator Prompt Injection Unescaped template variables

Protocols & architecture

MCP Deep Dive · A2A Protocol · Function-Calling Protocols · RAG Architecture · Vector Stores · Model Serving · Guardrail Systems · Model File Formats

Defenses

AI & Agent Defenses · Spotlighting