Skip to content

Kubernetes Security

Kubernetes is a distributed control plane that reconciles declared state (YAML manifests) into running containers across a fleet of nodes. From an appsec view it stacks three trust boundaries. The API server mediates all state changes, the kubelet on each node turns Pod specs into running containers, and the pod-to-pod network is flat and permissive by default. Every attack path either abuses an over-broad RBAC grant to talk to the API server, steals a service account token to impersonate a workload, or exploits a Pod spec field (hostPath, privileged, hostNetwork) that the admission layer failed to block. The 4 Cs model (Cloud, Cluster, Container, Code) exists because a hardened pod inside a soft cluster inside an over-permissive cloud IAM role still yields cluster takeover; the layers compose multiplicatively. Pod Security Policy is gone (removed in v1.25); Pod Security Admission is built in but coarse, so most production clusters run Kyverno or Gatekeeper for real policy. The interview bar is not "name the components" but "trace a stolen pod token through RBAC, kubelet, and network policy to explain why it did or did not become cluster-admin."

Interview frequency: Situational

See also: Authorization for RBAC and policy-engine design behind API server and admission-control authorization decisions, Secrets, Keys, and Data Protection for how Kubernetes Secrets and service account tokens should be custodied and rotated, and Multi-Tenancy and Isolation for the namespace, node, and cluster isolation layers that bound a compromised pod.

How it works

Control plane and node components

The API server (kube-apiserver) is the only component that talks to etcd, and the only front door for kubectl, controllers, and kubelets: every request passes authentication, authorization (RBAC by default), and admission control (mutating then validating) before the object is persisted. etcd holds all cluster state including secrets, so an etcd leak is a cluster leak. The controller manager runs reconciliation loops (Deployment, ReplicaSet, ServiceAccount token controller); the scheduler picks nodes based on resource requests, taints, and affinities.

On each node, the kubelet watches the API server for Pods assigned to it and instructs the container runtime (containerd or CRI-O; Docker shim was removed in v1.24) via CRI to pull images and start containers. kube-proxy programs iptables or IPVS for Service load balancing. The CNI plugin (Calico, Cilium, AWS VPC CNI) provisions pod networking and, if it supports it, enforces NetworkPolicy.

flowchart TB
  subgraph CP["Control plane"]
    API["kube-apiserver"]
    Etcd["etcd"]
    CM["Controller manager"]
    Sched["Scheduler"]
  end
  subgraph WN["Worker node"]
    Kubelet["kubelet"]
    Proxy["kube-proxy"]
    CNI["CNI plugin"]
    CRI["Container runtime<br/>(containerd / CRI-O)"]
  end
  Client["kubectl / controllers"] --> API
  API --> Etcd
  API --> CM
  API --> Sched
  Kubelet <--> API
  Kubelet --> CRI

The sequence below shows the same components in motion: a Pod create request flowing through the API server's authn/authz/admission chain to etcd, then the kubelet on the target node picking it up.

sequenceDiagram
  participant U as kubectl user
  participant API as kube-apiserver
  participant Adm as Admission chain
  participant Etcd as etcd
  participant Kubelet as kubelet on node
  participant CRI as containerd
  U->>API: POST /pods (Pod spec)
  API->>API: Authn (cert / OIDC / SA token)
  API->>API: Authz (RBAC / Node / ABAC)
  API->>Adm: Mutating webhooks (Kyverno, sidecar injector)
  API->>Adm: Validating webhooks (PSA, Gatekeeper)
  API->>Etcd: persist Pod object
  API-->>U: 201 Created
  Kubelet->>API: watch /pods?fieldSelector=nodeName=n1
  Kubelet->>CRI: RunPodSandbox, PullImage, CreateContainer
  CRI-->>Kubelet: containerID
  Kubelet->>API: PATCH /pods/status

Authentication modes

The API server has no user database. It accepts identities from multiple authenticators tried in order:

  • X.509 client certificates signed by the cluster CA. CN becomes the username, O becomes the group. Certificates cannot be revoked without rotating the CA, so cert-based user auth is discouraged for humans.
  • Service account tokens: JWTs signed by the API server, mounted into pods at /var/run/secrets/kubernetes.io/serviceaccount/token. Since v1.22 these are projected, short-lived, and audience-bound (the "BoundServiceAccountTokenVolume" feature is GA and default).
  • OIDC: the API server verifies an ID token from an external IdP (Okta, Google, Dex). Username and groups come from configured claims.
  • Webhook token auth: the API server POSTs the bearer token to an external service for verification. Used by managed clusters (EKS uses aws-iam-authenticator or the newer aws-auth mapping).
  • Bootstrap tokens for kubeadm node join, and static token files which are legacy and should not exist in production.

RBAC model

Four kinds: Role/RoleBinding are namespace-scoped, ClusterRole/ClusterRoleBinding are cluster-scoped. A ClusterRole granted via a namespaced RoleBinding is the standard way to give a service account read access to only its own namespace. Verbs are the usual get, list, watch, create, update, patch, delete, deletecollection plus the escalation-relevant ones: impersonate, bind, and escalate. escalate on roles/clusterroles lets a subject grant privileges they don't themselves have, turning a "role editor" into cluster-admin; bind does the same by letting a subject attach an existing high-privilege ClusterRole to themselves. Default system:* roles ship baked into the API server and can't be safely removed.

Pod Security Admission

Since v1.25 the API server ships with a built-in admission plugin implementing three Pod Security Standards[1]:

  • Privileged: no restrictions. Suitable for infrastructure workloads (CNI agents, storage drivers) that need host access.
  • Baseline: blocks the obviously dangerous fields (privileged: true, hostNetwork, hostPID, hostIPC, most hostPath volumes, adding capabilities beyond a small allow-list).
  • Restricted: baseline plus non-root, seccomp RuntimeDefault, drop ALL capabilities, no privilege escalation, read-only root filesystem is recommended but not required.

Each namespace can carry three labels: pod-security.kubernetes.io/enforce=<level> (rejects violating pods), .../audit=<level> (logs), .../warn=<level> (kubectl warning). The enforce policy applies only at pod creation; existing pods are not re-evaluated when the label changes. PSA does not cover CRD workloads, image provenance, or network policy; for those you need Kyverno[2], Gatekeeper[3], or Kubewarden.

NetworkPolicy

NetworkPolicy is a namespaced resource with pod selectors and ingress/egress rules. The default posture in a bare cluster is "any pod can reach any pod or external endpoint," including cross-namespace; a policy that selects a pod switches it to default-deny for whichever direction (ingress/egress) it has rules for, so true namespace-wide default-deny needs an empty-selector policy denying everything, then explicit allows layered on top. Enforcement depends on the CNI: flannel does not enforce policies at all, Calico and Cilium do — a common footgun, since the manifest applies successfully but has zero effect on a no-op CNI.

Secrets and etcd

Secret objects are stored in etcd as base64-encoded plaintext by default; encryption at rest needs EncryptionConfiguration on the API server with a KMS provider (AWS KMS, GCP KMS, HashiCorp Vault — kms v2, GA in v1.29, is the recommended one). Without it, anyone who reads etcd — backup files, disk snapshots, an etcd-role compromise — reads every secret in the cluster. Secret volumes mount into pods as tmpfs, so they never touch node disk, but the API server still serves them to any subject with get secrets in the namespace.

Workload identity

Long-lived service account tokens are a footgun; the modern pattern is workload identity federation — the pod's projected service account token is exchanged for a cloud IAM credential via OIDC federation (IRSA/Pod Identity on EKS, Workload Identity on GKE, Azure AD Workload Identity on AKS). The API server publishes an OIDC discovery document, the cloud IAM provider trusts that issuer, and pods get short-lived credentials tied to system:serviceaccount:<ns>:<name>. See 81-spiffe-spire.md for the generalized SPIFFE model this is a special case of.

Quick reference

# Wire: attacker compromises app pod, dumps projected SA token, hits API server
$ TOKEN=$(cat /var/run/secrets/kubernetes.io/serviceaccount/token)
$ CACERT=/var/run/secrets/kubernetes.io/serviceaccount/ca.crt
$ NS=$(cat /var/run/secrets/kubernetes.io/serviceaccount/namespace)
$ curl -sk --cacert $CACERT -H "Authorization: Bearer $TOKEN" \
    https://kubernetes.default.svc/api/v1/namespaces/$NS/secrets
# If this returns 200, the pod's SA has `get secrets` in its namespace.
# If it also works against /api/v1/secrets (cluster-wide), the SA is
# bound cluster-wide and every secret in the cluster is now readable.
Invariant Where enforced How violated Source
Every API request authenticated and authorized kube-apiserver authn + authz chain Anonymous auth left enabled (--anonymous-auth=true) with system:unauthenticated bound to a role [4]
Pods cannot mount host root or run privileged unless namespace opts in Pod Security Admission at enforce=restricted Namespace labeled enforce=privileged or PSA not labeled at all (defaults to no enforcement) [1]
Service account tokens are short-lived and audience-bound BoundServiceAccountTokenVolume (GA v1.22) automountServiceAccountToken: true on default SA plus legacy long-lived Secret-type token [5]
etcd contents encrypted at rest EncryptionConfiguration with KMS provider on API server Default install stores secrets as base64 in etcd; snapshot leak = full secret leak [6]
Pod-to-pod traffic default-deny NetworkPolicy + CNI that enforces it (Calico, Cilium) No NetworkPolicy present or CNI (flannel) does not enforce [7]
Kubelet API requires client cert authn and RBAC authz --authentication-token-webhook=true, --authorization-mode=Webhook on kubelet Kubelet on 10250 with --anonymous-auth=true and --authorization-mode=AlwaysAllow [4]
Admission webhooks fail closed for security policies failurePolicy: Fail on ValidatingWebhookConfiguration failurePolicy: Ignore means an attacker who DoSes the webhook bypasses policy [8]

Attack techniques

1. RBAC over-permissioning and cluster-admin sprawl

The most common cluster compromise is a service account granted cluster-admin — a Helm chart's default values shipped rbac.create: true and serviceAccount.clusterAdmin: true, or an operator installer bound its controller SA to cluster-admin "temporarily" during setup and no one ever narrowed it. Once an attacker gets code execution in any pod holding that SA, the projected token is a bearer credential to the API server: curl with it creates a privileged pod on any node, mounts the host filesystem, and reads every node's credentials.

The escalation surface is wider than "get secrets": impersonate, escalate, and bind are transitive privilege operators. impersonate users on the empty resource lets a subject send requests as any user, including system:admin. escalate roles on rbac.authorization.k8s.io lets a subject create a Role or ClusterRole with verbs they don't themselves have; bind rolebindings lets a subject attach an existing high-privilege ClusterRole to themselves. Kubernetes-goat and the "path to cluster-admin" community write-ups enumerate dozens of these one-step-away misconfigurations[9].

Black-box confirmation runs from inside a pod: kubectl auth can-i --list prints every verb/resource pair the current subject can perform, and kubectl auth can-i create pods --as=system:serviceaccount:default:default (from an admin context) tests specific paths. For a compromised pod, curl the API server with the mounted token against /apis/authorization.k8s.io/v1/selfsubjectrulesreviews to get the same output without kubectl installed.

2. Service account token theft from pod filesystem

Any code executing inside a pod can read /var/run/secrets/kubernetes.io/serviceaccount/token — the entire authentication story for in-cluster workloads, with no additional secret behind it. An SSRF or RCE that reads files on the pod filesystem is therefore an API server credential leak. Modern projected tokens are short-lived (~1 hour, refreshed by kubelet) and audience-bound to the API server, which narrows the window but doesn't close it: exfiltrating the token buys ~1 hour of API access, or persistent access if the attacker can re-read the file on each rotation.

automountServiceAccountToken: false on the pod spec is the mitigation most teams reach for, but it only helps pods that don't need the API server. Pods that do need access should bind a specific narrow ServiceAccount instead of default, enforced via admission policy. The pre-v1.22 pattern of long-lived, non-audience-bound tokens (kubernetes.io/service-account-token Secrets) never expires if any linger — a stolen legacy token is a permanent credential.

3. Dangerous Pod spec fields (privileged, hostPath, hostNetwork, hostPID)

Kubernetes exposes the Linux security model as Pod spec fields, and each dangerous one is a direct path to node compromise:

  • privileged: true grants all capabilities and disables seccomp/AppArmor — the container is functionally root on the node.
  • hostPath: / mounts the node root filesystem into the pod, so an attacker copies /etc/kubernetes/pki/* (control-plane certs on master nodes) or /var/lib/kubelet/pods/*/volumes/kubernetes.io~secret/* (every other pod's secrets on that node).
  • hostNetwork: true puts the pod on the node's network namespace, reaching localhost services including the kubelet's read-only port and, on some clouds, the metadata service — bypassing the pod network's egress restrictions.
  • hostPID: true lets the pod see and signal every process on the node, including the kubelet.

An attacker with create pods in any namespace can craft a pod with these fields and schedule it onto a target node using a nodeSelector or nodeName; from that pod, escaping to the node is one nsenter away (see 86-container-escape.md). This is the reason PSA restricted exists and the reason "just give the CI pipeline create pods" is a cluster-admin grant in disguise.

4. kubelet API and etcd exposure

Every node runs a kubelet with a TLS API on port 10250; configured with --anonymous-auth=true and --authorization-mode=AlwaysAllow (the historic default, still seen on hand-rolled clusters), any network peer can hit /pods to list pods, /exec/<ns>/<pod>/<container> to run a command, and /run/<ns>/<pod>/<container> to open a shell — a direct-to-shell primitive that bypasses the API server, RBAC, and audit logging entirely.

etcd listens on 2379 (client) and 2380 (peer); if the client port is reachable without mutual TLS (--client-cert-auth=false), an attacker with network access reads every object including secrets. Managed clusters (EKS, GKE, AKS) fence these ports off from the pod network by default — hand-rolled kubeadm clusters historically did not. The 2018 Tesla incident[10] and the recurring "exposed Kubernetes dashboard" incidents (dashboard v1 defaulted to skip login on some builds) are variants of the same "control plane on the internet" pattern.

5. Admission webhook bypass via fail-open

Kyverno and Gatekeeper install as ValidatingWebhookConfigurations, and each webhook's failurePolicy controls behavior when it's unreachable: Ignore admits the request (fail-open), Fail rejects it (fail-closed). Ignore on a policy webhook is a common misconfiguration, often chosen out of fear the webhook could take the cluster down — and it's a bypass primitive: an attacker with delete pods in the policy namespace kills the webhook pod or DoSes the service, then submits the previously-blocked manifest during the outage. Anyone who can update validatingwebhookconfigurations can flip failurePolicy or narrow the namespaceSelector themselves[8].

Confirm with kubectl get validatingwebhookconfigurations -o yaml | grep -E 'failurePolicy|namespaceSelector', which shows the enforcement posture. Bind this to admission-time policy: Kyverno itself can enforce that new webhooks must be failurePolicy: Fail.

6. Supply-chain: mutating webhooks and image tag mutability

Mutating admission webhooks run before validating ones and can rewrite any field of any object — the Istio and Linkerd sidecar injectors work this way, as do many "policy as code" tools. An attacker who installs one (create mutatingwebhookconfigurations is a cluster-admin verb by default, but sometimes granted to platform operators) can silently inject an extra container into every pod created cluster-wide: a sidecar that reads the shared volume, exfiltrates the service account token, or opens a reverse shell. This is Kubernetes-native persistence — a red-team implant that survives node reboots and namespace deletions.

Image tag mutability is the other supply-chain surface: image: myapp:latest resolves to whatever the registry currently points latest at, so an attacker who compromises the registry pushes a new latest and every restarting pod pulls the malicious image. Even image: myapp:v1.2.3 is mutable at the tag level unless the registry enforces immutability — only a digest pin (myapp@sha256:...) is cryptographic. Cosign/sigstore adds a signature layer verified at admission time via Kyverno's verifyImages or Sigstore's policy-controller[11].

7. Notable CVEs and CVE classes

CVE-2018-1002105 (API server proxy)[12]: the API server proxied user connections to backend services (aggregated API servers, kubelets, exec/attach). The proxy kept the backend connection open after the initial request completed and allowed the client to send arbitrary follow-up requests as the API server's own identity. Any user with exec permission on any pod became cluster-admin. Fixed by closing the proxied connection after the initial response. This is the archetypal "proxy state confusion" bug in Kubernetes and shows up in the CVSS 9.8 hall of fame.

CVE-2020-8558 (kube-proxy iptables masquerade)[13]: kube-proxy on some versions inserted an iptables rule that let containers reach services bound to the node's loopback interface (127.0.0.1) via the node IP. Local-only services (etcd on master nodes, cloud metadata proxies) that assumed 127.0.0.1 was safe were exposed to any pod. Fixed by adding a rule to drop -d 127.0.0.0/8 on non-loopback interfaces.

CVE-2022-0811 (cr8escape) in CRI-O[14]: the container runtime passed unsanitized sysctl values from Pod spec to the kernel. kernel.core_pattern=|/proc/self/exe ... set the core-dump handler to an attacker-controlled binary running on the host. Any subject with create pods and a namespace not enforced by PSA restricted (which blocks unsafe sysctls) could take the node.

CVE-2025-1974 (IngressNightmare) in ingress-nginx[15]: a chain of bugs in the ingress-nginx admission webhook and template rendering allowed unauthenticated attackers who could reach the admission webhook (which was cluster-internal by default but sometimes exposed) to inject Nginx config directives, achieving RCE in the ingress controller pod. The ingress controller typically has a service account with get secrets cluster-wide to read TLS material, so RCE in ingress-nginx frequently escalates to reading every Secret in the cluster.

Defense

Real fix

Enforce PSA restricted at every application namespace and PSA baseline at every infrastructure namespace. Set the label at namespace creation time via admission policy so a new namespace cannot exist without it. Reserve privileged for a short allow-list of infrastructure namespaces (CNI, CSI, monitoring). Verify with kubectl get ns -L pod-security.kubernetes.io/enforce[1].

RBAC least-privilege with no cluster-admin bindings outside break-glass. Every ServiceAccount gets a namespace-scoped Role bound via RoleBinding; workloads that need cross-namespace read get a ClusterRole granted via RoleBinding in each target namespace, never a ClusterRoleBinding. Audit periodically with kubectl get clusterrolebindings -o json | jq '.items[] | select(.roleRef.name=="cluster-admin")', which should return the empty set except for the system:masters group binding. Deny impersonate, escalate, bind, and create clusterrolebindings verbs to anyone outside the platform team.

NetworkPolicy default-deny on every namespace, on a CNI that enforces it. Ship a default policy as part of namespace bootstrap: podSelector: {} with empty ingress and egress arrays, then layer explicit allows. Confirm the CNI enforces (kubectl exec into a pod and curl a pod that should be blocked; expect timeout, not RST)[7].

Encrypt etcd at rest with a KMS provider (kms v2). The EncryptionConfiguration should list kms first for secrets and configmaps (and any custom resources holding sensitive data), with identity as the fallback for decryption during rotation. Rotate the KMS key on a schedule and re-encrypt existing objects with kubectl get secrets -A -o json | kubectl replace -f -[6].

Workload identity federation replaces long-lived tokens. On EKS use IRSA or Pod Identity; on GKE use Workload Identity; on AKS use Azure AD Workload Identity. Pods exchange the projected SA token for a cloud IAM credential scoped to system:serviceaccount:<ns>:<name>, so the SA name is the cloud IAM principal. Delete legacy kubernetes.io/service-account-token Secrets, they are non-expiring[5].

Admission policy blocks known-dangerous fields. Kyverno[2] or Gatekeeper[3] policies that deny privileged, hostPath outside a small allow-list, hostNetwork, hostPID, hostIPC, adding capabilities beyond a small set, mutable image tags, and pods without a securityContext.runAsNonRoot: true. failurePolicy: Fail on all such webhooks with liveness probes so the API server never talks to a dead webhook.

Image signing verified at admission. Cosign-signed images with a policy-controller (Sigstore's policy-controller or Kyverno's verifyImages) that rejects unsigned or wrong-signer images at pod creation. Combined with digest pinning this closes the mutable-tag path[11].

Defense in depth

Bind the API server to a private cluster endpoint. Use a VPC/VNet endpoint with no public IP, accessible only via bastion, VPN, or cloud-private connectivity. Managed clusters expose this as endpointPrivateAccess: true; disable public access unless you have a specific reason.

Run --authorization-mode=Node,RBAC. The Node authorizer restricts what a kubelet can do to only pods bound to its own node. Combined with the NodeRestriction admission plugin, this prevents a compromised kubelet from reading secrets for pods on other nodes.

Scan against the CIS Kubernetes Benchmark. kube-bench (or the vendor equivalent) runs the CIS controls against control-plane and worker configs. The interview-relevant controls are the API server flags (--anonymous-auth=false, --authorization-mode=Node,RBAC, --audit-log-path set), kubelet flags, and etcd flags. Alert on drift.

Runtime detection with Falco or Tetragon. Rules for "shell in container", "cat /var/run/secrets", "outbound to Kubernetes API from unexpected pod", "process wrote to /etc/kubernetes on a node", "sensitive syscall from container". Runtime detection catches the post-exploitation half that admission policy cannot see.

Ship audit logs to an external sink with alerting. --audit-log-path and an audit policy that captures RequestResponse for secrets, serviceaccounts/token, rolebindings, and mutatingwebhookconfigurations. Alert on create clusterrolebindings where roleRef.name=cluster-admin.

Turn automount off by default. Set automountServiceAccountToken: false at the ServiceAccount level for default in every namespace; pods that need API access opt in with a specific ServiceAccount.

Read-only root filesystem and drop ALL capabilities on application pods (securityContext.readOnlyRootFilesystem: true, capabilities.drop: [ALL]). It removes the ergonomic tooling an attacker relies on post-RCE, though it is not a fix on its own.

Detection and telemetry

Key audit-log signals from the API server:

  • verb=create resource=pods where the resulting pod has spec.hostNetwork=true, spec.hostPID=true, or a spec.containers[].securityContext.privileged=true. Log the user and namespace; alert on any non-infrastructure namespace.
  • verb=create resource=clusterrolebindings where roleRef.name is cluster-admin, admin, or contains edit. Alert always.
  • verb=create resource=serviceaccounts/token at high volume from one user or one IP. Legitimate workloads do this on rotation, but a scan-and-list from a compromised subject spikes this signal.
  • verb=exec resource=pods/exec from a non-human user. Automation should not exec into pods; if it does, log and require a change ticket ID in an annotation.
  • verb=update resource=validatingwebhookconfigurations or mutatingwebhookconfigurations. Alert on any change to failurePolicy or namespaceSelector.
  • Anonymous requests to any resource (user.username=system:anonymous) that return anything other than 401/403. If this fires at all outside the health endpoints, --anonymous-auth=true is enabled and needs to be turned off.

Node-level signals (Falco/Tetragon):

  • Process nsenter, unshare, or docker executed inside a container.
  • /proc/self/root/.. or /proc/1/root/.. path traversal from a container.
  • Container process opening /var/run/docker.sock, /var/run/containerd/containerd.sock, or /var/lib/kubelet/pods/.
  • Shell (bash, sh, zsh) exec'd inside a container with an image label indicating a scratch or distroless base (should not have a shell to begin with).

Canary shape: a dedicated namespace with a "honey secret" (Secret object with a plausible-looking database password) and an audit rule that alerts on any get secret/<name> from any subject. Read attempts indicate an attacker enumerating secrets cluster-wide.

Interviewer probes

Q: A pod is compromised via RCE in the app. Walk me through the blast radius. Mid: The attacker reads the projected service account token from /var/run/secrets/kubernetes.io/serviceaccount/token, hits the API server, and can do whatever that ServiceAccount is authorized for. Principal: Blast radius decomposes into identity and reachability. For identity, run curl against /apis/authorization.k8s.io/v1/selfsubjectrulesreviews to enumerate the SA's RBAC. If the SA has create pods in any namespace, that is cluster-admin-equivalent (schedule a privileged pod, mount hostPath, escape). If it has only get secrets in one namespace, blast is contained to that namespace. For reachability, default-deny NetworkPolicy prevents the pod from reaching other pods or the cloud metadata service (169.254.169.254), so cloud IAM escalation via IMDSv1 is blocked; and if the pod is not privileged and PSA restricted is enforced, the pod cannot escape to the node via well-known primitives, so the attacker is stuck at pod scope until a container-runtime CVE lands. The mitigation triage is: rotate the SA token, apply a NetworkPolicy quarantining the pod's labels, snapshot the pod for forensics, then delete.

Q: Why is automountServiceAccountToken: true on the default ServiceAccount a problem? Mid: Because every pod that does not specify a ServiceAccount gets a token mounted, even if it does not need to talk to the API server; RCE in that pod leaks a credential. Principal: The larger problem is the coupling: developers rarely audit ServiceAccount bindings on default, so a well-meaning platform engineer who binds a ClusterRole to system:serviceaccounts:my-ns (all SAs in the namespace) grants those privileges to every pod in the namespace, not just the ones explicitly opted in. Fix by defaulting automountServiceAccountToken: false on the default SA in every namespace, requiring workloads to declare a specific SA, and enforcing via admission policy that pods do not use default.

Q: PSA vs Kyverno vs Gatekeeper. When would you use which? Mid: PSA is built in and covers the three standard levels; Kyverno and Gatekeeper are external and more flexible. Principal: PSA is baseline coverage for pod-spec dangerous fields, and every cluster should have it on. It does not cover CRD resources, image provenance, resource limits, label requirements, or cross-resource invariants (e.g., "every Deployment must have a corresponding PodDisruptionBudget"). Kyverno uses native Kubernetes YAML which platform teams prefer for maintainability and diffing. Gatekeeper uses Rego which is more expressive for complex policies (multi-resource joins) but has a steeper learning curve. Kubewarden compiles policies to WebAssembly for portability and performance. In practice I would run PSA plus one of the three, with Kyverno being the default choice for a team without existing Rego expertise.

Q: Explain the difference between an X.509 cert-based user and a ServiceAccount from the API server's view. Mid: X.509 certs are for humans, ServiceAccounts are for pods; both authenticate to the API server. Principal: From the API server's view they are just two different authenticators feeding into the same authorization chain. A cert with CN=alice, O=devs becomes user alice in group devs, and RBAC applies. A ServiceAccount token JWT becomes user system:serviceaccount:<ns>:<name> in groups system:serviceaccounts and system:serviceaccounts:<ns>. The security-relevant difference is lifecycle. Certs cannot be revoked without rotating the CA (the API server does not check a CRL by default), so cert-based user auth is discouraged. ServiceAccount tokens post-v1.22 are short-lived and audience-bound. For humans, OIDC federation with an external IdP is the right answer; the IdP owns revocation and MFA.

Q: An engineer wants hostPath to mount the node's /var/log into a log-shipping pod. Is that OK? Mid: hostPath is dangerous; prefer a PersistentVolume or a sidecar log agent. Principal: hostPath: /var/log is one of the least-bad hostPath uses because logs are non-secret and read-only, but two things need to be true. First, the pod runs on the node's logging DaemonSet identity, not a general-purpose ServiceAccount, so a compromise of it does not grant broad API access. Second, the mount is readOnly: true; otherwise a compromised container can plant a symlink from /var/log/pwned to /etc/shadow and read it via the read-only mount, or write to /var/log which some node components read and act on. The general rule is that hostPath is a namespace-level opt-in via PSA baseline allowing specific host paths, not a per-pod free-for-all. If the log-shipping is Fluent Bit or Vector, they have well-defined DaemonSet manifests that show up in the small allow-list.

Q: How does IRSA/Workload Identity actually work end-to-end? Mid: The pod gets a short-lived cloud credential based on its ServiceAccount instead of using node IAM. Principal: The API server publishes an OIDC discovery document at /.well-known/openid-configuration with the cluster's public keys. On EKS, that discovery document is exposed to AWS IAM via a public S3 URL registered as an OIDC identity provider in the AWS account. Pods get a projected ServiceAccount token with audience sts.amazonaws.com (configured via annotations on the SA). The pod's process calls STS AssumeRoleWithWebIdentity with that token; STS verifies the signature against the OIDC provider's public keys and, if the role's trust policy trusts system:serviceaccount:<ns>:<name>, issues temporary credentials. The SDKs handle this automatically via the AWS_WEB_IDENTITY_TOKEN_FILE and AWS_ROLE_ARN env vars that the EKS admission webhook injects. The security property is that the cloud IAM principal is scoped per-SA per-cluster, not per-node, so a compromised pod on a node with other tenants cannot steal the node's IAM role.

Q: What is the point of failurePolicy: Fail on an admission webhook and when would you set Ignore? Mid: Fail rejects the request if the webhook is unreachable; Ignore admits it. Security policies should be Fail. Principal: The tradeoff is availability vs security. A sidecar injector webhook that is down and set to Fail breaks every pod creation in the cluster including the webhook's own restart, which is a bootstrap deadlock. A security policy webhook set to Ignore is a bypass primitive: an attacker who can DoS the webhook (kill the pod, exhaust its resource quota, or overwhelm it with junk requests) then submits the previously-blocked manifest during the outage. The right posture is Fail for security policies plus a namespaceSelector that excludes kube-system and the webhook's own namespace (so the webhook can restart itself), plus HA replicas and a PodDisruptionBudget so a rolling update does not create an outage. For non-security webhooks (mutating injectors that are convenience features), Ignore can be reasonable if the injected behavior is not security-critical.

War story

A platform team ran a Kyverno policy that denied privileged: true in every namespace except kube-system. A red-team exercise found that the ingress-nginx controller pod had get secrets cluster-wide (for TLS material) and shipped with an image whose tag was v1.9.6. The tag was still mutable on the internal registry mirror because tag immutability was set on the upstream registry but not the mirror. The red team pushed a modified image to v1.9.6 on the mirror, then triggered a rolling restart of the ingress controller via a legitimate config-map change. The new pods pulled the malicious image, the exfil sidecar read the projected SA token, listed every Secret in every namespace via the API server (allowed by RBAC), and posted them to an external URL. Kyverno never fired because no pod had privileged: true; PSA baseline was in force but the ingress controller pod's spec was already compliant. The fix switched the ingress controller image to a digest pin (@sha256:...), enabled tag immutability on the mirror, and narrowed the ingress controller's get secrets from cluster-wide to a named list of TLS-holding namespaces. The audit-log signal that would have caught this in the first place was create serviceaccounts/token at unusual volume plus list secrets cluster-wide from an unexpected user-agent; both were logged but not alerted.

Sources

[1] Pod Security Standards. Kubernetes Documentation. Retrieved 2026-08. https://kubernetes.io/docs/concepts/security/pod-security-standards/

[2] Kyverno Policy Engine Documentation. Retrieved 2026-08. https://kyverno.io/docs/

[3] OPA Gatekeeper Documentation. Retrieved 2026-08. https://open-policy-agent.github.io/gatekeeper/website/docs/

[4] Controlling Access to the Kubernetes API. Kubernetes Documentation. Retrieved 2026-08. https://kubernetes.io/docs/concepts/security/controlling-access/

[5] Configure Service Accounts for Pods, Bound Service Account Tokens. Kubernetes Documentation. Retrieved 2026-08. https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/

[6] Encrypting Confidential Data at Rest. Kubernetes Documentation. Retrieved 2026-08. https://kubernetes.io/docs/tasks/administer-cluster/encrypt-data/

[7] Network Policies. Kubernetes Documentation. Retrieved 2026-08. https://kubernetes.io/docs/concepts/services-networking/network-policies/

[8] Dynamic Admission Control. Kubernetes Documentation. Retrieved 2026-08. https://kubernetes.io/docs/reference/access-authn-authz/extensible-admission-controllers/

[9] Kubernetes-goat: Interactive Kubernetes Security Learning. Retrieved 2026-08. https://madhuakula.com/kubernetes-goat/

[10] Tesla's Kubernetes Console Not Password-Protected, Hackers Mined Cryptocurrency. Ars Technica. 2018-02. https://arstechnica.com/information-technology/2018/02/tesla-cloud-resources-are-hacked-to-run-cryptocurrency-mining-malware/

[11] Sigstore Cosign Documentation. Retrieved 2026-08. https://docs.sigstore.dev/cosign/overview/

[12] CVE-2018-1002105: Kubernetes API Server Privilege Escalation. Kubernetes Security Advisory. 2018-12. https://github.com/kubernetes/kubernetes/issues/71411

[13] CVE-2020-8558: Node setting allows for neighboring hosts to bypass localhost boundary. Kubernetes Security Advisory. 2020-07. https://github.com/kubernetes/kubernetes/issues/92315

[14] CVE-2022-0811 (cr8escape): CRI-O Container Escape via kernel.core_pattern. CrowdStrike Research. 2022-03. https://www.crowdstrike.com/blog/cr8escape-new-vulnerability-discovered-in-cri-o-container-engine-cve-2022-0811/

[15] CVE-2025-1974 (IngressNightmare): ingress-nginx Remote Code Execution. Wiz Research. 2025-03. https://www.wiz.io/blog/ingress-nginx-kubernetes-vulnerabilities