Most production Kubernetes clusters ship with defaults that favor convenience over security. I've seen clusters breached because a single pod ran as root or because default service accounts carried cluster-admin privileges. The attack surface shrinks fast when you tighten RBAC, lock down network paths, and enforce pod-level restrictions.
1. Lock down default service accounts
Every namespace gets a default service account that mounts a token into every pod. That token often has more permissions than the workload needs. Attackers who compromise a pod can use that token to query the API server and pivot.
Disable auto-mounting for the default account in each namespace:
apiVersion: v1
kind: ServiceAccount
metadata:
name: default
namespace: production
automountServiceAccountToken: false
Create dedicated service accounts for workloads that actually need API access. Most pods don't.
For pods that do need a token, scope the RBAC role to the minimum verbs and resources. A metrics collector might need get and list on pods in one namespace—nothing more.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: pod-reader
namespace: production
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list"]
Bind that role to the service account, then reference the account in your deployment spec. The blast radius drops to near zero if that pod is compromised.
2. Enforce Pod Security Standards with admission control
The built-in Pod Security Admission controller replaced PodSecurityPolicy. It enforces three levels: privileged, baseline, and restricted. Most workloads should run under the restricted profile.
Label your namespaces to enforce the policy:
apiVersion: v1
kind: Namespace
metadata:
name: production
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/audit: restricted
pod-security.kubernetes.io/warn: restricted
The restricted profile blocks privileged containers, host namespaces, host ports, and requires pods to drop all capabilities and run as non-root. Existing workloads that violate the policy will fail to deploy until you fix their manifests.
Set runAsNonRoot: true and specify a high UID in your pod security context:
securityContext:
runAsNonRoot: true
runAsUser: 10000
fsGroup: 10000
seccompProfile:
type: RuntimeDefault
Drop all Linux capabilities and add back only what you need:
containers:
- name: app
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
If a container needs NET_BIND_SERVICE to bind a low port, add that one capability. Don't leave ALL in place.
3. Apply network policies to isolate workloads
By default, every pod can talk to every other pod. An attacker who pops a front-end service can probe your database directly.
Network policies act as distributed firewalls. Start by denying all traffic, then allow only what each service needs.
Default-deny for a namespace:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: production
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
Then whitelist specific paths. Let your API pods reach the database on port 5432:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: api-to-db
namespace: production
spec:
podSelector:
matchLabels:
app: api
policyTypes:
- Egress
egress:
- to:
- podSelector:
matchLabels:
app: postgres
ports:
- protocol: TCP
port: 5432
Allow ingress to your front-end only from the ingress controller:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-from-ingress
namespace: production
spec:
podSelector:
matchLabels:
app: frontend
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
name: ingress-nginx
ports:
- protocol: TCP
port: 8080
Test policies in audit mode first. Most CNI plugins log denied connections, so you can tune rules before they break production traffic.
4. Restrict RBAC bindings to the minimum scope
Cluster-admin is a loaded gun. I've seen CI pipelines, monitoring agents, and operators all granted cluster-admin because it was easier than writing a narrow role.
Audit existing cluster role bindings:
kubectl get clusterrolebindings -o json | jq -r '.items[] | select(.roleRef.name=="cluster-admin") | .metadata.name'
For each binding, ask whether that identity really needs cluster-wide write access. Most don't.
Replace broad bindings with namespace-scoped roles. A deployment tool that manages apps in the staging namespace only needs a role in that namespace:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: deploy-manager
namespace: staging
rules:
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["get", "list", "create", "update", "patch", "delete"]
- apiGroups: [""]
resources: ["pods", "services", "configmaps", "secrets"]
verbs: ["get", "list", "create", "update", "patch", "delete"]
Bind it with a RoleBinding, not a ClusterRoleBinding. The identity can't touch other namespaces.
For cluster-scoped resources like nodes or persistent volumes, write a ClusterRole that grants only the needed verbs. A volume provisioner might need create and delete on PVs but not get or list on secrets.
5. Enable audit logging and watch for risky API calls
The API server audit log records every request. You can spot privilege escalation, secret access, and exec sessions in near real-time.
Define an audit policy that logs metadata for most calls and full request/response bodies for sensitive resources:
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
- level: RequestResponse
verbs: ["create", "update", "patch", "delete"]
resources:
- group: ""
resources: ["secrets", "configmaps"]
- level: Metadata
omitStages:
- RequestReceived
Pass the policy to the API server with --audit-policy-file and send logs to a file or webhook backend. Forward them to your log aggregator so you can alert on patterns like repeated 403s from a service account or unexpected exec calls into production pods.
Watch for:
- Service accounts hitting forbidden endpoints repeatedly (reconnaissance)
- Secrets or tokens accessed outside normal patterns
- Pod exec or attach commands from unfamiliar source IPs
- Role or ClusterRole modifications
6. Use admission webhooks to enforce custom policies
Pod Security Standards cover common cases, but you might need org-specific rules. Admission webhooks let you reject or mutate resources before they're persisted.
A validating webhook can block images from untrusted registries:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata:
name: image-policy
webhooks:
- name: validate-images.example.com
clientConfig:
service:
name: image-validator
namespace: kube-system
path: /validate
caBundle: <base64-ca-cert>
rules:
- operations: ["CREATE", "UPDATE"]
apiGroups: [""]
apiVersions: ["v1"]
resources: ["pods"]
admissionReviewVersions: ["v1"]
sideEffects: None
The webhook service inspects each pod's image field and returns an admission response. You can deny pods that don't pull from your internal registry or that lack a signature.
Mutating webhooks can inject sidecars, set resource limits, or add security contexts. Use them to enforce defaults without requiring every developer to remember the right YAML.
Tools like OPA Gatekeeper and Kyverno provide policy engines with pre-built rules. They're easier than writing webhook code from scratch.
7. Isolate the control plane and etcd
The API server and etcd hold the keys to your entire cluster. Isolate them on dedicated nodes or a separate network segment.
Taint control-plane nodes so workloads don't schedule there:
kubectl taint nodes control-plane-1 node-role.kubernetes.io/control-plane:NoSchedule
Use network policies or firewall rules to restrict access to the API server. Only your ingress controller, CI system, and operator workstations should reach port 6443.
Encrypt etcd at rest by passing --encryption-provider-config to the API server. The config file specifies which resources to encrypt and which key provider to use:
apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
- resources:
- secrets
providers:
- aescbc:
keys:
- name: key1
secret: <base64-32-byte-key>
- identity: {}
Rotate the encryption key periodically and re-encrypt all secrets:
kubectl get secrets --all-namespaces -o json | kubectl replace -f -
Restrict etcd client certificates to the API server. No other component should talk to etcd directly.
8. Scan images and block known vulnerabilities
Containers ship with OS packages and libraries that carry CVEs. Scanning catches high-severity flaws before they reach production.
Integrate a scanner into your CI pipeline. Tools like Trivy, Grype, or Clair can fail builds when they find critical vulnerabilities:
trivy image --severity HIGH,CRITICAL --exit-code 1 myapp:latest
Run scans in-cluster too. Admission webhooks can query a scanner API and reject images with unpatched CVEs.
Keep base images minimal. Distroless or Alpine-based images have fewer packages and a smaller attack surface. A Go binary doesn't need a full Debian userland.
Sign images with cosign or Notary so you can verify that an image came from your build pipeline and wasn't tampered with. Admission webhooks can enforce signature verification cluster-wide.
What happens after you harden a cluster?
Deployments that relied on permissive defaults will break. Pods that ran as root won't start. Services that talked to every endpoint will time out when network policies drop their packets.
Plan a migration window. Apply policies in audit or warn mode first, watch the logs, and fix manifests before you enforce. Most teams find that 80% of workloads can adopt the restricted profile with minor YAML changes.
Automation helps. Use a policy engine to generate network policies from observed traffic, or write a script that patches deployment specs with secure defaults.
Document which service accounts have API access and why. When the next breach attempt happens, you'll know exactly what the attacker can reach.
How do I test network policies without breaking production?
Set policyTypes to Ingress only and leave Egress open. Observe denied ingress in CNI logs, tune the rules, then add egress restrictions.
Some CNIs support a dry-run or audit mode that logs policy decisions without enforcing them. Cilium and Calico both offer this.
Can I use Pod Security Standards with older clusters?
Pod Security Admission shipped in Kubernetes 1.23 and became stable in 1.25. Older clusters can install the PSP replacement as an admission webhook or use OPA/Kyverno to enforce similar policies.
What if a workload legitimately needs privileged access?
Isolate it in a separate namespace with a permissive policy label. Apply stricter network policies to limit what that workload can reach. Use audit logging to track every action it takes.
How often should I rotate service account tokens?
Bound tokens expire automatically based on the pod's lifetime. Long-lived tokens in secrets should rotate every 90 days or whenever a pod is compromised.
Start with RBAC and network policies
You don't have to implement all eight fixes at once. Begin with RBAC—audit cluster-admin bindings and replace them with scoped roles. Then apply default-deny network policies to one namespace and whitelist only necessary traffic.
Those two changes block the most common attack paths: lateral movement after pod compromise and privilege escalation via over-permissioned service accounts. Add Pod Security Standards next, then layer in admission webhooks and image scanning as your process matures.
Hardening a cluster is incremental work, not a one-time checklist. Each fix narrows the window an attacker has to move from initial access to full control.
