Every production cluster I've touched in support has at least three of these misconfigurations. They're not obscure edge cases—they're the default settings that ship with most managed Kubernetes services and the shortcuts teams take when racing to deploy. The difference between a contained incident and a full cluster compromise usually comes down to whether you've locked down RBAC, isolated workloads with network policies, and enforced pod security standards.
Fix 1: Lock Down Default Service Account Permissions
Kubernetes creates a default service account in every namespace and automatically mounts its token into every pod. That token can talk to the API server. Most pods don't need that access, but they get it anyway.
Check what permissions your default service account has:
kubectl auth can-i --list --as=system:serviceaccount:default:default
If you see anything beyond basic self-inspection, you have a problem. An attacker who compromises a pod inherits those permissions. Disable auto-mounting at the namespace level:
apiVersion: v1
kind: ServiceAccount
metadata:
name: default
namespace: default
automountServiceAccountToken: false
For pods that legitimately need API access, create dedicated service accounts with narrow RBAC roles. Never grant cluster-admin to a service account unless you're writing a cluster operator, and even then think twice.
Fix 2: Enforce Pod Security Standards Across All Namespaces
The old PodSecurityPolicy API was deprecated because it was too complex and had weird corner cases. Pod Security Standards replaced it with three profiles: privileged, baseline, and restricted. Most workloads should run under baseline at minimum; restricted is better if your apps can handle it.
Enable PSS enforcement cluster-wide by adding labels to namespaces:
kubectl label namespace default \
pod-security.kubernetes.io/enforce=baseline \
pod-security.kubernetes.io/audit=restricted \
pod-security.kubernetes.io/warn=restricted
The baseline profile blocks the most dangerous things: privileged containers, host namespace sharing, and unsafe volume types. Restricted goes further by requiring non-root users, dropping all capabilities, and enforcing read-only root filesystems.
You'll break some legacy workloads when you turn this on. That's the point. Those workloads were running with more privilege than they needed.
Fix 3: Drop Unnecessary Linux Capabilities
Containers don't need all the capabilities that root gets on a Linux host. By default, the container runtime drops most of them, but a few dangerous ones stay enabled: CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, SETGID, SETUID, SETPCAP, NET_BIND_SERVICE, NET_RAW, SYS_CHROOT, MKNOD, AUDIT_WRITE, SETFCAP.
NET_RAW lets containers craft raw packets and perform ARP spoofing. CHOWN and DAC_OVERRIDE let them mess with file permissions in ways that can escalate privilege. Drop everything you don't explicitly need:
apiVersion: v1
kind: Pod
metadata:
name: secure-pod
spec:
containers:
- name: app
image: nginx:alpine
securityContext:
capabilities:
drop:
- ALL
add:
- NET_BIND_SERVICE
allowPrivilegeEscalation: false
runAsNonRoot: true
runAsUser: 1000
seccompProfile:
type: RuntimeDefault
Most apps only need NET_BIND_SERVICE if they bind to ports below 1024. Everything else can run with zero capabilities.
Fix 4: Implement Default-Deny Network Policies
Without network policies, every pod can talk to every other pod in the cluster. That's convenient for getting started, but it means a compromised frontend pod can directly query your database.
Start with a default-deny rule in every namespace:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: production
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
That blocks all traffic. Now add explicit allow rules for legitimate communication paths:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-frontend-to-backend
namespace: production
spec:
podSelector:
matchLabels:
app: backend
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app: frontend
ports:
- protocol: TCP
port: 8080
This requires you to understand your application's communication patterns. That's good—you should know who talks to whom. Network policies also block lateral movement after a breach, turning a full cluster compromise into an isolated incident.
Fix 5: Run a Validating Admission Controller
Admission controllers intercept API requests before they're persisted and can reject or modify them. A validating webhook lets you enforce policies that Pod Security Standards don't cover: base image sources, resource limits, required labels, banned environment variables.
Open Policy Agent and Kyverno are the two most common options. Kyverno is easier to get started with because policies are just Kubernetes YAML:
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-resource-limits
spec:
validationFailureAction: enforce
rules:
- name: check-limits
match:
resources:
kinds:
- Pod
validate:
message: "All containers must define CPU and memory limits"
pattern:
spec:
containers:
- resources:
limits:
memory: "?*"
cpu: "?*"
This policy blocks any pod that doesn't specify resource limits. No limits means a compromised container can consume all available CPU and memory, crashing other workloads. An admission controller enforces policies at creation time, so bad configs never make it into the cluster.
Fix 6: Enable Audit Logging and Ship It Offsite
Kubernetes audit logs record every API request: who made it, what they asked for, and whether it succeeded. You need these logs to investigate incidents and detect suspicious behavior, but they're not enabled by default in most managed clusters.
On self-hosted clusters, configure audit logging in the API server manifest:
apiVersion: v1
kind: Pod
metadata:
name: kube-apiserver
namespace: kube-system
spec:
containers:
- command:
- kube-apiserver
- --audit-policy-file=/etc/kubernetes/audit-policy.yaml
- --audit-log-path=/var/log/kubernetes/audit.log
- --audit-log-maxage=30
- --audit-log-maxbackup=10
- --audit-log-maxsize=100
Your audit policy should log authentication failures, privilege escalations, secret access, and exec/attach requests. Ship these logs to a separate system—if an attacker gets cluster-admin, they can delete local logs.
On managed clusters (EKS, GKE, AKS), enable audit logging through the cloud provider's console and route logs to CloudWatch, Cloud Logging, or Azure Monitor. Then forward them to your SIEM.
Fix 7: Restrict API Server Access by IP
The API server is the control plane for your entire cluster. If someone can reach it and has valid credentials—even a service account token stolen from a pod—they can do real damage. Restrict network access to known IP ranges.
On managed clusters, this is usually a firewall rule or authorized networks setting. For GKE:
gcloud container clusters update CLUSTER_NAME \
--enable-master-authorized-networks \
--master-authorized-networks 203.0.113.0/24,198.51.100.0/24
For self-hosted clusters, put the API server behind a firewall or VPN. Never expose it directly to the internet unless you have a very good reason and multiple layers of authentication.
Also disable anonymous authentication if it's still on:
--anonymous-auth=false
Anonymous requests default to the system:anonymous user with minimal permissions, but that's still more access than an unauthenticated attacker should have.
Fix 8: Scan Images and Block Vulnerable Base Layers
Running containers with known CVEs is asking for trouble. Most attacks start with a public exploit against an unpatched library. Set up automated image scanning in your CI pipeline and block deployments of images with high or critical vulnerabilities.
If you're on a managed registry (ECR, GCR, ACR), enable the built-in scanning feature. For ECR:
aws ecr put-image-scanning-configuration \
--repository-name myapp \
--image-scanning-configuration scanOnPush=true
Pair that with an admission controller policy that checks scan results:
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: block-vulnerable-images
spec:
validationFailureAction: enforce
rules:
- name: check-vulnerabilities
match:
resources:
kinds:
- Pod
validate:
message: "Image has high or critical vulnerabilities"
deny:
conditions:
any:
- key: "{{ images.*.vulnerabilities.critical || `0` }}"
operator: GreaterThan
value: 0
- key: "{{ images.*.vulnerabilities.high || `0` }}"
operator: GreaterThan
value: 5
This blocks images with any critical CVE or more than five high-severity ones. Adjust thresholds based on your risk tolerance, but don't let high-risk images into production.
So what if you inherit a cluster that has none of this?
Start with the default service account and Pod Security Standards. Those two fixes block the easiest privilege escalation paths and cost nothing to implement. Network policies come next—they take more planning but contain blast radius better than anything else on this list.
Run these changes in a staging environment first. You will break things. Some workloads depend on the loose defaults, and you'll need to adjust pod specs, create dedicated service accounts, or refactor how services discover each other. That's fine. Better to find out in staging than during an incident.
Document every exception you make. If a pod needs elevated privileges, write down why and set a review date. Exceptions pile up fast, and six months later nobody remembers why that one deployment runs as root.
How do I test if my RBAC rules are too permissive?
Use kubectl auth can-i --list --as=system:serviceaccount:NAMESPACE:SERVICEACCOUNT to see what a service account can do. If it can list secrets in other namespaces, create pods, or get nodes, you've granted too much. Service accounts should only access resources in their own namespace unless they're cluster infrastructure components.
Will network policies slow down my cluster?
No. The CNI plugin enforces network policies in the kernel using iptables or eBPF. The performance impact is negligible—measured in microseconds per packet. If you're seeing latency, it's not the network policies.
Can I enforce Pod Security Standards without breaking existing workloads?
Yes, using audit and warn modes. Set pod-security.kubernetes.io/audit=baseline and pod-security.kubernetes.io/warn=baseline on your namespaces first. This logs violations and shows warnings but doesn't block anything. Review the audit logs, fix the flagged workloads, then flip to enforce mode.
What's the fastest way to find overprivileged service accounts?
Use a tool like rbac-lookup or rakkess to query RBAC permissions across the cluster. Look for service accounts with verbs like *, resources like *, or cluster-wide roles bound to namespace-scoped accounts. Those are red flags.
What to check first
If you take nothing else from this, check whether your default service accounts have API access and whether Pod Security Standards are enforced. Those two misconfigurations show up in almost every cluster and are the first things attackers probe after compromising a pod. The rest of the list matters, but those two fixes block the most common attack paths with the least effort.
Kubernetes security isn't complicated—it's just not the default. Every fix here is a one-time configuration change or a policy you write once and apply everywhere. The hard part is remembering to do it before the first incident, not after.
