Moving a Kubernetes cluster to production without a systematic checklist invites outages, security breaches, and performance issues that could have been prevented. This guide walks you through the essential pre-launch steps that separate proof-of-concept clusters from production-grade infrastructure.
Pre-Flight: Cluster Foundation
Before diving into application workloads, verify your cluster's foundational configuration.
Control Plane High Availability
Production clusters require at least three control plane nodes distributed across availability zones or failure domains. Single control plane deployments create a catastrophic single point of failure.
Verify control plane node distribution:
kubectl get nodes -l node-role.kubernetes.io/control-plane --show-labels
Check that control plane components are healthy:
kubectl get componentstatuses
kubectl get pods -n kube-system
etcd Backup and Recovery
etcd stores all cluster state. Without working backups, cluster failures become data loss events.
Set up automated etcd snapshots. For managed Kubernetes services, verify the provider's backup schedule and test restoration. For self-managed clusters, implement regular snapshot automation:
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-snapshot-$(date +%Y%m%d-%H%M%S).db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
Document and test the full cluster restoration procedure before launch.
Node Operating System Hardening
Apply standard Linux hardening to all cluster nodes. Disable unnecessary services, configure automatic security updates, and ensure SSH access is restricted to bastion hosts or VPN endpoints.
Verify kernel parameters for network performance:
sysctl net.ipv4.ip_forward
sysctl net.bridge.bridge-nf-call-iptables
Both should return 1 for proper pod networking.
Security: Lock Down Before Launch
Kubernetes security requires layers of defense. Each control prevents specific attack vectors.
Network Policies
By default, pods can communicate with any other pod in the cluster. This lateral movement enables attackers who compromise one workload to reach others.
Implement default-deny network policies in every namespace:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: production-app
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
Then add explicit allow policies for legitimate traffic. Test network policies in staging with identical configurations before production deployment.
Pod Security Standards
Restrict pod capabilities to prevent container escapes and privilege escalation. Kubernetes Pod Security Standards replace the deprecated PodSecurityPolicy.
Enforce restricted policy on production namespaces:
apiVersion: v1
kind: Namespace
metadata:
name: production-app
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/audit: restricted
pod-security.kubernetes.io/warn: restricted
The restricted profile blocks privileged containers, host network access, and dangerous capabilities. Applications requiring elevated permissions need explicit justification and additional security controls.
RBAC Configuration
Role-Based Access Control governs who can perform actions in the cluster. Overly permissive RBAC grants developers cluster-admin rights they never needed.
Audit existing cluster role bindings:
kubectl get clusterrolebindings -o json | jq -r '.items[] | select(.roleRef.name=="cluster-admin") | .metadata.name'
Remove default cluster-admin bindings not required for cluster operation. Create namespace-scoped roles with minimum required permissions:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: app-developer
namespace: production-app
rules:
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["get", "list", "watch"]
- apiGroups: [""]
resources: ["pods", "pods/log"]
verbs: ["get", "list", "watch"]
Secrets Management
Never store sensitive data in ConfigMaps or environment variables visible in pod specifications. Use Kubernetes Secrets as a baseline, but recognize they are base64-encoded, not encrypted at rest by default.
Enable encryption at rest for secrets:
apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
- resources:
- secrets
providers:
- aescbc:
keys:
- name: key1
secret: <base64-encoded-32-byte-key>
- identity: {}
For higher security requirements, integrate external secret stores like HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault using the Secrets Store CSI driver.
Resource Management: Prevent Resource Contention
Without resource constraints, a single misbehaving application can starve the entire cluster.
Resource Requests and Limits
Every production pod must specify CPU and memory requests and limits. Requests guarantee resources; limits prevent overconsumption.
apiVersion: v1
kind: Pod
metadata:
name: production-app
spec:
containers:
- name: app
image: app:latest
resources:
requests:
memory: "256Mi"
cpu: "250m"
limits:
memory: "512Mi"
cpu: "500m"
Set requests based on observed usage in staging. Set limits with headroom for traffic spikes.
LimitRanges and ResourceQuotas
Enforce resource specifications at the namespace level. LimitRanges set defaults and boundaries for individual pods:
apiVersion: v1
kind: LimitRange
metadata:
name: production-limits
namespace: production-app
spec:
limits:
- max:
cpu: "2"
memory: "2Gi"
min:
cpu: "100m"
memory: "64Mi"
default:
cpu: "500m"
memory: "512Mi"
defaultRequest:
cpu: "250m"
memory: "256Mi"
type: Container
ResourceQuotas cap total namespace consumption:
apiVersion: v1
kind: ResourceQuota
metadata:
name: production-quota
namespace: production-app
spec:
hard:
requests.cpu: "10"
requests.memory: "20Gi"
limits.cpu: "20"
limits.memory: "40Gi"
pods: "50"
Quality of Service Classes
Kubernetes assigns QoS classes based on resource configuration. Guaranteed pods (requests equal limits) receive priority during resource pressure. Burstable pods can exceed requests up to limits. BestEffort pods (no requests or limits) are evicted first.
Critical applications should run as Guaranteed class.
Observability: Monitor Everything
You cannot troubleshoot what you cannot see. Production clusters require comprehensive monitoring and logging.
Metrics Collection
Deploy a metrics stack capable of storing and querying time-series data. Prometheus remains the de facto standard for Kubernetes monitoring.
Minimum required metrics:
- Node CPU, memory, disk, and network utilization
- Pod CPU and memory usage by namespace and workload
- Container restart counts and crash loop patterns
- API server request rates and latencies
- etcd performance metrics
- Persistent volume capacity and IOPS
Configure metric retention policies appropriate for your incident investigation timeframe.
Logging Infrastructure
Centralize logs from all pods and nodes. Without aggregated logs, troubleshooting multi-pod applications becomes impossible.
Common logging stacks include ELK (Elasticsearch, Logstash, Kibana), EFK (Elasticsearch, Fluentd, Kibana), or Loki. Select based on your scale and existing infrastructure.
Configure log retention based on compliance requirements and storage costs. Implement log rotation to prevent disk exhaustion on nodes.
Health Checks and Probes
Kubernetes cannot determine application health without explicit probes. Configure liveness and readiness probes for every container:
apiVersion: v1
kind: Pod
metadata:
name: production-app
spec:
containers:
- name: app
image: app:latest
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
timeoutSeconds: 5
failureThreshold: 3
readinessProbe:
httpGet:
path: /ready
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 3
failureThreshold: 2
Liveness probes restart unhealthy containers. Readiness probes remove unready pods from service endpoints. Startup probes handle slow-starting applications without extending liveness probe delays.
Alerting Rules
Define alerts for conditions requiring human intervention. Effective alerts are actionable, not informational.
Critical production alerts:
- Node not ready for more than five minutes
- Pod crash loop (more than three restarts in ten minutes)
- Persistent volume approaching capacity (over 85% full)
- API server error rate exceeds threshold
- Certificate expiration within 30 days
- Deployment rollout stuck or failing
Route alerts to on-call teams through appropriate channels.
High Availability and Resilience
Design for failure. Every component will eventually fail.
Pod Disruption Budgets
PodDisruptionBudgets prevent voluntary disruptions (node drains, cluster upgrades) from taking down too many pods simultaneously:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: production-app-pdb
namespace: production-app
spec:
minAvailable: 2
selector:
matchLabels:
app: production-app
Set minAvailable or maxUnavailable based on your availability requirements. Applications requiring high availability should maintain at least two healthy replicas at all times.
Anti-Affinity Rules
Distribute pod replicas across nodes and zones to survive node and zone failures:
apiVersion: apps/v1
kind: Deployment
metadata:
name: production-app
spec:
replicas: 3
template:
spec:
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: app
operator: In
values:
- production-app
topologyKey: kubernetes.io/hostname
For multi-zone clusters, add zone-level anti-affinity with preferredDuringScheduling rules.
Backup Strategy
Beyond etcd backups, implement application-level backup for stateful workloads. Use Velero or similar tools to back up persistent volumes, namespaces, and cluster resources.
Test restoration procedures regularly. Untested backups are not backups.
Incident Response Preparation
Plan your response before incidents occur.
Runbooks and Documentation
Document common failure scenarios and resolution steps. Runbooks should be executable by on-call engineers unfamiliar with specific applications.
Include:
- Cluster access procedures and credentials
- Log and metric query examples
- Rollback procedures for deployments
- Contact information for application owners
- Escalation paths for extended outages
Break-Glass Procedures
Define emergency access procedures for when normal authentication fails. Break-glass accounts bypass standard RBAC but must be audited and time-limited.
Communication Plans
Establish incident communication channels. Define who communicates with stakeholders, how status updates are delivered, and when post-incident reviews occur.
Pre-Launch Checklist Summary
Before deploying to production, verify:
- [ ] Control plane runs on at least three nodes across failure domains
- [ ] etcd backups run automatically and restoration is tested
- [ ] Network policies enforce default-deny in all namespaces
- [ ] Pod Security Standards enforced at namespace level
- [ ] RBAC follows least-privilege principle
- [ ] Secrets encrypted at rest or managed by external store
- [ ] All pods specify resource requests and limits
- [ ] LimitRanges and ResourceQuotas configured per namespace
- [ ] Metrics collection covers cluster and application layers
- [ ] Centralized logging captures all pod and node logs
- [ ] Health probes configured for all containers
- [ ] Critical alerts defined and routed to on-call teams
- [ ] PodDisruptionBudgets prevent excessive disruption
- [ ] Pod anti-affinity distributes replicas across nodes
- [ ] Backup and restore procedures tested
- [ ] Incident runbooks documented and accessible
- [ ] Break-glass procedures defined
Conclusion
Production Kubernetes deployments demand systematic preparation across security, resource management, observability, and resilience. This checklist provides the foundation for stable, secure clusters that survive real-world operational challenges. Work through each section methodically, test thoroughly in staging, and document your decisions. The upfront investment prevents costly outages and security incidents after launch. Your production users and on-call team will thank you.
