Skip to content
Back to Blog
Linux & Server12 min read

Kubernetes Production Checklist 2026: Deploy With Confidence

A comprehensive pre-launch checklist for Kubernetes deployments covering security policies, resource limits, monitoring, backup strategies, and incident response to ensure production readiness.

Written by Abdul AbrorTechnical Hosting Support Engineer
Kubernetes Production Checklist 2026: Deploy With Confidence
On this page

Moving a Kubernetes cluster to production without a systematic checklist invites outages, security breaches, and performance issues that could have been prevented. This guide walks you through the essential pre-launch steps that separate proof-of-concept clusters from production-grade infrastructure.

Pre-Flight: Cluster Foundation

Before diving into application workloads, verify your cluster's foundational configuration.

Control Plane High Availability

Production clusters require at least three control plane nodes distributed across availability zones or failure domains. Single control plane deployments create a catastrophic single point of failure.

Verify control plane node distribution:

kubectl get nodes -l node-role.kubernetes.io/control-plane --show-labels

Check that control plane components are healthy:

kubectl get componentstatuses
kubectl get pods -n kube-system

etcd Backup and Recovery

etcd stores all cluster state. Without working backups, cluster failures become data loss events.

Set up automated etcd snapshots. For managed Kubernetes services, verify the provider's backup schedule and test restoration. For self-managed clusters, implement regular snapshot automation:

ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-snapshot-$(date +%Y%m%d-%H%M%S).db \
  --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/server.crt \
  --key=/etc/kubernetes/pki/etcd/server.key

Document and test the full cluster restoration procedure before launch.

Node Operating System Hardening

Apply standard Linux hardening to all cluster nodes. Disable unnecessary services, configure automatic security updates, and ensure SSH access is restricted to bastion hosts or VPN endpoints.

Verify kernel parameters for network performance:

sysctl net.ipv4.ip_forward
sysctl net.bridge.bridge-nf-call-iptables

Both should return 1 for proper pod networking.

Security: Lock Down Before Launch

Kubernetes security requires layers of defense. Each control prevents specific attack vectors.

Network Policies

By default, pods can communicate with any other pod in the cluster. This lateral movement enables attackers who compromise one workload to reach others.

Implement default-deny network policies in every namespace:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-all
  namespace: production-app
spec:
  podSelector: {}
  policyTypes:
  - Ingress
  - Egress

Then add explicit allow policies for legitimate traffic. Test network policies in staging with identical configurations before production deployment.

Pod Security Standards

Restrict pod capabilities to prevent container escapes and privilege escalation. Kubernetes Pod Security Standards replace the deprecated PodSecurityPolicy.

Enforce restricted policy on production namespaces:

apiVersion: v1
kind: Namespace
metadata:
  name: production-app
  labels:
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/audit: restricted
    pod-security.kubernetes.io/warn: restricted

The restricted profile blocks privileged containers, host network access, and dangerous capabilities. Applications requiring elevated permissions need explicit justification and additional security controls.

RBAC Configuration

Role-Based Access Control governs who can perform actions in the cluster. Overly permissive RBAC grants developers cluster-admin rights they never needed.

Audit existing cluster role bindings:

kubectl get clusterrolebindings -o json | jq -r '.items[] | select(.roleRef.name=="cluster-admin") | .metadata.name'

Remove default cluster-admin bindings not required for cluster operation. Create namespace-scoped roles with minimum required permissions:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: app-developer
  namespace: production-app
rules:
- apiGroups: ["apps"]
  resources: ["deployments", "replicasets"]
  verbs: ["get", "list", "watch"]
- apiGroups: [""]
  resources: ["pods", "pods/log"]
  verbs: ["get", "list", "watch"]

Secrets Management

Never store sensitive data in ConfigMaps or environment variables visible in pod specifications. Use Kubernetes Secrets as a baseline, but recognize they are base64-encoded, not encrypted at rest by default.

Enable encryption at rest for secrets:

apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
  - resources:
    - secrets
    providers:
    - aescbc:
        keys:
        - name: key1
          secret: <base64-encoded-32-byte-key>
    - identity: {}

For higher security requirements, integrate external secret stores like HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault using the Secrets Store CSI driver.

Resource Management: Prevent Resource Contention

Without resource constraints, a single misbehaving application can starve the entire cluster.

Resource Requests and Limits

Every production pod must specify CPU and memory requests and limits. Requests guarantee resources; limits prevent overconsumption.

apiVersion: v1
kind: Pod
metadata:
  name: production-app
spec:
  containers:
  - name: app
    image: app:latest
    resources:
      requests:
        memory: "256Mi"
        cpu: "250m"
      limits:
        memory: "512Mi"
        cpu: "500m"

Set requests based on observed usage in staging. Set limits with headroom for traffic spikes.

LimitRanges and ResourceQuotas

Enforce resource specifications at the namespace level. LimitRanges set defaults and boundaries for individual pods:

apiVersion: v1
kind: LimitRange
metadata:
  name: production-limits
  namespace: production-app
spec:
  limits:
  - max:
      cpu: "2"
      memory: "2Gi"
    min:
      cpu: "100m"
      memory: "64Mi"
    default:
      cpu: "500m"
      memory: "512Mi"
    defaultRequest:
      cpu: "250m"
      memory: "256Mi"
    type: Container

ResourceQuotas cap total namespace consumption:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: production-quota
  namespace: production-app
spec:
  hard:
    requests.cpu: "10"
    requests.memory: "20Gi"
    limits.cpu: "20"
    limits.memory: "40Gi"
    pods: "50"

Quality of Service Classes

Kubernetes assigns QoS classes based on resource configuration. Guaranteed pods (requests equal limits) receive priority during resource pressure. Burstable pods can exceed requests up to limits. BestEffort pods (no requests or limits) are evicted first.

Critical applications should run as Guaranteed class.

Observability: Monitor Everything

You cannot troubleshoot what you cannot see. Production clusters require comprehensive monitoring and logging.

Metrics Collection

Deploy a metrics stack capable of storing and querying time-series data. Prometheus remains the de facto standard for Kubernetes monitoring.

Minimum required metrics:

  • Node CPU, memory, disk, and network utilization
  • Pod CPU and memory usage by namespace and workload
  • Container restart counts and crash loop patterns
  • API server request rates and latencies
  • etcd performance metrics
  • Persistent volume capacity and IOPS

Configure metric retention policies appropriate for your incident investigation timeframe.

Logging Infrastructure

Centralize logs from all pods and nodes. Without aggregated logs, troubleshooting multi-pod applications becomes impossible.

Common logging stacks include ELK (Elasticsearch, Logstash, Kibana), EFK (Elasticsearch, Fluentd, Kibana), or Loki. Select based on your scale and existing infrastructure.

Configure log retention based on compliance requirements and storage costs. Implement log rotation to prevent disk exhaustion on nodes.

Health Checks and Probes

Kubernetes cannot determine application health without explicit probes. Configure liveness and readiness probes for every container:

apiVersion: v1
kind: Pod
metadata:
  name: production-app
spec:
  containers:
  - name: app
    image: app:latest
    livenessProbe:
      httpGet:
        path: /healthz
        port: 8080
      initialDelaySeconds: 30
      periodSeconds: 10
      timeoutSeconds: 5
      failureThreshold: 3
    readinessProbe:
      httpGet:
        path: /ready
        port: 8080
      initialDelaySeconds: 10
      periodSeconds: 5
      timeoutSeconds: 3
      failureThreshold: 2

Liveness probes restart unhealthy containers. Readiness probes remove unready pods from service endpoints. Startup probes handle slow-starting applications without extending liveness probe delays.

Alerting Rules

Define alerts for conditions requiring human intervention. Effective alerts are actionable, not informational.

Critical production alerts:

  • Node not ready for more than five minutes
  • Pod crash loop (more than three restarts in ten minutes)
  • Persistent volume approaching capacity (over 85% full)
  • API server error rate exceeds threshold
  • Certificate expiration within 30 days
  • Deployment rollout stuck or failing

Route alerts to on-call teams through appropriate channels.

High Availability and Resilience

Design for failure. Every component will eventually fail.

Pod Disruption Budgets

PodDisruptionBudgets prevent voluntary disruptions (node drains, cluster upgrades) from taking down too many pods simultaneously:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: production-app-pdb
  namespace: production-app
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: production-app

Set minAvailable or maxUnavailable based on your availability requirements. Applications requiring high availability should maintain at least two healthy replicas at all times.

Anti-Affinity Rules

Distribute pod replicas across nodes and zones to survive node and zone failures:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: production-app
spec:
  replicas: 3
  template:
    spec:
      affinity:
        podAntiAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
          - labelSelector:
              matchExpressions:
              - key: app
                operator: In
                values:
                - production-app
            topologyKey: kubernetes.io/hostname

For multi-zone clusters, add zone-level anti-affinity with preferredDuringScheduling rules.

Backup Strategy

Beyond etcd backups, implement application-level backup for stateful workloads. Use Velero or similar tools to back up persistent volumes, namespaces, and cluster resources.

Test restoration procedures regularly. Untested backups are not backups.

Incident Response Preparation

Plan your response before incidents occur.

Runbooks and Documentation

Document common failure scenarios and resolution steps. Runbooks should be executable by on-call engineers unfamiliar with specific applications.

Include:

  • Cluster access procedures and credentials
  • Log and metric query examples
  • Rollback procedures for deployments
  • Contact information for application owners
  • Escalation paths for extended outages

Break-Glass Procedures

Define emergency access procedures for when normal authentication fails. Break-glass accounts bypass standard RBAC but must be audited and time-limited.

Communication Plans

Establish incident communication channels. Define who communicates with stakeholders, how status updates are delivered, and when post-incident reviews occur.

Pre-Launch Checklist Summary

Before deploying to production, verify:

  • [ ] Control plane runs on at least three nodes across failure domains
  • [ ] etcd backups run automatically and restoration is tested
  • [ ] Network policies enforce default-deny in all namespaces
  • [ ] Pod Security Standards enforced at namespace level
  • [ ] RBAC follows least-privilege principle
  • [ ] Secrets encrypted at rest or managed by external store
  • [ ] All pods specify resource requests and limits
  • [ ] LimitRanges and ResourceQuotas configured per namespace
  • [ ] Metrics collection covers cluster and application layers
  • [ ] Centralized logging captures all pod and node logs
  • [ ] Health probes configured for all containers
  • [ ] Critical alerts defined and routed to on-call teams
  • [ ] PodDisruptionBudgets prevent excessive disruption
  • [ ] Pod anti-affinity distributes replicas across nodes
  • [ ] Backup and restore procedures tested
  • [ ] Incident runbooks documented and accessible
  • [ ] Break-glass procedures defined

Conclusion

Production Kubernetes deployments demand systematic preparation across security, resource management, observability, and resilience. This checklist provides the foundation for stable, secure clusters that survive real-world operational challenges. Work through each section methodically, test thoroughly in staging, and document your decisions. The upfront investment prevents costly outages and security incidents after launch. Your production users and on-call team will thank you.

FAQ

How many control plane nodes do I need?

Three control plane nodes provide high availability with tolerance for one node failure. Five nodes tolerate two failures but increase management complexity. Clusters smaller than three control plane nodes are not production-ready.

Should I use namespace-level or cluster-level network policies?

Start with namespace-level default-deny policies. Cluster-level policies risk unintended consequences across multiple teams. As you mature, cluster-level policies can enforce organization-wide standards while namespaces add application-specific rules.

What resource request-to-limit ratio should I use?

For predictable workloads, set requests equal to limits (Guaranteed QoS). For bursty workloads, set limits 1.5-2x higher than requests. Avoid setting limits much higher than requests as this enables resource contention.

How long should I retain metrics and logs?

Retain high-resolution metrics for at least seven days and lower-resolution metrics for 30-90 days. Retain logs for 30-90 days minimum, longer if compliance requires. Balance retention against storage costs and investigation needs.

Do I need a service mesh for production?

Service meshes add observability, security, and traffic management but increase complexity. Start without a mesh. Add one when you need mTLS between services, sophisticated traffic routing, or distributed tracing that application instrumentation cannot provide.