Kubernetes has become the standard for container orchestration, but deploying a production-grade cluster requires more than running a few install commands. This guide walks you through building a secure, maintainable Kubernetes environment from scratch, covering architecture decisions, tooling choices, and hardening steps that matter in real hosting environments.
Understanding Kubernetes Architecture
Before deploying anything, understand what you're building. A Kubernetes cluster consists of control plane nodes (managing the cluster state) and worker nodes (running your workloads). The control plane runs the API server, scheduler, controller manager, and etcd (the key-value store holding cluster state). Worker nodes run the kubelet (node agent), container runtime, and kube-proxy (network rules).
For production, you need at least three control plane nodes for high availability and multiple worker nodes based on your workload requirements. Single control plane setups are acceptable for development but represent a single point of failure.
Choosing Your Deployment Method
Several tools can bootstrap Kubernetes clusters, each with different trade-offs:
kubeadm remains the standard for manual cluster deployment. It handles certificate generation, control plane component setup, and node joining. Use kubeadm when you need full control over the cluster configuration and understand the underlying components.
Kubespray uses Ansible playbooks to deploy production-ready clusters with sensible defaults for networking, storage, and security. It's excellent for on-premises deployments where you manage the infrastructure.
Managed Kubernetes services (EKS, GKE, AKS) handle control plane management for you. Consider these when you want to focus on applications rather than cluster operations, though you sacrifice some control over underlying infrastructure.
For this guide, we'll use kubeadm as it provides the clearest path to understanding Kubernetes internals while remaining production-capable.
Infrastructure Prerequisites
Before installing Kubernetes, prepare your infrastructure:
Server Requirements
- Control plane nodes: 2 CPU cores minimum, 4 GB RAM minimum, 20 GB disk
- Worker nodes: 2 CPU cores minimum, 4 GB RAM minimum, storage based on workload
- Operating system: Ubuntu 22.04 LTS, Rocky Linux 9, or similar supported distributions
- Network connectivity: All nodes must reach each other, with specific ports open
Network Planning
- Pod network CIDR (default 10.244.0.0/16, adjust if conflicts exist)
- Service CIDR (default 10.96.0.0/12)
- Load balancer IP for control plane API (for multi-master setups)
- DNS resolution between all nodes
Required Ports
Control plane nodes need these ports open:
- 6443: Kubernetes API server
- 2379-2380: etcd client and peer communication
- 10250: Kubelet API
- 10259: kube-scheduler
- 10257: kube-controller-manager
Worker nodes need:
- 10250: Kubelet API
- 30000-32767: NodePort services (if used)
Installing Container Runtime
Kubernetes requires a container runtime. containerd has become the standard choice following Docker's deprecation as a runtime.
On each node, install containerd:
# Load required kernel modules
cat <<EOF | sudo tee /etc/modules-load.d/containerd.conf
overlay
br_netfilter
EOF
sudo modprobe overlay
sudo modprobe br_netfilter
# Set required sysctl parameters
cat <<EOF | sudo tee /etc/sysctl.d/99-kubernetes-cri.conf
net.bridge.bridge-nf-call-iptables = 1
net.ipv4.ip_forward = 1
net.bridge.bridge-nf-call-ip6tables = 1
EOF
sudo sysctl --system
# Install containerd
sudo apt-get update
sudo apt-get install -y containerd
# Configure containerd
sudo mkdir -p /etc/containerd
containerd config default | sudo tee /etc/containerd/config.toml
# Enable systemd cgroup driver
sudo sed -i 's/SystemdCgroup = false/SystemdCgroup = true/' /etc/containerd/config.toml
sudo systemctl restart containerd
sudo systemctl enable containerd
The systemd cgroup driver configuration ensures compatibility with kubeadm and modern init systems.
Installing Kubernetes Components
Install kubelet, kubeadm, and kubectl on all nodes:
# Add Kubernetes repository
sudo apt-get update
sudo apt-get install -y apt-transport-https ca-certificates curl gpg
curl -fsSL https://pkgs.k8s.io/core:/stable:/v1.30/deb/Release.key | sudo gpg --dearmor -o /etc/apt/keyrings/kubernetes-apt-keyring.gpg
echo 'deb [signed-by=/etc/apt/keyrings/kubernetes-apt-keyring.gpg] https://pkgs.k8s.io/core:/stable:/v1.30/deb/ /' | sudo tee /etc/apt/sources.list.d/kubernetes.list
# Install Kubernetes packages
sudo apt-get update
sudo apt-get install -y kubelet kubeadm kubectl
sudo apt-mark hold kubelet kubeadm kubectl
# Disable swap (Kubernetes requirement)
sudo swapoff -a
sudo sed -i '/ swap / s/^/#/' /etc/fstab
Holding the package versions prevents accidental upgrades that could break cluster compatibility.
Initializing the Control Plane
On your first control plane node, initialize the cluster:
sudo kubeadm init \
--pod-network-cidr=10.244.0.0/16 \
--control-plane-endpoint=loadbalancer.example.com:6443 \
--upload-certs
Replace loadbalancer.example.com with your load balancer DNS name or the first control plane node's IP if building a single-master cluster initially.
The command outputs a kubeadm join command for both control plane and worker nodes. Save these commands securely.
Configure kubectl access:
mkdir -p $HOME/.kube
sudo cp -i /etc/kubernetes/admin.conf $HOME/.kube/config
sudo chown $(id -u):$(id -g) $HOME/.kube/config
Installing a Pod Network
Kubernetes requires a Container Network Interface (CNI) plugin for pod-to-pod communication. Cilium, Calico, and Flannel are popular choices.
For Calico (robust security policies and good performance):
kubectl create -f https://raw.githubusercontent.com/projectcalico/calico/v3.27.0/manifests/tigera-operator.yaml
kubectl create -f https://raw.githubusercontent.com/projectcalico/calico/v3.27.0/manifests/custom-resources.yaml
Wait for all pods to reach Running state:
kubectl get pods -n calico-system -w
Joining Additional Nodes
For high availability, add two more control plane nodes using the join command from initialization:
sudo kubeadm join loadbalancer.example.com:6443 \
--token <token> \
--discovery-token-ca-cert-hash sha256:<hash> \
--control-plane \
--certificate-key <certificate-key>
Join worker nodes with the worker-specific command:
sudo kubeadm join loadbalancer.example.com:6443 \
--token <token> \
--discovery-token-ca-cert-hash sha256:<hash>
Verify all nodes joined successfully:
kubectl get nodes
Security Hardening
RBAC Configuration
Kubernetes uses Role-Based Access Control by default. Never use the cluster-admin ClusterRole for application service accounts. Create least-privilege roles:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
namespace: production
name: app-reader
rules:
- apiGroups: [""]
resources: ["pods", "services"]
verbs: ["get", "list", "watch"]
Pod Security Standards
Enable Pod Security Standards at the namespace level to restrict dangerous pod configurations:
kubectl label namespace production \
pod-security.kubernetes.io/enforce=restricted \
pod-security.kubernetes.io/audit=restricted \
pod-security.kubernetes.io/warn=restricted
This prevents privilege escalation, host namespace access, and dangerous volume types.
Network Policies
Implement default-deny network policies and explicitly allow required traffic:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: production
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
API Server Audit Logging
Enable audit logging to track API access. Create an audit policy:
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
- level: Metadata
omitStages:
- RequestReceived
Update kube-apiserver configuration to use this policy by editing /etc/kubernetes/manifests/kube-apiserver.yaml.
Secrets Encryption at Rest
Encrypt secrets in etcd. Create an encryption configuration:
apiVersion: apiserver.config.k8s.io/v1
kind: EncryptionConfiguration
resources:
- resources:
- secrets
providers:
- aescbc:
keys:
- name: key1
secret: <base64-encoded-32-byte-key>
- identity: {}
Reference this file in kube-apiserver startup arguments.
Storage Configuration
Production clusters need persistent storage. Configure a storage class based on your infrastructure:
For local SSDs (development and specific workloads):
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: local-ssd
provisioner: kubernetes.io/no-provisioner
volumeBindingMode: WaitForFirstConsumer
For production, use dynamic provisioners like Longhorn (on-premises), or CSI drivers for cloud block storage. Install Longhorn:
kubectl apply -f https://raw.githubusercontent.com/longhorn/longhorn/v1.6.0/deploy/longhorn.yaml
This provides replicated, distributed block storage across your cluster nodes.
Ingress Controller Setup
Deploy an ingress controller to expose services externally. NGINX Ingress Controller is widely used:
kubectl apply -f https://raw.githubusercontent.com/kubernetes/ingress-nginx/controller-v1.10.0/deploy/static/provider/baremetal/deploy.yaml
For LoadBalancer service type support on bare metal, install MetalLB:
kubectl apply -f https://raw.githubusercontent.com/metallb/metallb/v0.14.0/config/manifests/metallb-native.yaml
Configure MetalLB with your available IP range:
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
name: default-pool
namespace: metallb-system
spec:
addresses:
- 192.168.1.240-192.168.1.250
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
name: default
namespace: metallb-system
Monitoring and Observability
Deploy Prometheus and Grafana for cluster monitoring:
kubectl create namespace monitoring
kubectl apply -f https://raw.githubusercontent.com/prometheus-operator/kube-prometheus/main/manifests/setup/ -n monitoring
kubectl apply -f https://raw.githubusercontent.com/prometheus-operator/kube-prometheus/main/manifests/ -n monitoring
Access Grafana through port-forwarding initially:
kubectl port-forward -n monitoring svc/grafana 3000:3000
Configure ingress for external access once your ingress controller is running.
Backup Strategy
Back up etcd regularly as it contains all cluster state:
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-snapshot.db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
Automate this with a CronJob and store backups off-cluster.
For application-level backups, consider Velero, which backs up Kubernetes resources and persistent volumes.
Cluster Upgrades
Plan upgrades carefully. Kubernetes supports version skew of one minor version between control plane and kubelet. Upgrade control plane nodes first, then workers:
# On first control plane node
sudo apt-mark unhold kubeadm
sudo apt-get update && sudo apt-get install -y kubeadm=1.30.x-*
sudo apt-mark hold kubeadm
sudo kubeadm upgrade plan
sudo kubeadm upgrade apply v1.30.x
# Upgrade kubelet and kubectl
sudo apt-mark unhold kubelet kubectl
sudo apt-get update && sudo apt-get install -y kubelet=1.30.x-* kubectl=1.30.x-*
sudo apt-mark hold kubelet kubectl
sudo systemctl daemon-reload
sudo systemctl restart kubelet
Repeat for remaining control plane nodes, then workers. Drain nodes before upgrading to move workloads:
kubectl drain <node-name> --ignore-daemonsets
Uncordon after upgrade:
kubectl uncordon <node-name>
Production Checklist
Before going live, verify:
- [ ] High availability: Three or more control plane nodes
- [ ] etcd backups automated and tested
- [ ] Resource limits set on all workloads
- [ ] Pod Security Standards enforced
- [ ] Network policies implemented
- [ ] Secrets encrypted at rest
- [ ] RBAC roles follow least privilege
- [ ] Monitoring and alerting configured
- [ ] Log aggregation in place
- [ ] Disaster recovery plan documented and tested
- [ ] Node auto-scaling configured (if on cloud)
- [ ] Certificate expiration monitoring enabled
Conclusion
Deploying production Kubernetes requires attention to architecture, security, and operational tooling. Start with a solid foundation using kubeadm, harden security through RBAC and network policies, implement proper storage and networking, and establish monitoring before running production workloads. This systematic approach creates a maintainable cluster that scales with your needs while remaining secure and observable. Remember that Kubernetes is complex by nature—invest time in understanding each component rather than rushing to production.
