Every week I see tickets where clusters deployed straight from quickstart tutorials fall over under real load or get compromised because nobody locked down the control plane. This walkthrough covers the five configuration layers that separate a working cluster from one that runs in production.
We'll provision three Ubuntu 22.04 VMs—one control plane node, two workers—and wire up containerd, Calico networking, pod security admission, resource quotas, and horizontal pod autoscaling. By the end you'll have a cluster that enforces baseline security posture and scales workloads automatically.
Before you provision nodes
Each VM needs at least 2 CPU cores and 2 GB RAM for the control plane, 1 core and 1 GB for workers. Disk: 20 GB minimum. Disable swap on every node because kubelet refuses to start with swap enabled.
sudo swapoff -a
sudo sed -i '/ swap / s/^/#/' /etc/fstab
Open these ports in your firewall: 6443 (API server), 2379-2380 (etcd), 10250-10252 (kubelet, scheduler, controller-manager) on the control plane. Workers need 10250 and 30000-32767 for NodePort services. If you're behind a cloud provider security group, create rules for the entire pod CIDR and service CIDR you'll define in step two.
Install a container runtime on all nodes. Kubernetes dropped direct Docker support in 1.24, so we're using containerd.
sudo apt update
sudo apt install -y containerd
sudo mkdir -p /etc/containerd
containerd config default | sudo tee /etc/containerd/config.toml
Edit /etc/containerd/config.toml and set SystemdCgroup = true under [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runc.options]. Restart containerd.
sudo systemctl restart containerd
sudo systemctl enable containerd
Load the kernel modules Kubernetes needs and set sysctl parameters.
cat <<EOF | sudo tee /etc/modules-load.d/k8s.conf
overlay
br_netfilter
EOF
sudo modprobe overlay
sudo modprobe br_netfilter
cat <<EOF | sudo tee /etc/sysctl.d/k8s.conf
net.bridge.bridge-nf-call-iptables = 1
net.bridge.bridge-nf-call-ip6tables = 1
net.ipv4.ip_forward = 1
EOF
sudo sysctl --system
Step 1: Install kubeadm and initialize the control plane
Add the Kubernetes package repository and install kubeadm, kubelet, and kubectl on all three nodes.
sudo apt install -y apt-transport-https ca-certificates curl gpg
curl -fsSL https://pkgs.k8s.io/core:/stable:/v1.30/deb/Release.key | sudo gpg --dearmor -o /etc/apt/keyrings/kubernetes-apt-keyring.gpg
echo 'deb [signed-by=/etc/apt/keyrings/kubernetes-apt-keyring.gpg] https://pkgs.k8s.io/core:/stable:/v1.30/deb/ /' | sudo tee /etc/apt/sources.list.d/kubernetes.list
sudo apt update
sudo apt install -y kubelet kubeadm kubectl
sudo apt-mark hold kubelet kubeadm kubectl
On the control plane node, run kubeadm init with an explicit pod network CIDR. We're using 192.168.0.0/16 because that's Calico's default; if your node IPs overlap that range, pick a different CIDR.
sudo kubeadm init --pod-network-cidr=192.168.0.0/16 --control-plane-endpoint=<CONTROL_PLANE_IP>:6443
Replace <CONTROL_PLANE_IP> with the actual private IP of your control plane node. The init command prints a join token at the end. Copy the entire kubeadm join line; you'll run it on the worker nodes in a minute.
Set up kubectl access for your regular user.
mkdir -p $HOME/.kube
sudo cp -i /etc/kubernetes/admin.conf $HOME/.kube/config
sudo chown $(id -u):$(id -g) $HOME/.kube/config
Verify the control plane components are running.
kubectl get pods -n kube-system
You'll see coredns pods stuck in Pending state. That's expected until we install a CNI.
Step 2: Deploy Calico for pod networking
Kubernetes doesn't ship with a network plugin. Without one, pods can't talk to each other or reach the internet. Calico is a good default—it handles pod routing and gives you network policy enforcement later.
Download the Calico manifest and apply it.
kubectl create -f https://raw.githubusercontent.com/projectcalico/calico/v3.28.0/manifests/tigera-operator.yaml
kubectl create -f https://raw.githubusercontent.com/projectcalico/calico/v3.28.0/manifests/custom-resources.yaml
Watch the calico-system and calico-apiserver namespaces until all pods reach Running.
kubectl get pods -n calico-system -w
Once Calico is up, coredns pods will start. Check that all kube-system pods are healthy.
kubectl get pods -n kube-system
Now join the worker nodes. SSH into each worker and paste the kubeadm join command you copied earlier. It looks like this:
sudo kubeadm join <CONTROL_PLANE_IP>:6443 --token <TOKEN> --discovery-token-ca-cert-hash sha256:<HASH>
Back on the control plane, verify all nodes are Ready.
kubectl get nodes
If a node stays NotReady for more than two minutes, check kubelet logs with journalctl -u kubelet -f.
Step 3: Enable pod security admission
Kubernetes replaced PodSecurityPolicy with Pod Security Standards in 1.25. The new model has three levels: privileged (unrestricted), baseline (blocks known privilege escalations), and restricted (hardened, drop all capabilities).
We'll enforce baseline cluster-wide and restricted in a dedicated namespace for untrusted workloads. Create a namespace with restricted enforcement.
kubectl create namespace restricted-apps
kubectl label namespace restricted-apps \
pod-security.kubernetes.io/enforce=restricted \
pod-security.kubernetes.io/audit=restricted \
pod-security.kubernetes.io/warn=restricted
Any pod you deploy into restricted-apps must run as non-root, drop all capabilities, and use a read-only root filesystem. If you try to deploy a pod that violates restricted, the API server blocks it.
Test with a baseline-compliant nginx pod.
apiVersion: v1
kind: Pod
metadata:
name: nginx-baseline
namespace: restricted-apps
spec:
securityContext:
runAsNonRoot: true
runAsUser: 1000
seccompProfile:
type: RuntimeDefault
containers:
- name: nginx
image: nginx:1.25
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop:
- ALL
readOnlyRootFilesystem: true
volumeMounts:
- name: cache
mountPath: /var/cache/nginx
- name: run
mountPath: /var/run
volumes:
- name: cache
emptyDir: {}
- name: run
emptyDir: {}
Apply it and confirm it starts.
kubectl apply -f nginx-baseline.yaml
kubectl get pod nginx-baseline -n restricted-apps
For the default namespace, enforce baseline and warn on restricted violations.
kubectl label namespace default \
pod-security.kubernetes.io/enforce=baseline \
pod-security.kubernetes.io/warn=restricted
Now privileged pods are blocked in default, but you still get a warning if a pod violates restricted rules without blocking deployment.
Step 4: Configure resource quotas and limit ranges
Without quotas, a single buggy deployment can consume all cluster CPU. Create a ResourceQuota in the default namespace.
apiVersion: v1
kind: ResourceQuota
metadata:
name: default-quota
namespace: default
spec:
hard:
requests.cpu: "4"
requests.memory: 8Gi
limits.cpu: "8"
limits.memory: 16Gi
pods: "20"
Apply it.
kubectl apply -f resource-quota.yaml
Now the total of all pod requests in default can't exceed 4 CPU and 8 GB. Any pod without resource requests will be rejected, so set a LimitRange to provide defaults.
apiVersion: v1
kind: LimitRange
metadata:
name: default-limits
namespace: default
spec:
limits:
- max:
cpu: "2"
memory: 2Gi
min:
cpu: 100m
memory: 128Mi
default:
cpu: 500m
memory: 512Mi
defaultRequest:
cpu: 250m
memory: 256Mi
type: Container
Apply the LimitRange.
kubectl apply -f limit-range.yaml
New pods without explicit requests get 250m CPU and 256 MB memory requests, 500m CPU and 512 MB limits. Test by deploying a pod with no resources defined; check its effective requests with kubectl describe pod.
Step 5: Set up horizontal pod autoscaling
The Horizontal Pod Autoscaler scales replica count based on CPU or memory utilization. It needs the metrics-server to read resource usage.
Install metrics-server.
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
If you're using self-signed certificates on your nodes, the metrics-server will fail TLS verification. Edit the deployment and add --kubelet-insecure-tls to the container args.
kubectl edit deployment metrics-server -n kube-system
Add this line under spec.template.spec.containers[0].args:
- --kubelet-insecure-tls
Wait for metrics-server to start, then check that node and pod metrics are available.
kubectl top nodes
kubectl top pods -A
Deploy a sample app with resource requests.
apiVersion: apps/v1
kind: Deployment
metadata:
name: php-apache
namespace: default
spec:
replicas: 1
selector:
matchLabels:
app: php-apache
template:
metadata:
labels:
app: php-apache
spec:
containers:
- name: php-apache
image: registry.k8s.io/hpa-example
ports:
- containerPort: 80
resources:
requests:
cpu: 200m
limits:
cpu: 500m
---
apiVersion: v1
kind: Service
metadata:
name: php-apache
namespace: default
spec:
ports:
- port: 80
selector:
app: php-apache
Apply it.
kubectl apply -f php-apache.yaml
Create a HorizontalPodAutoscaler that scales from 1 to 10 replicas when CPU crosses 50%.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: php-apache-hpa
namespace: default
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: php-apache
minReplicas: 1
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50
Apply the HPA.
kubectl apply -f hpa.yaml
Generate load to trigger scaling. Open a second terminal and run a load generator pod.
kubectl run -i --tty load-generator --rm --image=busybox --restart=Never -- /bin/sh -c "while sleep 0.01; do wget -q -O- http://php-apache; done"
In your first terminal, watch the HPA.
kubectl get hpa php-apache-hpa --watch
Within a minute or two, CPU utilization will climb past 50% and the HPA will increase replicas. Stop the load generator with Ctrl+C and watch replicas scale back down after five minutes of low utilization.
What about network policies?
Calico installed the network policy controller, but by default all pods can talk to all pods. Lock down traffic by creating a default-deny policy in each namespace, then allow only necessary connections.
Deny all ingress in the default namespace.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-ingress
namespace: default
spec:
podSelector: {}
policyTypes:
- Ingress
Apply it.
kubectl apply -f deny-ingress.yaml
Now no external traffic reaches pods in default unless you explicitly allow it. For the php-apache service, allow ingress from a specific namespace or pod label.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-php-apache
namespace: default
spec:
podSelector:
matchLabels:
app: php-apache
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
role: frontend
ports:
- protocol: TCP
port: 80
This allows ingress only from pods labeled role: frontend on port 80. Repeat for every service that needs inbound traffic.
Keeping the cluster patched
Kubernetes releases a new minor version every four months. Plan to upgrade at least twice a year. The upgrade path is control plane first, then workers one at a time.
Before upgrading, check the release notes for API deprecations. Drain each worker before upgrading its kubelet so pods reschedule to healthy nodes.
kubectl drain <NODE_NAME> --ignore-daemonsets --delete-emptydir-data
sudo apt update && sudo apt install -y kubelet kubeadm kubectl
sudo systemctl restart kubelet
kubectl uncordon <NODE_NAME>
Test the upgrade process in a non-production cluster first. I've seen custom controllers break after a minor version bump because they relied on a beta API that got removed.
Do I need a service mesh?
Not unless you're running dozens of microservices that need mutual TLS between every pod. Network policies handle the basics. Add a mesh like Istio or Linkerd when you need traffic splitting, retries, or observability at the service layer.
How many control plane nodes do I need?
Three for high availability. Run an odd number so etcd can maintain quorum if one node fails. A single control plane is fine for dev and staging environments.
Can I run the control plane on a worker?
Yes, remove the NoSchedule taint from the control plane node with kubectl taint nodes <NODE_NAME> node-role.kubernetes.io/control-plane:NoSchedule-. In production, keep the control plane isolated so workload churn doesn't starve etcd or the API server.
What if metrics-server still won't start?
Check that DNS is working inside the cluster. Create a test pod with kubectl run busybox --image=busybox --restart=Never -- sleep 3600, exec into it with kubectl exec -it busybox -- sh, and try to resolve kubernetes.default.svc.cluster.local with nslookup. If DNS fails, check coredns logs.
What actually makes this production-ready
The difference between a demo cluster and a production one is enforcement. Pod security admission blocks privilege escalation. Resource quotas prevent runaway pods from starving the cluster. Network policies stop lateral movement. Autoscaling handles traffic spikes without manual intervention.
You still need monitoring, backup, and a disaster recovery plan, but this five-step setup gives you a cluster that won't fold the first time someone deploys a memory leak or a pod requests root privileges. Run through these steps on a staging environment, break things on purpose, and confirm your guardrails actually work before you point production traffic at it.
