When a critical CVE drops, you need a plan that moves fast but doesn't break production. I've seen teams scramble for hours because they never rehearsed the sequence. This guide walks you through triage, staging tests, and a rollout that keeps your sites live while closing the hole.
Triage in the first thirty minutes
You hear about a new CVE from a vendor advisory, a monitoring tool, or someone on your team. First question: does it affect software you're actually running?
Check the affected package and version range against your live inventory. If you manage a fleet, automate this with a configuration-management database or a simple script that queries rpm -qa, dpkg -l, or yum list installed across all boxes. For a handful of servers, SSH in and verify manually.
Not every "critical" rating deserves a four-hour sprint. Read the vendor's description and ask:
- Is the vulnerable service exposed to the internet or only localhost?
- Does exploitation require authentication?
- Is there a known public exploit in the wild?
If the service listens only on 127.0.0.1 and exploitation needs root, you have breathing room. If it's a remote code execution in a public-facing daemon with a Metasploit module already published, you patch now.
Document your decision in a shared chat channel or ticket. Write down the CVE identifier, affected packages, and which servers need the update. This clarity prevents duplicate work when three people all try to fix the same box.
Snapshot before you touch anything
Before applying any update, take a snapshot of each affected VM or create a backup of the running configuration. Cloud providers let you snapshot a disk in seconds. On bare metal, at minimum back up /etc and note the current kernel and package versions.
rpm -qa > /root/packages-before-patch.txt
uname -r > /root/kernel-before-patch.txt
tar -czf /root/etc-backup-$(date +%F).tar.gz /etc
If you're patching a production database server or a box with stateful services, coordinate a maintenance window or prepare a failover node. For stateless web servers behind a load balancer, you can pull one node out, patch it, test, then rotate through the pool.
Stage the update in a non-production clone
Never apply a patch to production first. Spin up a staging VM that mirrors your production OS version and installed software. If you don't have one, provision it now—it takes ten minutes and will save you from an outage.
Install the patched package on the staging box:
# RHEL/CentOS/AlmaLinux
sudo yum update <package-name> -y
# Debian/Ubuntu
sudo apt update && sudo apt install --only-upgrade <package-name> -y
Restart the affected service and watch the logs for errors:
sudo systemctl restart httpd
sudo journalctl -u httpd -f
Run a smoke test: hit the main application endpoint, check a few pages, confirm the service responds. If you have automated tests, run them now. If not, manually verify the top three workflows your users depend on.
Check for configuration-file conflicts. Package managers sometimes leave .rpmnew or .dpkg-dist files when a config has local changes. Compare them:
find /etc -name '*.rpmnew' -o -name '*.dpkg-dist'
diff /etc/httpd/conf/httpd.conf /etc/httpd/conf/httpd.conf.rpmnew
If the new config introduces breaking changes, merge carefully or stick with your current config if the CVE doesn't require config adjustments.
Roll out across production with a canary pattern
Once staging is stable, move to production in phases. Pull one server out of rotation, patch it, verify it, then put it back. Repeat for the next server.
For a load-balanced web cluster:
- Mark the first node as down in your load balancer (HAProxy, nginx upstream, or cloud LB).
- Apply the update and restart services.
- Curl the node directly (bypass the LB) to confirm it responds.
- Re-enable the node in the load balancer and watch error rates and response times for five minutes.
- If metrics look normal, proceed to the next node. If you see a spike in 500s or timeouts, roll back the canary and investigate.
For single-server setups, schedule a brief maintenance window during low-traffic hours. Apply the patch, restart, and monitor closely. Have the snapshot or backup ready in case you need to revert.
What if a reboot is required?
Kernel patches and certain library updates need a reboot to take effect. Check if needs-restarting (RHEL-based) or /var/run/reboot-required (Debian-based) signals a reboot.
# RHEL/CentOS
sudo needs-restarting -r
# Debian/Ubuntu
ls /var/run/reboot-required
If you must reboot, do it in the same canary pattern: one node at a time, verify it comes back healthy, then move on. For single-server environments, announce downtime, reboot, and be ready to roll back from the snapshot if the box doesn't come up.
Some teams use live-patching tools like kpatch or Ksplice to apply kernel fixes without rebooting. These are worth evaluating if your SLA forbids downtime, but test them in staging first—they don't cover every CVE.
Verify the patch closed the vulnerability
After rollout, confirm the CVE is actually fixed. Check the package version:
rpm -q <package-name>
dpkg -l | grep <package-name>
Compare the installed version to the vendor's advisory. If the advisory says "fixed in version 2.4.52" and you're running 2.4.51, the patch didn't apply—track down why.
Run a vulnerability scanner if you have one (OpenVAS, Nessus, or Qualys). Some CVEs have public proof-of-concept exploits; you can test defensively on your staging box to confirm the exploit now fails.
Document the patch in your change log: which servers were updated, what version was installed, and the timestamp. This record is critical during audits or when another CVE hits the same software.
When to skip the patch temporarily
Sometimes a patch introduces worse problems than the CVE. If the update breaks a business-critical feature and you've confirmed the vulnerability requires local access or authentication, consider a temporary workaround:
- Firewall the vulnerable service so only trusted IPs can reach it.
- Disable the affected feature if it's not in use.
- Apply strict input validation or rate limiting to reduce exploit likelihood.
Schedule the proper patch for the next maintenance window and document the compensating controls. Never skip a remote-code-execution fix without layered mitigations.
Communication checklist during the response
Keep your team and stakeholders informed at each stage:
- Triage complete: Post which CVE, which servers, and the plan.
- Staging tested: Share the test results and any issues found.
- Production rollout started: Announce which nodes are being patched.
- Rollout complete: Confirm all servers are patched and verified.
If you manage client servers, notify them before and after patching. A short email saying "we applied a critical security update to your server; no downtime occurred" builds trust.
Automate what you can for next time
After the fire drill, codify the process. Write a runbook that lists every step, every command, and every check. Next time a CVE drops, you'll execute faster.
Automate inventory checks with Ansible, Puppet, or a shell script that queries all servers and outputs which ones need patching. Automate snapshot creation in your cloud provider's API. Automate smoke tests with a curl script or a monitoring synthetic.
The less you have to remember under pressure, the fewer mistakes you'll make.
What you'll repeat every time
Critical CVE patching isn't a one-time event—it's a drill you'll run again. Triage fast, snapshot everything, test in staging, roll out carefully, and verify the fix. Document each step so the next engineer on call can follow the same sequence. Speed matters, but so does not breaking production. With this plan, you'll close the vulnerability before attackers get a foothold and keep your uptime promise intact.
