When a critical CVE drops, the pressure is immediate. Your production servers need patching, but you cannot afford downtime or a broken service. The key is a repeatable process that balances speed with safety: triage fast, test thoroughly, deploy confidently.
This playbook walks you through the operational steps for managing critical patches from the moment a vulnerability is announced to the moment your production fleet is secured.
Understanding CVE Severity and Impact
Not every CVE requires an emergency response. Start by understanding what you are dealing with.
Check the CVSS score. Critical vulnerabilities typically score 9.0 or higher. High severity ranges from 7.0 to 8.9. These numbers give you a baseline, but context matters more.
Assess exploitability. Is there a public exploit? Is it being actively exploited in the wild? A critical CVE with proof-of-concept code in circulation demands faster action than one that is purely theoretical.
Identify affected components. Which services, packages, or libraries are vulnerable? Cross-reference your production stack. If the vulnerable component is not installed or is disabled, your urgency drops.
Determine exposure. Is the vulnerable service internet-facing or internal-only? Is it behind a firewall or WAF? Public-facing services require immediate attention.
Document your findings in a shared location. Your team needs visibility into what is at risk and why you are prioritizing certain patches.
Building Your Triage Workflow
A consistent triage process saves time and reduces mistakes.
Subscribe to security advisories. Follow your distribution's security mailing list (RHEL/CentOS, Debian, Ubuntu) and vendor advisories for software you run. Set up alerts for CVE databases if you need broader coverage.
Maintain an asset inventory. Know what is running where. Use configuration management tools or simple spreadsheets to track package versions, service configurations, and dependencies across your fleet.
Create a severity matrix. Define what critical, high, medium, and low mean for your organization. Map each severity level to a response timeline. For example, critical patches might require deployment within 24 hours, while high severity allows 72 hours.
Assign ownership. Designate who triages, who tests, and who approves production changes. Ambiguity slows response.
Testing Patches Before Production
Never apply a critical patch directly to production without testing. Even vendor-supplied patches can introduce regressions.
Set Up a Staging Environment
Your staging environment should mirror production as closely as possible: same OS version, same package versions, same service configurations. If your production runs WordPress with specific plugins, staging needs identical plugins.
For infrastructure managed with Ansible, Puppet, or Terraform, apply the same configuration to staging. This ensures you catch integration issues before they hit production.
Apply the Patch in Staging
For Debian and Ubuntu:
sudo apt update
sudo apt list --upgradable | grep <package-name>
sudo apt install --only-upgrade <package-name>
For RHEL, CentOS, and Fedora:
sudo yum check-update <package-name>
sudo yum update <package-name>
For individual services like OpenSSL, Apache, or Nginx, verify the installed version after patching:
openssl version
apache2 -v
nginx -v
Run Your Test Suite
Execute automated tests if you have them. If not, run manual smoke tests:
- Verify service startup and restart behavior
- Check log files for errors or warnings
- Test key user flows: authentication, API calls, database queries
- Monitor resource usage for unexpected spikes
- Validate SSL/TLS connections if patching cryptographic libraries
Document any issues. If the patch breaks functionality, decide whether to proceed with workarounds or wait for a vendor fix.
Deploying to Production Safely
Once staging validates the patch, plan your production rollout.
Schedule Maintenance Windows
If downtime is unavoidable, schedule it during low-traffic periods. Notify users in advance. For critical vulnerabilities, shorter notice is acceptable, but communicate clearly.
For 24/7 services, rolling deployments minimize risk.
Use Rolling Deployments
If you run multiple servers behind a load balancer, patch one node at a time.
- Remove the first server from the load balancer pool
- Apply the patch
- Restart the service
- Run health checks
- Return the server to the pool
- Repeat for remaining servers
This approach keeps your service online throughout the process. Monitor each step. If issues appear, stop the rollout and investigate before continuing.
Leverage Live Patching Where Available
Some Linux distributions offer live patching for kernel vulnerabilities, allowing you to apply patches without rebooting.
Ubuntu offers Livepatch:
sudo snap install canonical-livepatch
sudo canonical-livepatch enable <token>
RHEL provides kpatch, and SUSE offers kGraft. These tools are ideal for kernel CVEs but do not cover user-space applications.
Apply the Patch
Use the same commands you tested in staging. Log each step for audit purposes.
After patching, restart the affected services:
sudo systemctl restart apache2
sudo systemctl restart nginx
sudo systemctl restart mysql
Verify the service is running:
sudo systemctl status apache2
Validate Production Health
After deployment, monitor your systems closely:
- Check application logs for errors
- Monitor response times and error rates
- Verify SSL/TLS handshakes if certificates or cryptographic libraries were updated
- Test critical user flows
- Review system metrics (CPU, memory, disk I/O)
Keep your rollback plan ready. If issues appear, be prepared to revert.
Handling Rollbacks
Despite careful testing, production issues can still occur.
Before You Patch
Take snapshots of critical systems if your infrastructure supports it. Cloud providers and virtualization platforms make this straightforward.
For package-level rollbacks, know your options:
Debian/Ubuntu:
sudo apt install <package-name>=<previous-version>
RHEL/CentOS:
sudo yum downgrade <package-name>
During a Rollback
If the patch causes issues, revert immediately:
- Remove the patched server from rotation
- Downgrade to the previous package version
- Restart services
- Validate functionality
- Return the server to rotation only after confirming stability
Document what went wrong. Share findings with your team and the vendor if appropriate.
Automating Patch Management
Manual patching does not scale. Automation reduces human error and speeds response.
Use Configuration Management
Ansible, Puppet, Chef, and SaltStack can automate patch deployment across your fleet.
Example Ansible playbook for updating a specific package:
---
- name: Apply critical security patch
hosts: webservers
become: yes
tasks:
- name: Update vulnerable package
apt:
name: openssl
state: latest
update_cache: yes
when: ansible_os_family == "Debian"
- name: Restart affected service
systemd:
name: apache2
state: restarted
Run the playbook against staging first, validate, then deploy to production.
Enable Unattended Upgrades for Low-Risk Patches
For non-critical security updates, consider automatic patching.
On Debian/Ubuntu, configure unattended-upgrades:
sudo apt install unattended-upgrades
sudo dpkg-reconfigure --priority=low unattended-upgrades
Edit /etc/apt/apt.conf.d/50unattended-upgrades to control which updates install automatically. Restrict automatic updates to security patches only, and exclude mission-critical services that require manual oversight.
Monitor Patch Status
Use tools to track which servers have been patched and which are still vulnerable. Configuration management platforms provide reporting. For smaller environments, a simple script that checks package versions and outputs to a dashboard works.
Coordinating with Your Team
Patching production is a team effort.
Establish a communication channel. Use Slack, Microsoft Teams, or a dedicated incident channel for real-time coordination during critical patching.
Document your process. Maintain a runbook with commands, rollback steps, and contact information. When an urgent CVE appears, you want a reference guide, not uncertainty.
Conduct post-patch reviews. After deploying a critical patch, gather your team for a brief retrospective. What went well? What slowed you down? Update your playbook based on lessons learned.
Train your team. Ensure multiple people know how to triage, test, and deploy patches. Relying on one person creates a bottleneck and a single point of failure.
Special Considerations for Specific Services
Web Servers (Apache, Nginx)
Test configuration syntax after patching:
sudo nginx -t
sudo apachectl configtest
If syntax checks pass, reload (graceful restart) instead of a hard restart when possible:
sudo systemctl reload nginx
sudo systemctl reload apache2
Database Servers (MySQL, PostgreSQL)
Database patches may require downtime. Test backup and restore procedures before patching. Verify replication health after patching replicas.
For MySQL:
sudo systemctl stop mysql
sudo apt install --only-upgrade mysql-server
sudo systemctl start mysql
mysql_upgrade -u root -p
Control Panels (cPanel, Plesk)
Control panels handle patching through their own interfaces. Use their built-in update mechanisms rather than bypassing them with manual package updates. Breaking a control panel with an unsupported patch creates more problems than it solves.
Conclusion
Managing critical CVE patches is about process, not panic. Triage based on real risk, test in staging, deploy methodically, and monitor closely. Build automation where it makes sense, but keep human oversight where it matters.
The servers you patch today stay secure. The process you build today makes the next critical CVE manageable instead of chaotic. Document your playbook, train your team, and refine your workflow after every incident. Your production environment will be more resilient for it.
