When a critical CVE drops, you need to act fast—but rushing without a plan breaks production systems. This playbook walks you through triaging severity, testing patches in isolation, deploying safely, and documenting every step so your next critical response is faster and smoother.
The First 15 Minutes: Assess and Contain
Your first job is to understand what you're dealing with and whether you need to take immediate defensive action before patching.
Verify the CVE Applies to Your Stack
Not every critical CVE affects your systems. Check:
- Affected software and versions: Does the vulnerability target software you run? Confirm installed versions.
- Attack prerequisites: Does exploitation require authentication, local access, or specific configurations you don't use?
- Exposure surface: Is the vulnerable service public-facing, internal-only, or unused?
Quick version checks:
# Package versions on Debian/Ubuntu
dpkg -l | grep package-name
apt-cache policy package-name
# Package versions on RHEL/CentOS/AlmaLinux
rpm -qa | grep package-name
yum info package-name
# Application versions
nginx -v
apache2 -v
php -v
mysql --version
If the vulnerable version isn't installed or the service isn't exposed, document that fact and monitor for updates. You can deprioritize the response.
Evaluate True Severity
CVSS scores provide a starting point, but context determines your actual risk:
- Exploitability: Is there public exploit code? Active scanning? Proof-of-concept attacks in the wild?
- Impact: Does successful exploitation mean remote code execution, data exfiltration, privilege escalation, or denial of service?
- Compensating controls: Do firewalls, access control lists, WAFs, or authentication layers reduce exposure?
A remote code execution vulnerability in an internet-facing web server is more urgent than a local privilege escalation bug on a single-tenant VPS with no shell users.
Immediate Containment Options
If the CVE is actively exploited and patching will take hours, consider temporary mitigations:
- Firewall rules: Block external access to the vulnerable service.
- Disable features: Turn off unused modules or endpoints.
- WAF rules: Deploy signatures to block known exploit patterns.
- Access restrictions: Require VPN or IP whitelisting.
Example firewall rule to restrict access:
# Block external access to a service, allow from trusted IPs only
iptables -A INPUT -p tcp --dport 8080 -s 203.0.113.0/24 -j ACCEPT
iptables -A INPUT -p tcp --dport 8080 -j DROP
# Save rules
netfilter-persistent save # Debian/Ubuntu
service iptables save # RHEL/CentOS
Containment buys time but is not a substitute for patching.
Build a Patch Testing Pipeline
Never apply patches directly to production. Test in an environment that mirrors production configuration, load, and dependencies.
Set Up an Isolated Test Environment
Your test environment should match production:
- Same OS version and patch level
- Same application versions and configurations
- Representative workload and traffic patterns
- Dependent services (databases, caches, APIs)
Options for快速 test environments:
- Staging servers: Pre-existing mirrors of production.
- Containers: Spin up Docker or LXC containers with production images.
- Virtual machines: Clone production VMs or use snapshots.
- Cloud instances: Launch temporary instances from production AMIs or images.
Snapshot production before building your test:
# LVM snapshot for quick rollback
lvcreate -L 10G -s -n prod-snapshot /dev/vg0/prod-lv
# VirtualBox snapshot
VBoxManage snapshot "ProdVM" take "pre-patch-snapshot"
# AWS EC2 snapshot
aws ec2 create-snapshot --volume-id vol-xxxxx --description "Pre-patch backup"
Obtain and Verify Patches
Use official distribution repositories or vendor channels. Avoid third-party patch sources.
# Update package lists
apt update # Debian/Ubuntu
yum check-update # RHEL/CentOS
# View available updates for specific package
apt-cache policy openssl
yum list updates openssl
# Download without installing (for inspection)
apt download openssl
yumdownloader openssl
Verify package signatures:
# Debian/Ubuntu
dpkg-sig --verify package.deb
# RHEL/CentOS
rpm --checksig package.rpm
Test Patch Installation
Apply the patch in your test environment and verify:
- Installation succeeds: No dependency conflicts or errors.
- Services restart cleanly: All daemons come back up.
- Application functions: Core workflows, APIs, and user-facing features work.
- Performance baseline: Response times and resource usage stay normal.
- Logs are clean: No new errors or warnings.
Test installation:
# Debian/Ubuntu
apt install --only-upgrade openssl
systemctl restart nginx
systemctl status nginx
# RHEL/CentOS
yum update openssl
systemctl restart httpd
systemctl status httpd
Run Automated Tests
If you have automated test suites, run them after patching:
- Unit tests
- Integration tests
- End-to-end tests
- Load tests
For web applications, quick smoke tests:
# Check HTTP response codes
curl -I https://testapp.example.com
# Test API endpoint
curl -X POST https://testapp.example.com/api/health \
-H "Content-Type: application/json" \
-d '{"check": "status"}'
# Load test with Apache Bench
ab -n 1000 -c 10 https://testapp.example.com/
Document any failures or anomalies. If tests fail, investigate before proceeding.
Deploy to Production Safely
Once testing confirms the patch is safe, deploy to production with a rollback plan and monitoring in place.
Pre-Deployment Checklist
Before touching production:
- [ ] Snapshot or back up all affected systems
- [ ] Notify your team and stakeholders of the maintenance window
- [ ] Prepare rollback procedures
- [ ] Verify monitoring and alerting are active
- [ ] Have communication channels open (Slack, PagerDuty, etc.)
- [ ] Schedule during low-traffic periods if possible
Staged Rollout Strategy
Deploy in phases to limit blast radius:
- Single canary server: Patch one production server, monitor for 15-30 minutes.
- Small cohort: Patch 10-20% of your fleet, monitor for anomalies.
- Full deployment: Patch remaining servers in batches.
For multi-server environments:
# Remove server from load balancer before patching
# (example with HAProxy via socket)
echo "disable server backend/server1" | socat stdio /var/run/haproxy.sock
# Patch the server
apt install --only-upgrade package-name
systemctl restart service-name
# Verify service health
curl -f http://localhost/health || echo "Health check failed"
# Return server to load balancer
echo "enable server backend/server1" | socat stdio /var/run/haproxy.sock
Minimize Downtime
For critical services, use zero-downtime strategies:
- Rolling restarts: Restart services one node at a time behind a load balancer.
- Blue-green deployment: Patch a parallel environment, then switch traffic.
- Live patching: Use kpatch or kGraft for kernel patches without reboots (when supported).
Example rolling restart with health checks:
for server in server1 server2 server3; do
echo "Patching $server"
ssh $server "apt install --only-upgrade openssl && systemctl restart nginx"
# Wait for service to be healthy
for i in {1..30}; do
if curl -sf http://$server/health > /dev/null; then
echo "$server is healthy"
break
fi
sleep 2
done
# Brief pause before next server
sleep 10
done
Monitor Post-Deployment
Watch these metrics closely for the first hour:
- Error rates: Application logs, HTTP 5xx responses
- Response times: API latency, page load times
- Resource usage: CPU, memory, disk I/O
- Service health: Daemon status, port listeners
- Security logs: Failed authentication, exploit attempts
Set up temporary alerts for anomalies:
# Watch error log for new issues
tail -f /var/log/nginx/error.log | grep -i "error\|crit\|alert"
# Monitor system load
watch -n 5 'uptime && free -m'
# Check for segfaults or crashes
journalctl -f -p err
If errors spike or services degrade, roll back immediately.
Rollback Procedures
Have a tested rollback plan before you start patching.
Package Rollback
Revert to the previous package version:
# Debian/Ubuntu - reinstall specific version
apt install package-name=previous-version
# RHEL/CentOS - downgrade package
yum downgrade package-name
# View available versions
apt-cache showpkg package-name
yum list package-name --showduplicates
Snapshot Rollback
Restore from pre-patch snapshots:
# LVM snapshot restore
lvconvert --merge /dev/vg0/prod-snapshot
# (requires reboot to take effect)
# VM snapshot restore
VBoxManage snapshot "ProdVM" restore "pre-patch-snapshot"
# Container rollback
docker stop container-name
docker run --name container-name previous-image:tag
Configuration Rollback
If the patch required config changes, revert them:
# Restore from backup
cp /etc/nginx/nginx.conf.backup /etc/nginx/nginx.conf
systemctl restart nginx
# Git-tracked configs
cd /etc/nginx
git checkout HEAD~1 nginx.conf
systemctl restart nginx
Test rollback procedures in staging before you need them in production.
Document Everything
Post-incident documentation makes future responses faster and trains your team.
Create a Response Timeline
Record key events:
- CVE disclosure time
- When you learned about it
- Triage and assessment completion
- Patch testing start and finish
- Production deployment start and finish
- Any issues encountered and how they were resolved
Write a Runbook
Capture the steps you took:
- Affected systems and versions
- Patch source and version
- Testing procedures
- Deployment commands
- Rollback steps
- Monitoring queries and thresholds
Store runbooks in your team wiki or documentation system. Tag them by CVE ID and software component.
Conduct a Brief Retrospective
Within a day or two, review:
- What went well?
- What took longer than expected?
- What would speed up the next response?
- What automation or tooling would help?
Use insights to improve your patch management process and runbooks.
Automate What You Can
Repeatable tasks should be scripted or automated to reduce response time and human error.
Automated Patch Testing
Build CI/CD pipelines that:
- Spin up test environments
- Apply patches
- Run test suites
- Report results
Example with simple shell script:
#!/bin/bash
# quick-patch-test.sh
TEST_VM="test-server-01"
PACKAGE="$1"
ssh $TEST_VM "apt update && apt install --only-upgrade -y $PACKAGE"
ssh $TEST_VM "systemctl restart nginx && systemctl is-active nginx"
ssh $TEST_VM "curl -f http://localhost/health"
if [ $? -eq 0 ]; then
echo "Patch test passed for $PACKAGE"
else
echo "Patch test failed for $PACKAGE"
exit 1
fi
Monitoring and Alerting
Set up alerts for:
- New CVE announcements for your software stack
- Available security updates in your repos
- Failed patch installations
- Service health degradation post-patch
Tools like Nagios, Prometheus, Zabbix, or Datadog can watch package versions and service status.
Configuration Management
Use Ansible, Puppet, Chef, or Salt to:
- Maintain consistent configurations across servers
- Deploy patches to fleets
- Roll back changes programmatically
Example Ansible playbook:
---
- name: Apply security patch
hosts: webservers
become: yes
serial: 1
tasks:
- name: Update package
apt:
name: openssl
state: latest
update_cache: yes
- name: Restart service
systemd:
name: nginx
state: restarted
- name: Wait for service to be ready
wait_for:
port: 80
delay: 5
Automation reduces the time from CVE disclosure to production patch.
Conclusion
Critical CVE response is a balance between speed and stability. Triage quickly to understand your exposure, test patches thoroughly in isolated environments, deploy in stages with rollback plans ready, and monitor closely. Document your process so each incident makes the next one smoother. With a solid playbook and the right tooling, you can patch fast without breaking production.
