A critical CVE drops at 3 PM on a Friday. Your monitoring tools flag affected packages across twenty production servers. Management wants a timeline. Customers are asking questions. You need a repeatable process that gets the patch out fast without turning the weekend into a war room.
I've handled dozens of emergency patches—OpenSSL, kernel bugs, PHP zero-days. The teams that ship fixes in under 24 hours all follow the same basic pattern: triage hard, test once, stage carefully, deploy in waves, verify everything. No shortcuts, no cowboy moves, no surprises.
Stage One: Triage and Scope (First 2 Hours)
Your first job is figuring out what you're actually dealing with. Not every CVE with a scary CVSS score needs an emergency patch.
Start by reading the advisory. What's the attack vector—remote unauthenticated, local privilege escalation, or something that requires existing access? Does it affect your specific software version and configuration? A vulnerability in Apache's mod_http2 doesn't matter if you terminated TLS at the load balancer and disabled HTTP/2 months ago.
Inventory your exposure fast:
# Find package versions across your fleet
ansible all -m shell -a "dpkg -l | grep openssl"
# or for RPM systems
ansible all -m shell -a "rpm -qa | grep openssl"
Document which servers run the vulnerable version, what services depend on the affected package, and whether those services are customer-facing. A vulnerability in a background job processor is different from one in your edge web servers.
Check if a patch exists yet. Sometimes advisories drop before the vendor has released fixed packages. If there's no patch, your workflow pivots to mitigation—firewall rules, service isolation, or disabling features. That's a different playbook.
Set a target. For remote code execution or privilege escalation bugs with public exploits, I aim to have production patched within 24 hours. For lower-severity issues or vulnerabilities that need local access, you have more breathing room.
Stage Two: Test in Isolation (Hours 3-6)
Never apply a security patch directly to production, even under time pressure. I've seen well-intentioned patches break PHP applications, cause kernel panics, and kill database replication. Test first.
Spin up an isolated test environment that mirrors production as closely as possible—same OS version, same software stack, same configuration files. If you maintain staging servers, use those. Otherwise, clone a production server or launch a fresh instance.
Apply the patch:
# Debian/Ubuntu example
apt update
apt install --only-upgrade openssl libssl3
# RHEL/CentOS
yum update openssl
Run through your verification checklist:
- Does the service restart cleanly?
- Do dependent services still connect (PHP-FPM to Apache, application to database)?
- Do basic transactions work—page loads, API calls, user logins?
- Are there new warnings or errors in logs (
/var/log/syslog,/var/log/apache2/error.log, application logs)?
If you have automated tests, run them now. Load tests are even better—patches occasionally introduce performance regressions that only show up under load.
Document any issues you find and their fixes. When something breaks, don't just revert—figure out why. Maybe you need to restart dependent services in a specific order, or update a configuration file that referenced the old library path. Write down the steps.
Stage Three: Stage and Prepare (Hours 7-12)
Staging means getting everything ready so the actual deployment is fast and boring.
Update your change documentation. I keep a simple template:
CVE: CVE-YYYY-NNNNN
Affected: OpenSSL 3.0.0 - 3.0.7
Patch: openssl 3.0.8-1
Servers: web-01 through web-20, api-01 through api-05
Rollback: Snapshot IDs or package versions
Downtime: None expected (rolling restart)
Verification: curl https://example.com/health, check error logs
Owner: Your Name
Scheduled: 2024-03-15 22:00 UTC
Create rollback snapshots or backups. For cloud VMs, snapshot the root disk. For bare metal, at minimum back up the package database and configuration files. Rollback needs to be a single command you can execute half-asleep.
Prepare your deployment tooling. Whether you use Ansible, Puppet, Salt, or a bash script with SSH loops, test the playbook against your test environment one more time. Verify that it targets the correct servers and includes your post-patch verification steps.
Schedule a maintenance window if you need one. For patches that require service restarts but no downtime (rolling restarts behind a load balancer), you can often skip the window. For anything that requires taking services offline, communicate early.
Brief your team. Even if you're doing the work solo, tell someone else what you're doing and when. If something goes wrong at 2 AM, you want another engineer who can read your notes and understand the situation.
Stage Four: Deploy in Waves (Hours 13-20)
Never patch everything at once. Deploy in waves so you catch problems before they hit your entire infrastructure.
Start with a canary—one or two servers that handle real production traffic but aren't critical. If you have geographically distributed servers, pick one in a low-traffic region. Apply the patch, restart services, and watch for 30 minutes.
What are you watching? Error rates in your monitoring, response times, CPU and memory usage, and anything unusual in logs. If your canary looks stable, proceed to the next wave.
For a typical setup, I batch like this:
- Two canary servers (wait 30 min)
- 25% of remaining servers (wait 15 min)
- Another 50% (wait 15 min)
- Final 25%
Adjust based on your architecture. If you have active-active database replicas, patch secondaries first and verify replication before touching the primary. If you have layered services (load balancers, app servers, background workers), patch from the outside in.
Use rolling restarts to avoid downtime:
# Remove server from load balancer
curl -X POST https://lb.example.com/api/drain/web-03
# Wait for existing connections to finish
sleep 30
# Apply patch and restart
apt install --only-upgrade openssl
systemctl restart apache2
# Verify health
curl -f https://web-03.internal/health || exit 1
# Return to load balancer
curl -X POST https://lb.example.com/api/enable/web-03
If something breaks, stop immediately. Don't push forward hoping it will work itself out. Roll back the broken batch, investigate, fix the issue in your test environment, and start the wave progression again.
Stage Five: Verify and Document (Hours 21-24)
Once all servers are patched, verify that the vulnerability is actually fixed. The patched package version should appear in your inventory:
ansible all -m shell -a "dpkg -l | grep openssl" | grep 3.0.8
Check that dependent services restarted and are running the new library. For shared libraries like OpenSSL, processes need a restart to load the updated code. A lingering process running the old version still has the vulnerability:
# Find processes using deleted libraries (old versions)
lsof | grep DEL | grep libssl
Run a targeted vulnerability scan if you have one. Some organizations require proof that the CVE is no longer detectable before closing the ticket.
Update your asset inventory and configuration management database. Mark which servers got patched, when, and by whom. This audit trail matters when the next CVE drops and you need to know your baseline.
Write a brief post-mortem for your team. What went well? What took longer than expected? What would you do differently next time? Even successful patches have lessons.
Notify stakeholders that the patch is complete. Include basic metrics—how many servers, how long it took, whether there was any impact. Management likes numbers.
What Slows Down Emergency Patches
In my experience, three things turn a 24-hour patch into a 72-hour nightmare.
First, no test environment. If you're testing patches for the first time on production, you will break things. Maintain at least one server or VM that mirrors production for exactly this scenario.
Second, manual deployment processes. SSHing into 30 servers one by one to run apt upgrade is slow and error-prone. Even a basic Ansible playbook or bash script will cut deployment time by 80 percent.
Third, no rollback plan. When a patch breaks something at midnight, you need to revert in minutes. If your rollback process is "reinstall the old package and hope all the config files are still compatible," you're in trouble. Snapshots and backaged package versions make rollback a non-event.
What Works in Production
Emergency patching doesn't have to be chaotic. You need a test environment that mirrors production, a rollback plan that works when you're tired, and a deployment process that catches problems early.
I've shipped critical patches in under 20 hours using this workflow. The triage phase forces you to understand the real risk instead of panicking. Testing in isolation catches breaking changes before they hit production. Wave-based deployment gives you circuit breakers at every stage.
Document your process now, before the next CVE drops. Write down your server inventory, your deployment commands, and your verification steps. When you're under pressure at 2 AM, you want a checklist you can follow half-asleep, not a decision tree you have to invent on the spot.
