You bought an NVMe VPS expecting single-digit millisecond latency and hundreds of thousands of IOPS. Instead, disk waits spike under load and your database crawls. The marketing worked, but the default config didn't.
Most NVMe VPS performance advice stops at "use the right filesystem" or "enable TRIM." That's table stakes. The real speed killers live deeper: wrong I/O schedulers for your workload, CPU affinity conflicts, IRQ storms on a single core, and NUMA topology mismatches that turn local memory access into remote fetches. I've traced these problems through support tickets where customers measured 4x throughput gaps between identical VPS configs after fixing one tuning mistake.
Wrong I/O scheduler for your NVMe workload pattern
The default I/O scheduler on many VPS images is still mq-deadline or even the legacy cfq. NVMe SSDs have no moving parts and can handle thousands of parallel operations. Queue-based schedulers designed to minimize disk head movement add pure latency overhead.
Check your current scheduler:
cat /sys/block/nvme0n1/queue/scheduler
If you see [mq-deadline] or [cfq] in brackets, you're reordering I/O requests that should go straight to hardware. For random-read workloads like database queries, the reordering costs you 20-40% of your available IOPS. For sequential writes, the benefit is marginal at best.
Switch to none (also called noop on older kernels):
echo none > /sys/block/nvme0n1/queue/scheduler
Make it permanent by adding elevator=none to your kernel command line in /etc/default/grub and running update-grub. The exception: if you run a write-heavy append-only log workload and you see better throughput with mq-deadline, keep it. Profile first.
Scheduler choice interacts with queue depth. NVMe drives expose multiple hardware queues. If your scheduler is single-queue (like old noop), you bottleneck at the software layer before the hardware queues even fill. Use none on multi-queue kernels (4.10+) to let the block layer distribute requests directly across NVMe submission queues.
CPU pinning without matching NVMe NUMA node
You pinned your database process to specific CPU cores for consistent cache behavior. Good start. But if those cores sit on a different NUMA node than the NVMe controller, every I/O request crosses the inter-socket link and adds 50-100ns of latency per operation. At 100k IOPS, that's measurable.
Find which NUMA node owns your NVMe device:
cat /sys/block/nvme0n1/device/numa_node
If it returns -1, NUMA awareness is disabled or the device is equidistant (rare on modern hardware). If it returns 0 or 1, pin your I/O-bound processes to CPUs on that same node:
numactl --cpunodebind=0 --membind=0 your-database-command
This keeps CPU, memory allocations, and NVMe controller on the same memory bus. I've seen PostgreSQL query latency drop by 15% after fixing a cross-node pinning mistake on a 32-core VPS.
Virtualized environments complicate this. KVM and Xen expose virtual NUMA topologies to the guest, but the mapping to physical nodes depends on how the hypervisor scheduled your vCPUs. Check lscpu output for the NUMA map inside your VPS, then align your process pinning. If the hypervisor didn't pass through NUMA topology (common on smaller VPS plans), you won't see distinct nodes and this optimization doesn't apply.
IRQ affinity sending all NVMe interrupts to one core
NVMe controllers generate an interrupt completion for every finished I/O operation. Default IRQ affinity often routes all interrupts for a device to a single CPU core. That core becomes a bottleneck: it spends all its cycles handling interrupts while your application threads on other cores sit idle waiting for I/O.
Check current IRQ affinity for your NVMe device:
grep nvme /proc/interrupts
Note the IRQ numbers, then check their affinity:
cat /proc/irq/<irq_number>/smp_affinity_list
If you see a single core (like 0 or 3), spread the load. For NVMe devices with multiple MSI-X vectors (most modern controllers), distribute interrupts across cores:
echo 0-7 > /proc/irq/<irq_number>/smp_affinity_list
Better yet, use irqbalance daemon to handle this dynamically, but verify it's actually distributing NVMe IRQs. On some distributions, the default policy keeps NVMe on a single core anyway. Watch /proc/interrupts under load to confirm interrupt counts increment across multiple CPUs.
Some KVM hypervisors assign vCPU affinity in ways that cluster interrupts regardless of guest-side tuning. If you're on a VPS and irqbalance changes nothing, the issue lives at the hypervisor layer and you'll need to raise a ticket with your host.
Virtio-scsi vs virtio-blk: choosing the wrong backend
Your VPS probably presents the NVMe volume to your guest OS through either virtio-blk or virtio-scsi. Virtio-blk is simpler and lower-latency for single-queue workloads, but it doesn't support SCSI command passthrough or advanced features like TRIM in all configurations. Virtio-scsi adds a SCSI emulation layer that supports multiple LUNs, hotplug, and better scaling across queues—but it adds a few microseconds per I/O.
Check your block device type:
lsblk -o NAME,TRAN
If you see virtio under TRAN, that's not enough detail. Check the actual driver:
ls -l /sys/block/vda/device/driver
If it points to virtio-pci, look deeper:
lsmod | grep virtio
You'll see either virtio_blk or virtio_scsi loaded. For single-disk VPS with high random I/O (databases, key-value stores), virtio-blk usually wins. For multi-disk setups or environments where you need SCSI commands (like UNMAP for TRIM), virtio-scsi is correct.
The problem: many VPS templates default to virtio-scsi for flexibility, even when you're running a single NVMe volume. That costs you latency for features you don't use. If you control the VM definition (on KVM-based hosts where you have API access), switch to virtio-blk and re-test.
On managed VPS where you can't change the hypervisor config, at least confirm your guest kernel is using virtio-blk drivers if the host supports it. The driver negotiation happens at boot; mismatched expectations between host and guest can fall back to slower modes.
Queue depth too low for parallel workloads
NVMe drives handle dozens or hundreds of parallel operations, but your block layer queue depth might cap at 32 or 64. If your application issues concurrent I/O (databases with multiple worker threads, async file servers), a shallow queue depth serializes requests that the hardware could process in parallel.
Check current queue depth:
cat /sys/block/nvme0n1/queue/nr_requests
Default is often 128. For workloads with high concurrency, increase it:
echo 1024 > /sys/block/nvme0n1/queue/nr_requests
Watch I/O stats under load:
iostats -x 1
Look at the aqu-sz column (average queue size). If it consistently sits near your nr_requests limit, you're queue-depth bound. If it's far below, increasing the limit won't help—your application isn't issuing enough concurrent I/O to saturate the hardware.
There's a ceiling: beyond 2048, you're unlikely to see gains and you increase memory pressure from queued bios. Profile your specific workload. Random small reads benefit more from deeper queues than sequential large writes.
On virtualized NVMe, the host-side queue depth matters too. If the hypervisor's virtio queue is shallow, your guest-side tuning hits a wall. You can't control that on a VPS, but understanding it helps you interpret why queue depth increases stop helping past a certain point.
Not aligning partitions to NVMe erase block boundaries
Misaligned partitions force the SSD controller to perform read-modify-write cycles when updating data that spans erase block boundaries. NVMe drives typically use 4KB physical sectors, but some use larger erase blocks (128KB or more). If your partition starts at an odd sector, every write incurs extra overhead.
Check partition alignment:
sudo parted /dev/nvme0n1 align-check optimal 1
If it returns "not aligned," your writes are slower than they should be. For new VPS deployments, ensure partitions start at a multiple of 1MiB (sector 2048 with 512-byte sectors). Most modern partitioning tools do this automatically, but older templates or dd-cloned images sometimes carry misaligned layouts.
Fixing alignment on an existing filesystem requires repartitioning and restoring data. Not practical on a running production system. Verify alignment before you migrate data to a new VPS. If you're setting up a fresh instance, use parted or gdisk and accept the defaults—they'll align to 1MiB boundaries.
Filesystem alignment matters too: XFS and ext4 will auto-detect stripe width on most NVMe drives, but if you override with mkfs.xfs -d su=X,sw=Y, make sure those values match the actual device geometry. For NVMe, alignment is less critical than on RAID arrays, but getting it wrong still costs you 5-10% throughput on write-heavy workloads.
Disabled or misconfigured NVMe multipath in virtualized environments
Some hypervisors expose the same NVMe namespace through multiple paths for redundancy or load distribution. Without NVMe multipath enabled in the guest kernel, the OS sees multiple devices (/dev/nvme0n1, /dev/nvme1n1) but can't aggregate them. You get path failover but no I/O load balancing.
Check if NVMe multipath is active:
nvme list
If you see duplicate entries for the same namespace (same serial/model), multipath should be handling them. Check kernel support:
cat /sys/module/nvme_core/parameters/multipath
If it returns N, multipath is disabled. Enable it by adding nvme_core.multipath=Y to your kernel command line and reboot. After enabling, nvme list should show a single /dev/nvme0c0n1 controller with multiple paths underneath.
On VPS platforms, multipath is rare unless you're on enterprise-grade infrastructure with dual NVMe controllers. Most consumer VPS plans present a single path. Still, I've debugged cases where a hypervisor migration or live snapshot created duplicate NVMe namespaces in the guest, and without multipath enabled the OS picked one path arbitrarily and ignored the other—cutting available bandwidth in half.
If you're not sure whether your VPS uses multipath, check with your provider. Enabling it when you have only one path does no harm but adds a tiny bit of code in the I/O path.
Swapping to NVMe instead of disabling swap entirely
You have fast NVMe, so swap-to-disk is okay, right? Wrong. Any swapping at all means the kernel is evicting active memory under pressure. Even NVMe swap adds milliseconds compared to RAM access. If your VPS is swapping, you're memory-constrained and no amount of NVMe tuning will fix the fundamental problem.
Check swap usage:
free -h
cat /proc/swaps
If swap is in use, find the memory hog:
ps aux --sort=-%mem | head -n 10
Option one: add more RAM. Option two: tune the application to use less memory. Option three: disable swap entirely and let the OOM killer handle it:
swapoff -a
Remove swap entries from /etc/fstab to persist across reboots. The counterargument: swap prevents OOM kills. True, but in a latency-sensitive production environment, a service that's swapping is already failing its SLA. It's better to crash and restart than to limp along with 50ms query latencies.
If you must keep swap (for kernel memory compression or emergency headroom), set vm.swappiness=1 to make the kernel prefer dropping cache over swapping process memory:
sysctl vm.swappiness=1
echo "vm.swappiness=1" >> /etc/sysctl.conf
But seriously: if you're using swap on an NVMe VPS, you need more RAM, not faster swap.
Ignoring NVMe idle power state transitions under low load
NVMe drives support multiple power states (PS0, PS1, etc.) to save energy when idle. The controller can drop into a lower-power state after a few milliseconds of inactivity, then take microseconds to wake up on the next I/O. For latency-sensitive workloads with bursty patterns, those wake-up delays add up.
Check supported power states:
nvme id-ctrl /dev/nvme0 | grep -A 5 "^ps"
You'll see exit latencies for each state. Some drives wake from PS1 in under 10µs; others take 500µs. If your workload has strict P99 latency requirements, disable automatic power state transitions:
nvme set-feature /dev/nvme0 -f 0x0c -v 0
This forces the drive to stay in PS0 (maximum performance). Power consumption increases slightly, but you eliminate tail latency spikes from power state transitions. On a VPS where you don't pay the power bill directly, there's little downside for performance-critical workloads.
Some hypervisors override guest-side NVMe power management settings or don't pass through the relevant commands. If nvme set-feature returns an error, the host is blocking it. In that case, you can't control power states from the guest, and you'll need to ask your provider whether they disable APST (Autonomous Power State Transition) at the hypervisor level.
For general-purpose workloads, leaving power states enabled is fine. For databases or real-time services where P99 latency matters more than power draw, disabling APST cuts tail latency by 10-20% in my load tests.
What to check first when NVMe speed drops
Start with iostat -x 1 and watch %util and await. High utilization with low await means you're saturating IOPS but latency is fine. High await with low utilization means something upstream (CPU, scheduler, IRQ affinity) is the bottleneck.
Next, confirm your I/O scheduler is none and queue depth is adequate. Check dmesg | grep -i nvme for controller errors or resets—these indicate hardware or driver issues the hypervisor should fix.
If metrics look normal but performance is bad, profile the application layer. Use perf or bpftrace to trace I/O latency inside your application. Often the problem is synchronous writes that should be async, or single-threaded I/O that can't use NVMe's parallel queues.
FAQ: NVMe VPS tuning edge cases
Does TRIM matter on virtualized NVMe?
Depends on the hypervisor. If the host runs TRIM on the underlying physical NVMe and uses thin provisioning, guest-initiated TRIM helps. Many VPS platforms disable TRIM passthrough for security or simplicity. Run fstrim -v / to test—if it returns freed space, TRIM works. If not, the host handles it at the storage layer.
Can I benchmark my VPS NVMe reliably?
Yes, but use fio with direct I/O (--direct=1) to bypass cache, and run tests long enough (60s+) to account for noisy neighbor effects. Synthetic benchmarks won't match real workload performance; profile your actual application under load.
Should I use a separate partition for swap even if I disable swap?
No. If swap is off, the partition wastes space. If you need emergency swap later, create a swapfile on your root filesystem—it's more flexible and performs identically on NVMe.
How do I tell if my VPS actually uses NVMe or just claims it?
Run lsblk -d -o name,rota. If ROTA is 0, it's SSD (could be SATA or NVMe). Then check cat /sys/block/nvme0n1/queue/rotational—should be 0. Finally, nvme list will error if the device isn't real NVMe. Some providers present SATA SSDs as virtio-blk and call it "NVMe" in marketing.
Fixing the layers you control
You can't change the hypervisor's NVMe presentation or the physical drive model your host chose. You can tune the guest kernel, align workloads to NUMA topology, fix IRQ affinity, and choose the right I/O scheduler. Those changes combined often recover 30-50% of the performance gap between marketed specs and actual throughput.
Profile first. Change one variable. Re-test. Don't tune blindly—half the "optimizations" floating around forums are cargo-cult configs that helped one specific workload and hurt yours. Measure before and after every change, and revert anything that doesn't show measurable improvement.
![9 NVMe VPS Performance Mistakes That Kill Speed [2026]](/images/blog/9-nvme-vps-performance-mistakes-that-kill-speed-2026.jpg)