Skip to content
Back to Blog
Performance11 min read

Advanced NVMe VPS Performance Tuning: 7 Providers Tested

Deep optimization techniques for NVMe VPS instances—queue depth tuning, IRQ affinity, filesystem parameters, and kernel tweaks that move the needle beyond vendor defaults.

Written by Abdul AbrorTechnical Hosting Support Engineer
Advanced NVMe VPS Performance Tuning: 7 Providers Tested
On this page

Most NVMe VPS guides stop at picking the right plan. That's half the story. The other half lives in kernel parameters, queue management, IRQ routing, and filesystem mount options that vendors leave at defaults—often tuned for spinning rust or generic cloud workloads, not the low-latency characteristics of NVMe.

I tested seven providers over three months, pushing identical workloads through each. What separated the top performers from the middle of the pack wasn't always raw hardware. It was how well the host configuration let you exploit NVMe's parallelism and whether you could tune the guest without hitting virtualization walls.

Block device scheduler matters more than you think

The Linux kernel ships multiple I/O schedulers. Most distributions default to mq-deadline or bfq for backward compatibility. NVMe devices with dozens of hardware queues don't need the same single-queue scheduling logic that SATA drives required.

Check your current scheduler:

cat /sys/block/nvme0n1/queue/scheduler

You'll see something like [mq-deadline] kyber bfq none. The bracketed entry is active. For NVMe under virtualized workloads, none often wins. It bypasses software scheduling entirely and pushes I/O straight to the device's multi-queue dispatch.

Switch it live:

echo none > /sys/block/nvme0n1/queue/scheduler

Make it permanent by adding elevator=none to your kernel command line in /etc/default/grub, then run update-grub and reboot. Three of the seven providers I tested showed a 15–22% latency drop at the 99th percentile just from this change. One provider's images already used none; the rest defaulted to mq-deadline.

Queue depth and nr_requests

NVMe supports deep queues—often 64K commands per queue. Virtualized instances cap that, but you still control the block layer queue depth on the guest side.

Inspect current values:

cat /sys/block/nvme0n1/queue/nr_requests
cat /sys/block/nvme0n1/queue/nr_hw_queues

Default nr_requests is typically 256 or 1024. For databases with heavy random I/O, increasing it to 2048 or 4096 can smooth out bursty writes. For sequential workloads that already saturate bandwidth, a deeper queue just adds latency.

Set it temporarily:

echo 2048 > /sys/block/nvme0n1/queue/nr_requests

Persist it with a udev rule in /etc/udev/rules.d/60-nvme-tuning.rules:

ACTION=="add|change", KERNEL=="nvme[0-9]n[0-9]", ATTR{queue/nr_requests}="2048"

One provider's default was 128, which choked under parallel PostgreSQL writes. Bumping it to 2048 cut transaction commit latency by a third.

IRQ affinity and CPU pinning

NVMe devices generate interrupts for I/O completion. By default, the kernel spreads these across all CPUs. That's fine for physical servers with local NVMe, but in a VPS you might have only 2–4 vCPUs. Spreading interrupts can cause cache thrashing and context switches.

List NVMe IRQ numbers:

grep nvme /proc/interrupts

You'll see entries like nvme0q0, nvme0q1, etc.—one per queue. Pin them to specific cores. If you have four vCPUs and four NVMe queues, assign each queue to one CPU:

echo 1 > /proc/irq/130/smp_affinity_list  # nvme0q0 -> CPU 0
echo 2 > /proc/irq/131/smp_affinity_list  # nvme0q1 -> CPU 1
echo 4 > /proc/irq/132/smp_affinity_list  # nvme0q2 -> CPU 2
echo 8 > /proc/irq/133/smp_affinity_list  # nvme0q3 -> CPU 3

(The value is a bitmask; CPU 0 is bit 0, so 1; CPU 1 is bit 1, so 2; CPU 2 is 4; CPU 3 is 8.)

For persistence, use irqbalance exclusions or a systemd service that reapplies the affinity on boot. Two providers in my test set showed measurable latency improvement with pinned IRQs; five showed no change, likely because the hypervisor already handled queue-to-vCPU mapping.

Filesystem mount options

Ext4 and XFS both have mount flags that trade durability for speed. If your workload can tolerate a few seconds of data loss on crash (caches, session stores, CI build artifacts), you can skip some of the journal overhead.

For ext4, try:

noatime,nodiratime,data=writeback,barrier=0,commit=60

Add those to /etc/fstab:

/dev/nvme0n1p1  /  ext4  noatime,nodiratime,data=writeback,barrier=0,commit=60  0 1
  • noatime stops updating access time on every read.
  • data=writeback decouples data writes from journal commits (faster, less safe).
  • barrier=0 disables write barriers (only safe if your storage has battery-backed cache or you don't care about crashes).
  • commit=60 stretches journal commits to once a minute instead of every five seconds.

For XFS, the equivalent:

noatime,nodiratime,logbufs=8,logbsize=256k,nobarrier

XFS already uses delayed allocation aggressively, so the gains are smaller. I saw 8–12% throughput improvement on write-heavy workloads with data=writeback on ext4. XFS was closer to 4–6%. Your mileage varies with fsync frequency in the application.

Read-ahead tuning

The kernel prefetches blocks when it detects sequential reads. Default read-ahead is conservative—128 KB or 256 KB. NVMe can deliver multi-megabyte bursts without breaking a sweat.

Check current setting:

blockdev --getra /dev/nvme0n1

It returns the value in 512-byte sectors. A result of 256 means 128 KB. Raise it to 2048 (1 MB) or 4096 (2 MB):

blockdev --setra 4096 /dev/nvme0n1

Persist it by adding a line to /etc/rc.local or a udev rule. For streaming logs, backups, or media transcoding, this makes a visible difference. For random-access databases, it does nothing or wastes memory.

So what if you're CPU-bound instead?

NVMe is fast enough that the bottleneck often shifts to CPU—specifically, the overhead of handling interrupts, context switches, and memory copies. If iostat shows low %util but your application still feels slow, profile CPU time.

Install perf:

apt-get install linux-tools-generic  # Debian/Ubuntu
yum install perf                     # RHEL/CentOS

Record 30 seconds of activity:

perf record -a -g sleep 30
perf report

Look for nvme_irq, blk_mq_run_hw_queue, or high time in softirq. If interrupt handling dominates, IRQ affinity or switching to polling mode can help.

Enable polling mode (busy-wait instead of interrupts) for ultra-low latency:

echo 1 > /sys/module/nvme/parameters/poll_queues

This burns CPU cycles but can shave microseconds off latency. Only worth it if you have spare cores and latency matters more than throughput.

Virtual memory and dirty page writeback

Linux buffers writes in RAM before flushing to disk. The default thresholds let dirty pages pile up, which is great for throughput but murder for latency spikes when the kernel finally flushes.

Check current settings:

sysctl vm.dirty_ratio
sysctl vm.dirty_background_ratio

Defaults are often 20 and 10—meaning the kernel starts background writeback at 10% of RAM and blocks writes at 20%. On a 16 GB VPS, that's 1.6 GB of dirty data before a blocking flush.

For lower-latency writes, tighten them:

sysctl -w vm.dirty_ratio=10
sysctl -w vm.dirty_background_ratio=5

Persist in /etc/sysctl.conf:

vm.dirty_ratio = 10
vm.dirty_background_ratio = 5

This trades some peak throughput for predictable latency. In my tests, the 99th-percentile write latency dropped by 40–60% under bursty workloads with these settings.

Transparent Huge Pages and memory locality

Transparent Huge Pages (THP) can hurt NVMe performance if the kernel spends time compacting memory or splitting pages during I/O. Databases like MongoDB and Redis often recommend disabling THP.

Check status:

cat /sys/kernel/mm/transparent_hugepage/enabled

If it says [always], disable it:

echo never > /sys/kernel/mm/transparent_hugepage/enabled
echo never > /sys/kernel/mm/transparent_hugepage/defrag

Add to /etc/rc.local or a systemd unit to persist. I saw no change on pure I/O benchmarks, but PostgreSQL query latency improved slightly—probably from reduced page fault overhead.

Provider-specific quirks I hit

One provider exposed NVMe as virtio-blk instead of native NVMe passthrough. The nvme CLI tools didn't work, and I couldn't tune queue depth the same way. Performance was still good, but you lose visibility.

Another provider's default kernel was 4.19, which lacks some NVMe multi-queue improvements from 5.x. Upgrading to a 5.10+ kernel from backports shaved 10% off random read latency.

A third provider rate-limited IOPS at the hypervisor level—no amount of guest tuning moved the needle past that cap. Check your plan's IOPS limit before you waste time tuning. Run fio with high queue depth and see if IOPS plateau.

fio --name=test --rw=randread --bs=4k --iodepth=64 --numjobs=4 --runtime=60 --time_based --filename=/dev/nvme0n1

If IOPS hit a round number (like exactly 10,000 or 50,000), you're hitting a hypervisor throttle.

What about NVMe-specific features?

NVMe supports features like namespace management, multiple namespaces, and per-namespace I/O queues. In a VPS, you rarely get access to these—the hypervisor abstracts them. One provider let me query SMART data with nvme smart-log /dev/nvme0, which was useful for catching early drive wear. Most block that command.

If your provider exposes the full NVMe admin interface, you can check firmware version and feature support:

nvme id-ctrl /dev/nvme0 | grep -i firmware

No provider in my test set let me update firmware or change controller settings, which makes sense—they manage the physical hardware.

Testing methodology

I ran identical workloads across all seven VPSs: fio for raw block I/O, sysbench for filesystem throughput, and a containerized PostgreSQL instance with pgbench for real-world latency. Each test ran three times at different times of day to catch noisy-neighbor effects.

The spread was wider than I expected. Same NVMe drive model, same CPU cores, but the default kernel and I/O stack settings created a 30% performance gap between the fastest and slowest. After tuning, the gap narrowed to about 10%—mostly explained by hypervisor-level caps.

How do I know if tuning is working?

Benchmark before and after. Use fio for raw I/O and application-specific tools for real workloads. Watch iostat -x 1 during load and look for await (average I/O completion time) and %util. If await drops after a change, it worked.

Will these tweaks break things?

Most are safe to revert. The risky ones are mount options like barrier=0 and data=writeback, which can lose data on crash. Test on non-production first. Scheduler and queue depth changes are safe—worst case, performance doesn't improve.

Do I need to retune after kernel updates?

Scheduler and sysctl settings persist across kernel updates. IRQ affinity and some udev rules might need rechecking if device names or IRQ numbers change, but that's rare in VPS environments.

Can I use these tweaks on cloud block storage too?

Some apply, some don't. Cloud block storage (EBS, Persistent Disks) is network-attached, so IRQ affinity and local NVMe tuning don't help. Scheduler, mount options, and writeback settings still matter.

What to tune first

Start with the I/O scheduler—it's a one-line change with the biggest consistent return. Then tighten dirty page ratios if you see latency spikes. Add queue depth and IRQ affinity only after profiling confirms they're bottlenecks. Filesystem mount options come last; they're the highest risk and often the smallest gain unless your workload is write-heavy and fault-tolerant.

Measure before and after every change. Tuning without data is just superstition.