Skip to content
Back to Blog
Performance11 min read

9 NVMe VPS Hosting Performance Mistakes That Kill Speed

Deep optimization errors even experienced admins make on NVMe VPS setups—from scheduler misconfigurations to interrupt affinity problems that silently throttle throughput.

Written by Abdul AbrorTechnical Hosting Support Engineer
9 NVMe VPS Hosting Performance Mistakes That Kill Speed
On this page

The hidden costs of default NVMe configs

You paid for NVMe. The marketing promised microsecond latency and half a million IOPS. Reality? Your database still crawls during backups, and compiling a medium-sized codebase takes longer than it should. The drive isn't the problem—your stack is.

Most NVMe VPS performance issues live in layers you configured once and forgot. I've seen setups where a single sysctl knob or a wrong I/O scheduler cut throughput by 40%. The defaults shipped by hosting providers assume spinning rust or generic SSDs, not the beast you're running.

This isn't about basic tuning. We're past "use XFS" and "enable TRIM." These are the edge cases and second-order effects that separate a fast VPS from one that merely looks fast in benchmarks.

Wrong I/O scheduler for your workload

The kernel defaults to mq-deadline or bfq on most distros, and that's fine for SATA. NVMe doesn't need complex scheduling—it has dozens of hardware queues and can handle thousands of parallel operations.

Check what you're running:

cat /sys/block/nvme0n1/queue/scheduler

If you see anything other than none, you're adding latency. The none scheduler (also called noop in older kernels) hands requests straight to the hardware without reordering or merging.

Switch it:

echo none > /sys/block/nvme0n1/queue/scheduler

Make it permanent in /etc/udev/rules.d/60-ioschedulers.rules:

ACTION=="add|change", KERNEL=="nvme[0-9]n[0-9]", ATTR{queue/scheduler}="none"

Exception: if you're running a mix of foreground web traffic and heavy batch jobs, bfq can keep interactive latency low. Test both under your actual load. The difference shows up in tail latency, not averages.

Queue depth set too low

NVMe thrives on depth. A single queue with depth 1 can't saturate the bus, no matter how fast the drive.

Check current settings:

cat /sys/block/nvme0n1/queue/nr_requests

Default is often 128 or 256. NVMe drives handle 1024+ without breaking a sweat. Bump it:

echo 1024 > /sys/block/nvme0n1/queue/nr_requests

Also check queue_depth for the NVMe namespace itself:

cat /sys/block/nvme0n1/device/queue_depth

You can't change this without reloading the driver, but you should know the limit. If your workload is random 4K reads and you're only getting 30K IOPS, shallow queues are likely the culprit.

One gotcha: virtualized NVMe on certain hypervisors has a hard ceiling. KVM with virtio-blk typically caps at 128 or 256 no matter what you set. Check dmesg after changing values—the kernel will complain if it ignores your request.

Interrupt affinity spreading load poorly

NVMe uses MSI-X interrupts, one per queue. By default the kernel spreads these across all CPUs, which sounds smart but often isn't. NUMA topology and CPU cache matter.

See current mappings:

grep nvme /proc/interrupts

Then check which CPUs handle each IRQ:

cat /proc/irq/*/smp_affinity_list | head -20

If your VPS has 8 vCPUs and interrupts are bouncing between all of them, you're burning cycles on cache misses. Pin NVMe interrupts to a subset of cores and keep your application on others.

Example: pin NVMe to cores 0-3, run your app on 4-7:

for irq in $(awk '/nvme/ {print $1}' /proc/interrupts | tr -d ':'); do
    echo 0-3 > /proc/irq/$irq/smp_affinity_list
done

Then start your database or web server with taskset -c 4-7.

On NUMA systems, keep the NVMe interrupts and the consuming process on the same node. Cross-node memory access kills low-latency workloads.

Mount options ignoring NVMe characteristics

Your filesystem mount probably has relatime and discard=async, and you called it done. Not enough.

For ext4:

noatime,nodiratime,discard=async,commit=60,journal_async_commit
  • noatime stops the kernel from writing back access times on every read. nodiratime does the same for directories.
  • commit=60 syncs metadata every minute instead of every 5 seconds. Safe on a VPS with battery-backed or capacitor-protected storage (most cloud NVMe qualifies).
  • journal_async_commit lets the journal write continue while the commit block is being written.

For XFS:

noatime,nodiratime,discard=async,logbsize=256k,largeio,inode64
  • logbsize=256k enlarges the journal buffer, reducing write amplification.
  • largeio hints that the workload uses large sequential I/O. Helps with databases and VM images.
  • inode64 allows inodes anywhere on the disk, not just the first TB. Matters on larger volumes.

Add these in /etc/fstab and remount. You should see fewer write syscalls in iostat output.

Filesystem block size misaligned with workload

Default block size is 4K. NVMe drives work in 4K pages internally, so that sounds fine. But if you're storing database files or VM images averaging 128K or larger, you're fragmenting unnecessarily.

For XFS you can set a larger block size at format time:

mkfs.xfs -b size=16384 /dev/nvme0n1p1

For ext4 the same:

mkfs.ext4 -b 16384 /dev/nvme0n1p1

Catch: you can't change block size after creation, and some tools choke on non-4K blocks. Weigh this before reformatting a live system. The win shows up in reduced metadata overhead and better read-ahead behavior.

If your workload is truly random small I/O (Redis, small file hosting), stick with 4K.

Transparent Huge Pages fighting small writes

THP collapses 4K pages into 2M huge pages to reduce TLB pressure. Good for apps that allocate big contiguous chunks. Terrible for databases that do small random writes.

Check if it's enabled:

cat /sys/kernel/mm/transparent_hugepage/enabled

If you see [always], the kernel aggressively promotes pages. This causes stalls when it tries to compact memory or split huge pages during writes.

Set it to madvise:

echo madvise > /sys/kernel/mm/transparent_hugepage/enabled

Now only programs that explicitly request huge pages get them. MySQL, PostgreSQL, and MongoDB all recommend disabling THP entirely or setting it to madvise.

Make it permanent via sysctl or a systemd unit that runs at boot.

Swap on NVMe eating write endurance

You added swap because why not—NVMe is fast and you had 20GB free. Then your monitoring shows the drive writing 50GB a day, and you can't figure out why.

Swap on NVMe burns write cycles for temporary data that could live in memory or not exist at all. Check current usage:

swapon --show
cat /proc/swaps

If swappiness is above 10, the kernel is paging out cold memory even when RAM isn't full:

cat /proc/sys/vm/swappiness

Lower it:

echo 1 > /proc/sys/vm/swappiness

Persist in /etc/sysctl.conf:

vm.swappiness=1

Better: disable swap entirely if you have enough RAM. Most VPS workloads should never swap. If your app needs more memory than you have, add RAM—don't slowly kill the NVMe.

One exception: if you run containers or VMs that hibernate, you need swap. In that case, put swap on a separate LVM volume so you can monitor its write amplification separately.

Not monitoring per-queue metrics

You watch overall disk throughput and call it done. That misses imbalance across NVMe queues.

Install nvme-cli:

apt install nvme-cli

Dump per-queue stats:

nvme show-regs /dev/nvme0

Also check queue utilization:

cat /sys/block/nvme0n1/mq/*/cpu*

If one queue is handling 80% of requests while others sit idle, your IRQ affinity or application thread pinning is broken. Rebalance as described earlier.

Per-namespace stats:

nvme smart-log /dev/nvme0n1

Look at data_units_written and unsafe_shutdowns. The first tells you if you're burning through endurance faster than expected. The second warns about power issues or kernel panics that could corrupt data.

Running ancient kernels without NVMe optimizations

If you're on Ubuntu 18.04 LTS or CentOS 7 with the stock kernel, you're missing years of NVMe improvements. Multi-queue block layer (blk-mq) was stabilized in 4.X kernels, and every release since added polling modes, better interrupt coalescing, and lower-latency code paths.

Check your kernel:

uname -r

Anything below 5.4 is leaving performance on the table. Anything below 4.11 is actively bad for NVMe.

Upgrade to a current LTS kernel (5.15 or 6.1 at time of writing). On Ubuntu:

apt install linux-generic-hwe-22.04

On RHEL/Rocky/Alma, switch to the elrepo mainline kernel or use a newer minor release.

Reboot, verify the new kernel loaded, then recheck your scheduler and queue settings—they often reset to defaults after a kernel change.

What to check first

Start with the I/O scheduler. It's a one-line fix and consistently yields the biggest single improvement. Then verify your queue depth and interrupt affinity. Those three changes catch most of the low-hanging fruit.

Mount options and block size require planning and sometimes downtime, so test them on a staging VPS first. Measure before and after with fio using your actual workload pattern—random 4K reads for databases, sequential 1M writes for backups, whatever matches reality.

Swap and THP are defensive fixes. They prevent slowdowns rather than speeding things up, but on a busy VPS the difference feels dramatic.

Finally, monitor per-queue metrics weekly. NVMe performance doesn't degrade like spinning disks, but your workload changes and unmasks bottlenecks you didn't have six months ago.

FAQ

Does the none scheduler hurt latency consistency?
No. NVMe hardware queues handle fairness internally, and removing kernel overhead actually improves tail latency. Test it—your P99 and P99.9 metrics should drop.

How do I know if my VPS uses real NVMe or emulated?
Run lsblk -d -o name,rota and check the ROTA column. 0 means non-rotational. Then nvme list will show actual NVMe devices. If nvme list returns nothing but lsblk shows an SSD, you're probably on virtio-scsi over NVMe, which is slower.

Can I set queue depth above 1024?
Depends on the drive and driver. Most NVMe SSDs support up to 64K, but the kernel's nr_requests caps at 4096 in recent versions. Going higher rarely helps because you hit other bottlenecks first.

Should I disable the journal entirely on NVMe?
No. The performance gain is tiny, and the risk isn't worth it. Async commit gives you most of the benefit without the danger.

Why does fio show 500K IOPS but my app only gets 80K?
fio bypasses the filesystem and page cache when you use direct=1. Your app probably doesn't. Also check if your app is single-threaded or using synchronous I/O—NVMe needs parallel requests to shine.

Where speed actually hides

The gap between marketing and reality lives in these details. NVMe gives you the hardware capacity, but the kernel, filesystem, and runtime configuration decide whether you actually use it. I've seen production VPSes running at 20% of their theoretical throughput because no one questioned the defaults.

Fix the scheduler, tune the queues, pin the interrupts. Measure the result with iostat -x 1 during peak load and watch your await and %util numbers drop. That's where the speed was all along.