Skip to content
Back to Blog
Performance11 min read

Advanced NVMe VPS Performance Tuning: 7 Providers Tested

Real-world NVMe VPS optimization beyond marketing benchmarks—queue depth tuning, filesystem choices, and provider comparison for maximum throughput.

Written by Abdul AbrorTechnical Hosting Support Engineer
Advanced NVMe VPS Performance Tuning: 7 Providers Tested
On this page

Most NVMe VPS comparison articles stop at manufacturer datasheets and synthetic benchmarks. That tells you almost nothing about real workload performance once you layer on virtualization overhead, kernel I/O schedulers, filesystem metadata operations, and noisy neighbors.

I've spent years migrating high-traffic WordPress sites and database servers to NVMe VPS hosts. The difference between a poorly tuned NVMe instance and an optimized one can be 3x in actual application throughput, even when the underlying hardware is identical. This guide focuses on the configuration details that matter and which providers give you the control to apply them.

Why raw IOPS numbers lie

Vendor spec sheets love to advertise million-IOPS NVMe drives. You'll never see those numbers on a shared VPS. Virtualization adds latency. The hypervisor's I/O scheduler sits between your VM and the physical drive. Block size, queue depth, and access patterns shift everything.

A 4KB random read benchmark might show 100K IOPS, but your database does 16KB reads with fsync after every transaction. Different workload, different result. The filesystem you choose—ext4, XFS, Btrfs—changes metadata overhead and write amplification. Then there's the noisy neighbor problem: another VM hammering the same physical NVMe can crater your performance during their backup window.

What actually matters for hosting workloads? Latency consistency under mixed read-write load, not peak sequential throughput. I've watched a VPS with "300K IOPS" grind to single-digit IOPS during neighbor activity while a properly isolated competitor held steady at 15K.

Queue depth and the virtio-blk trap

Most VPS providers default to virtio-blk or virtio-scsi for block device virtualization. Check yours:

lsblk -d -o name,rota,disc-max
cat /sys/block/vda/queue/scheduler

If you see vda or vdb, you're on virtio. Default queue depth is often 128 or lower. NVMe drives excel with queue depths of 256-1024, but the virtio layer caps what the guest can submit. You can tune the guest side:

echo 1024 > /sys/block/vda/queue/nr_requests
echo 512 > /sys/block/vda/queue/read_ahead_kb

Add those to a systemd service or rc.local to persist across reboots. Some providers let you adjust virtio queue parameters at the hypervisor level via support tickets—Vultr and Linode have done this for me in the past.

The scheduler matters too. none or noop works well for NVMe-backed VMs because the physical drive's internal scheduler is smarter. mq-deadline is fine if you have mixed SSD and rotational setups, but most pure NVMe hosts should skip it:

echo none > /sys/block/vda/queue/scheduler

Make it permanent in /etc/default/grub:

GRUB_CMDLINE_LINUX="elevator=none"

Then update-grub and reboot.

Filesystem choice: ext4 vs XFS for database hosts

Ext4 is the safe default. XFS shines on large files and parallel writes. For MySQL or PostgreSQL on NVMe, XFS usually wins because it handles delayed allocation better and has lower metadata overhead on big tablespaces.

Format with tuned parameters:

mkfs.xfs -f -d agcount=8 -l size=128m /dev/vdb

Set aggressive mount options in /etc/fstab:

/dev/vdb  /var/lib/mysql  xfs  noatime,nodiratime,logbufs=8,logbsize=256k,largeio,swalloc  0 0

noatime skips access-time updates on reads—huge for databases. logbufs=8 and logbsize=256k improve transaction log throughput. largeio tells XFS you're doing big I/O, so it optimizes readahead.

For ext4 on NVMe, disable the journal if you can tolerate a filesystem check after a crash:

tune2fs -O ^has_journal /dev/vdb

Or keep the journal but put it on a separate fast device. Most VPS setups don't give you that luxury, so journaling stays on the same NVMe.

Real-world benchmarking: fio configs that match your workload

Synthetic benchmarks tell you if something is catastrophically broken. To predict actual application performance, write an fio job that mimics your database or web server pattern.

For a MySQL server with InnoDB and the doublewrite buffer:

[mysql-simulation]
rw=randrw
rwmixread=70
bs=16k
iodepth=64
direct=1
fsync=1
runtime=300
time_based=1
filename=/dev/vdb

Run it:

fio mysql-simulation.fio --output=results.json --output-format=json

Parse the JSON for 99th-percentile latency, not average IOPS. Averages hide stalls. A workload that spikes to 500ms latency every thirty seconds will make your site feel sluggish even if average latency is 2ms.

For WordPress or static site generators, test small random reads with low queue depth:

[wordpress-sim]
rw=randread
bs=4k
iodepth=4
direct=1
runtime=180
time_based=1
filename=/dev/vdb

Most PHP-FPM workers don't queue deep I/O—they're blocked waiting on a file read. Low queue depth is realistic.

Seven providers compared: what you actually get

I tested comparable NVMe VPS plans (4 vCPU, 8GB RAM, NVMe storage) across seven hosts over three months. I ran the fio profiles above every week at random times to catch noisy-neighbor variance.

Vultr gives consistent performance. Queue depth tuning worked immediately, and 99th-percentile latency stayed under 5ms even during evening traffic. Their NVMe instances are on dedicated drives, not shared with SATA VMs. Control panel exposes decent metrics.

Linode (now Akamai) has slightly higher baseline latency—around 3-4ms—but variance is tight. I never saw a latency spike above 20ms in ninety days. Their block storage uses NVMe but goes over the network, so local NVMe instances are faster. Support responded to a virtio tuning request in under four hours.

DigitalOcean premium NVMe droplets perform well in benchmarks but show occasional 100ms+ stalls during neighbor backups (I assume). Standard NVMe droplets share more aggressively. Managed databases run on separate NVMe pools, which is smart. API and automation tooling are first-rate.

Hetzner Cloud offers the best price-to-performance ratio in Europe. NVMe instances are fast and stable. Latency percentiles rival Vultr. Downside: fewer regions, and US latency from their German datacenters is 100ms+. If your audience is in the EU, this is a top pick.

OVHcloud NVMe VPS plans underperform on small random I/O—my WordPress fio test showed 30% lower IOPS than the others. Sequential writes are fine. They oversubscribe aggressively. For batch processing or media encoding it's cheap and adequate. For latency-sensitive databases, skip it.

Contabo advertises NVMe but I found it's often NVMe behind network block storage with caching. Latency is inconsistent. The price is unbeatable, but you're trading performance predictability. Good for dev environments, not production databases.

AWS Lightsail (NVMe instances) is expensive for raw specs, but noisy neighbors are almost nonexistent because AWS isolates block I/O bandwidth per instance. If you need guaranteed performance and AWS integration, it's worth the premium. If you're cost-sensitive and can tune yourself, Vultr or Linode deliver more IOPS per dollar.

Tuning the Linux I/O stack for multi-threaded apps

If you run a multi-threaded application (Java app server, Node cluster, database with parallel query workers), enable multi-queue block layer support. Modern kernels do this by default, but verify:

cat /sys/block/vda/queue/nomerges
cat /sys/block/vda/mq/0/cpu_list

If you see multiple mq/* directories, blk-mq is active. Tune nr_requests per queue:

for queue in /sys/block/vda/mq/*; do
    echo 1024 > "$queue/../nr_requests"
done

For databases, disable merging entirely to reduce latency:

echo 2 > /sys/block/vda/queue/nomerges

2 means no merges at all. Trades some throughput for lower tail latency.

When should you add more NVMe or scale vertically?

Monitor iostat -x 1 during peak load:

iostat -x 1 60

Watch %util and await. If %util stays above 80% and await climbs above 10ms, your workload is I/O-bound. Check if CPU is idle—if so, adding vCPUs won't help. You need either faster storage, better tuning, or a workload redesign (caching, read replicas, sharding).

Sometimes the fix is application-level. I cut database IOPS by 60% once by switching from JSON columns (which rewrite the whole blob) to normalized tables. The NVMe VPS was never the bottleneck; the schema was.

Common pitfalls and how to avoid them

TRIM/discard support: Many VPS hosts enable discard by default, which is great for SSD longevity but adds latency on every delete. Test with and without:

# check current state
lsblk -D
# disable for a mount
mount -o nodiscard /dev/vdb /mnt/test

For write-heavy workloads, nodiscard can improve throughput. Run fstrim manually during maintenance windows instead.

Swapping on NVMe: Swap is faster on NVMe than rotational drives, but it's still slower than RAM. If vmstat 1 shows si/so columns with non-zero values, you're swapping under load. Add RAM or reduce memory usage. Don't let NVMe swap trick you into thinking it's acceptable—latency still spikes.

Backup-induced I/O starvation: Schedule backups during low-traffic windows. Use ionice -c3 to run backup scripts as idle I/O class:

ionice -c3 tar czf backup.tar.gz /var/www

Better yet, snapshot at the hypervisor level if your provider supports it. Block-level snapshots don't involve the guest filesystem at all.

What to measure after tuning

Benchmarks before and after changes prove value. Track these for a week:

  • 50th, 95th, 99th percentile I/O latency (from fio JSON output)
  • Application response time (from access logs or APM tools)
  • Database query time distribution (slow query log analysis)
  • CPU I/O wait percentage (iostat or top)

If 99th-percentile latency drops and I/O wait falls, your tuning worked. If application response time doesn't improve, the bottleneck is elsewhere—probably database queries or external API calls.

FAQ: advanced NVMe VPS tuning

Does over-provisioning help on VPS NVMe?
No. Over-provisioning is a physical-drive feature where the manufacturer reserves extra NAND for wear leveling. You can't configure it at the guest level, and the hypervisor or provider handles it.

Should I use LVM or partition directly?
LVM adds a tiny latency cost (sub-millisecond) but gives you snapshots and resizing. For a single static partition, skip LVM. For flexibility, the overhead is negligible on NVMe.

Can I run RAID on NVMe VPS storage?
You can software RAID inside the guest, but you're RAIDing virtual devices backed by the same physical NVMe in many cases. The provider usually handles redundancy at the hypervisor level. RAID 0 for performance makes sense if you have multiple NVMe volumes; RAID 1 for redundancy is pointless—use provider snapshots instead.

What if fio shows great IOPS but my app is still slow?
Your app might be doing synchronous I/O with small transactions or waiting on network. Profile with strace -c or an APM tool. NVMe can't fix single-threaded code that serializes every operation.

Does enabling swap on NVMe hurt the drive lifespan?
Modern NVMe drives handle hundreds of terabytes written before wear-out. Swap activity on a VPS is trivial compared to that. Lifespan isn't the problem—latency is. Avoid swapping by sizing RAM correctly.

Measure twice, tune once

The best NVMe VPS for you depends on workload, region, and budget. Vultr and Linode deliver reliable performance with knobs you can turn. Hetzner wins on price in Europe. DigitalOcean fits if you're already in their ecosystem. OVH and Contabo work for non-latency-sensitive tasks.

After choosing a provider, tune the I/O stack: adjust queue depth, pick the right filesystem, disable unnecessary features like atime updates, and test with realistic fio workloads. Monitor latency percentiles, not averages. If 99th-percentile latency stays low under load, your users won't notice the underlying hardware—they'll just see a fast site.