Skip to content
Back to Blog
Performance11 min read

Advanced NVMe VPS Hosting Performance: 7 Proven Providers

Beyond raw speed—deep dive into NVMe VPS optimization, kernel tuning, I/O schedulers, and provider differences that matter for production workloads.

Written by Abdul AbrorTechnical Hosting Support Engineer
Advanced NVMe VPS Hosting Performance: 7 Proven Providers
On this page

Raw NVMe speed means nothing if your VPS kernel is misconfigured or the hypervisor throttles IOPS behind the scenes. Most NVMe VPS reviews stop at synthetic benchmarks—this one focuses on tuning, edge cases, and what actually breaks under load.

Why NVMe alone won't solve your performance problem

I've seen dozens of support tickets where clients moved from SATA to NVMe and still hit bottlenecks. The storage layer is only fast if every component in the stack cooperates—hypervisor queue depth, kernel I/O scheduler, filesystem mount options, and application-level buffering all matter more than the marketing spec sheet.

NVMe delivers sub-100 microsecond latency and parallelism through multiple queues. Traditional SATA SSDs use a single queue with 32 commands max; NVMe supports 64k queues with 64k commands each. But virtualization adds layers: your VPS block device is abstracted through virtio-blk or virtio-scsi, and the host's storage stack sits between you and bare metal.

Check your actual queue depth before anything else:

cat /sys/block/vda/queue/nr_requests
cat /sys/block/vda/queue/scheduler

If nr_requests is 128 or lower and you're running database or high-concurrency workloads, you're leaving performance on the table. Many providers ship conservative defaults.

Kernel I/O scheduler tuning for NVMe

The mq-deadline, bfq, kyber, and none schedulers behave completely differently under NVMe. The none scheduler (also called noop in older kernels) makes sense for NVMe because the device itself handles request ordering more efficiently than the kernel can. For most VPS workloads I recommend testing none first:

echo none > /sys/block/vda/queue/scheduler

Make it persistent by adding elevator=none to your kernel boot parameters in /etc/default/grub and running update-grub.

mq-deadline helps when you have mixed read/write workloads and need fairness—useful for shared hosting environments or WordPress multisites where one tenant's backup job shouldn't starve others. kyber adjusts queue depth dynamically based on latency targets; it's excellent for NVMe but needs tuning. Default target latencies are often too conservative.

bfq prioritizes interactive I/O and works well for desktop-like workloads but adds CPU overhead. Skip it on headless servers.

Benchmark each scheduler under your actual workload. Synthetic tools like fio help, but real application behavior matters more:

fio --name=randrw --ioengine=libaio --iodepth=32 --rw=randrw \
    --bs=4k --direct=1 --size=1G --numjobs=4 --runtime=60 \
    --group_reporting --filename=/mnt/test

Run this with each scheduler and log IOPS and latency percentiles. The 99th percentile latency (lat (usec) in fio output) matters more than average—one slow request can cascade through your application.

Filesystem and mount options that matter

Ext4, XFS, and Btrfs all expose NVMe speed differently. Ext4 with noatime,nodiratime is the safe default. noatime skips updating file access timestamps on every read—irrelevant for most server workloads and saves write cycles.

For databases, add data=writeback if you have application-level crash recovery (like PostgreSQL WAL or MySQL InnoDB). It trades strict ordering for speed. Never use it for the root filesystem.

# Example /etc/fstab entry for a database volume
/dev/vdb1 /var/lib/mysql ext4 noatime,nodiratime,data=writeback 0 2

XFS performs better than ext4 for large files and high-concurrency writes. Hosting providers often default to ext4 because it's familiar, but XFS shines with NVMe's parallelism. If you're running object storage, media transcoding, or log aggregation, test XFS.

Btrfs adds snapshots and compression at a CPU cost. Compression (zstd level 3) can paradoxically improve performance on NVMe by reducing write volume—if your workload is CPU-light and write-heavy. Check with:

sudo btrfs filesystem df /mnt/data
sudo compsize /mnt/data

What separates serious NVMe VPS providers from the rest

Marketing claims are useless. The differences that matter:

1. IOPS throttling and burst policies
Many providers advertise "full NVMe speed" but implement silent IOPS caps. Vultr, Linode, and DigitalOcean publish baseline and burst IOPS per plan tier. Others don't. Run sustained write tests for 10+ minutes to detect throttling:

fio --name=sustained --ioengine=libaio --iodepth=16 --rw=write \
    --bs=4k --direct=1 --size=10G --runtime=600 --filename=/mnt/test

If IOPS drop sharply after the first minute, you've hit a burst limit.

2. Storage backend architecture
Ceph, local NVMe, or hybrid? Distributed storage like Ceph adds latency—typically 1-3ms even with NVMe OSDs—because writes must replicate across nodes. Local NVMe is faster but lacks live migration and built-in redundancy. Ask support directly what they use. Hetzner Cloud uses local NVMe with Ceph for block storage volumes (two-tier approach). DigitalOcean uses distributed block storage everywhere.

For latency-sensitive workloads (Redis, real-time analytics), local NVMe wins. For high-availability apps that need instant failover, distributed storage is the trade-off.

3. Hypervisor and virtio tuning
KVM with virtio-scsi generally outperforms virtio-blk for multi-queue NVMe. Check which your provider uses:

lsblk -d -o NAME,ROTA,DISC-GRAN
dmesg | grep -i virtio

If you see virtio-blk, you're on an older or simpler stack. virtio-scsi with scsi-mq gives you per-vCPU I/O queues, critical for NVMe parallelism.

Some providers let you request virtio-scsi explicitly during provisioning. Others hardcode virtio-blk. This isn't in the marketing materials—you have to test or ask.

4. Network throughput to storage nodes
For cloud block storage (non-local NVMe), the internal network between compute and storage nodes becomes the bottleneck. Providers with 25Gb or 100Gb backend networks deliver consistent performance. Those on 10Gb or oversubscribed links show jitter under load.

You can't measure this directly, but sustained sequential write speed is a proxy. If you pay for "NVMe performance" but sequential writes cap at 500 MB/s, the network is the limit.

Seven providers worth testing for production NVMe VPS

I'm listing these because they publish technical details and support staff can answer hypervisor questions—not because of affiliate programs or sponsorships.

Vultr – Per-plan baseline IOPS documented; local NVMe on high-frequency compute; virtio-scsi available; good latency in most regions. They don't throttle aggressively.

Linode (Akamai) – Distributed block storage uses NVMe OSDs; published IOPS limits; excellent uptime; some latency overhead from Ceph but predictable. Dedicated CPU plans get higher queue depth.

Hetzner Cloud – Local NVMe on CCX line; cheapest per GB; no hidden throttling in my tests; limited regions outside Europe. Block storage volumes use Ceph, so separate them by workload.

DigitalOcean – Consistent distributed NVMe; good documentation; IOPS scale with droplet size; Premium plans get better burst. Support is fast when you need stack-level details.

OVHcloud – Bare-metal-like VPS with local NVMe on some lines; complex product matrix; once you find the right SKU, performance is excellent. Their "Advance" range uses local storage.

UpCloud (MaxiOps) – Marketed specifically on storage performance; tiered IOPS; published architecture docs; higher price but less variance. Support responds to kernel-level questions.

Kamatera – Flexible IOPS allocation; pay-per-GB storage; good for custom tuning; slightly higher latency than top-tier but never throttled in sustained tests. More expensive than Hetzner, less than UpCloud.

Test before committing. Spin up a trial VPS, run fio, check /proc/diskstats under load, and measure 99th percentile latency with your actual app stack.

Benchmarking your actual workload, not synthetic tests

Synthetic benchmarks like fio tell you the ceiling. Real workloads rarely hit it. Your application's I/O pattern—block size, queue depth, read/write ratio—defines actual performance.

For databases, use the database's own benchmarking tool. PostgreSQL's pgbench, MySQL's sysbench, MongoDB's YCSB. For WordPress or similar PHP apps, measure time-to-first-byte under load with ab or wrk while monitoring iostat:

iostat -x 1

Watch %util and await (average wait time). If %util stays below 80% but your app is slow, storage isn't the problem—look at CPU or network. If await spikes above 10ms on NVMe, something is misconfigured or the provider is throttling.

Kernel tuning beyond the I/O scheduler

Two sysctl parameters make a measurable difference on NVMe VPS:

# Increase max queue depth for block devices
echo 1024 > /sys/block/vda/queue/nr_requests

# Reduce swappiness if you have enough RAM
sysctl vm.swappiness=10

# Increase dirty page writeback for better write coalescing
sysctl vm.dirty_ratio=40
sysctl vm.dirty_background_ratio=10

The dirty page settings let the kernel batch writes instead of flushing constantly. This helps NVMe because it can handle large parallel writes efficiently. Don't go above dirty_ratio=40 or you risk long stalls during writeback.

For VPS with less than 4 GB RAM, keep defaults—you need memory pressure signals to avoid OOM.

When NVMe isn't enough: identifying other bottlenecks

If you've tuned everything above and performance still lags, storage probably isn't the issue. Check CPU steal time first:

top
# Look at the %st column

High steal time (above 5% sustained) means the hypervisor is overcommitted. No amount of NVMe tuning fixes noisy neighbors. Move to a dedicated or "guaranteed CPU" plan.

Network-bound apps won't benefit from faster storage. If you're serving static files or proxying requests, measure network throughput separately:

iperf3 -c <remote_host> -P 10

Application-level locking can make storage speed irrelevant. WordPress with 50 plugins that all write to the database on every page load will be slow even on NVMe. Profile with Query Monitor or New Relic before blaming infrastructure.

Check throttling first, tune second

Most NVMe VPS performance problems come from hidden IOPS caps or conservative kernel defaults, not the hardware. Before switching providers, benchmark sustained I/O for 10+ minutes, verify your I/O scheduler matches your workload, and check CPU steal time. If the provider throttles or the hypervisor is overcommitted, no amount of tuning saves you. Pick a provider that publishes IOPS limits and test under load before moving production workloads.

FAQ

Does NVMe help WordPress caching plugins?

Only if your cache backend (Redis, Memcached, or object cache) writes frequently to disk. Most in-memory caches won't see much difference. Page caching that writes static HTML benefits, but the gain is small unless you regenerate cache constantly.

Can I measure IOPS throttling from inside the VPS?

Yes. Run a sustained random-write fio test for 10 minutes and plot IOPS over time. If it drops sharply, you're throttled. Use --output-format=json and graph the per-second IOPS data.

Is XFS faster than ext4 on NVMe?

For large files and parallel writes, yes. For small random I/O (typical WordPress or small databases), the difference is negligible. Test with your workload.

Do I need a special kernel for NVMe VPS?

No. Kernel 4.4+ has stable NVMe and virtio-scsi support. Ubuntu 20.04+, Debian 10+, and recent CentOS/AlmaLinux ship with everything you need.

Will local NVMe survive a host failure?

No. Local storage is fast but not redundant. Back up externally or use block storage volumes if uptime matters more than latency.