You paid for NVMe. Your VPS specs look great on paper. But page loads drag, database queries crawl, and disk wait times sit higher than they should.
The problem usually isn't the drive—it's how you're using it. Over the years handling VPS performance tickets, I've seen the same nine mistakes kill speed on otherwise solid hardware. Here's what to fix.
Using the wrong I/O scheduler
Linux ships with multiple I/O schedulers. The default varies by distro and kernel version. Many older systems still use cfq or deadline, both designed for spinning rust. NVMe drives have no seek time and handle massive queue depths. Those legacy schedulers add latency for no benefit.
Check your current scheduler:
cat /sys/block/nvme0n1/queue/scheduler
You'll see something like [mq-deadline] none. The brackets show what's active. For NVMe, use none or noop on older kernels. Modern kernels (5.0+) ship with bfq and kyber as multi-queue options—both work, but for pure speed none still wins on dedicated VPS workloads.
Set it immediately:
echo none > /sys/block/nvme0n1/queue/scheduler
Make it permanent by adding elevator=none to your kernel boot parameters in /etc/default/grub, then run update-grub or grub2-mkconfig.
Running a tiny read-ahead value
The kernel prefetches sequential data using a read-ahead buffer. Default values sit between 128 KB and 256 KB. That's fine for spinning disks. NVMe can deliver multi-GB/s throughput, so a 128 KB buffer is like sipping from a firehose through a straw.
Check it:
sudo blockdev --getra /dev/nvme0n1
If you see 256 or lower, bump it. For web and database workloads, 2048 KB (2 MB) works well. For large file streaming or log processing, go higher—4096 KB or 8192 KB.
sudo blockdev --setra 2048 /dev/nvme0n1
Add it to /etc/rc.local or a systemd service to survive reboots. The performance jump on sequential reads can hit 30-40% depending on workload.
Ignoring queue depth limits
NVMe supports deep queues—often 64K commands per queue. Your system might not be using that. Check the current queue depth:
cat /sys/block/nvme0n1/queue/nr_requests
Many systems default to 128 or 256. Under heavy concurrent load (multiple databases, web workers, or log writes), the queue fills and requests stall. Raise it:
echo 1024 > /sys/block/nvme0n1/queue/nr_requests
There's a ceiling where gains flatten, usually around 2048 for most VPS workloads. Test with fio under your real access patterns before going higher.
Skipping filesystem tuning for NVMe
Ext4 and XFS both default to settings that favor HDDs. Two quick wins: disable access time updates and enable discard.
Access time writes (atime) trigger a metadata update on every file read. That's a write penalty you don't need. Mount with noatime:
# In /etc/fstab
/dev/nvme0n1p1 / ext4 defaults,noatime,discard 0 1
The discard option tells the filesystem to issue TRIM commands, keeping the SSD's free block pool healthy. Some people prefer fstrim on a cron instead of continuous discard—both work, but I've seen fewer surprises with the mount option on modern kernels.
Remount to apply:
sudo mount -o remount /
For XFS, also consider allocsize=64m if you write large files, and logbsize=256k to speed up metadata-heavy operations.
Running swappiness at default
Linux defaults vm.swappiness to 60, meaning it starts swapping pages to disk fairly aggressively. On an NVMe VPS with enough RAM for your workload, swapping is almost always the wrong move. Yes, NVMe swap is faster than HDD swap. It's still slower than RAM by orders of magnitude.
Check current value:
cat /proc/sys/vm/swappiness
Set it low:
sudo sysctl vm.swappiness=10
Make it permanent in /etc/sysctl.conf:
vm.swappiness=10
If you're running databases, drop it to 1. The kernel will still swap under memory pressure, but it'll exhaust cache eviction first.
Leaving dirty page writeback at stock settings
The kernel buffers writes in RAM (dirty pages) before flushing them to disk. Default thresholds allow up to 10-20% of RAM to sit dirty. On a VPS with 8 GB, that's 800 MB to 1.6 GB of buffered writes. Under a write spike, the flush can stall everything.
Two sysctl knobs matter:
vm.dirty_ratio=10
vm.dirty_background_ratio=5
These set the percentage of total memory that can be dirty before the kernel starts blocking writes (dirty_ratio) or begins background writeback (dirty_background_ratio). For NVMe, lower them:
vm.dirty_ratio=5
vm.dirty_background_ratio=2
Add to /etc/sysctl.conf and reload with sysctl -p. This keeps write latency predictable and prevents sudden I/O storms.
Forgetting to tune MySQL or PostgreSQL for NVMe
Database defaults assume spinning disks. If you dropped MySQL or PostgreSQL onto an NVMe VPS without touching config files, you're leaving speed on the table.
For MySQL (or MariaDB), key changes in /etc/my.cnf:
innodb_flush_method=O_DIRECT
innodb_io_capacity=2000
innodb_io_capacity_max=4000
innodb_read_io_threads=8
innodb_write_io_threads=8
O_DIRECT bypasses the OS page cache, avoiding double-buffering. innodb_io_capacity tells InnoDB how many IOPS the storage can handle for background tasks—2000 is conservative for NVMe; modern drives do tens of thousands. Bump innodb_io_capacity_max to match.
For PostgreSQL, edit postgresql.conf:
effective_io_concurrency=200
random_page_cost=1.1
shared_buffers=2GB
random_page_cost defaults to 4.0 (the penalty for random vs. sequential reads on HDDs). NVMe has almost no penalty, so drop it to 1.1. effective_io_concurrency controls how many I/O operations the planner thinks the disk can handle in parallel—set it high.
Restart the database after changes.
Running a bloated kernel with unused modules
Many VPS templates ship with generic kernels that load dozens of drivers and modules you'll never use. Each one adds memory overhead and interrupt latency. If you're chasing microseconds, a leaner kernel helps.
List loaded modules:
lsmod | wc -l
If you see 80+ modules and you're running a simple LAMP stack, you're carrying dead weight. Consider switching to a minimal or tuned kernel build. Some distros offer low-latency or performance-oriented kernel packages.
For Ubuntu:
sudo apt install linux-image-lowlatency
For a truly clean build, compile your own kernel with only the drivers and features you need. That's overkill for most workloads, but I've seen it shave 5-10% off p99 latencies in high-throughput environments.
Ignoring monitoring and I/O wait metrics
You can't optimize what you don't measure. Many people only check CPU and memory. I/O wait (iowait) matters just as much. If processes spend time blocked waiting for disk, your NVMe speed is wasted.
Watch it live:
top
Look at the %wa column. Anything above 5% consistently means something's wrong—misconfigured scheduler, queue saturation, or an application hammering the disk with small random writes.
For deeper visibility, use iostat:
iostat -x 1
Watch %util, await, and svctm. If %util sits near 100% but await is high (above 10 ms), the queue is saturated. Go back and check queue depth and scheduler settings.
Install iotop to see which processes generate I/O:
sudo iotop -o
The -o flag shows only active processes. I've found rogue log rotation scripts, backup tools, and poorly indexed database queries this way.
What's the real fix?
Most of these mistakes compound. You might have the wrong scheduler and a low read-ahead and default database tuning. Each one cuts speed by 10-20%. Stacked up, your NVMe performs like a SATA SSD or worse.
Start with the I/O scheduler and read-ahead. Those take 30 seconds and deliver immediate wins. Then tune your main application (database, web server, whatever hammers the disk hardest). Finally, adjust kernel parameters and monitor the results.
Speed isn't just about hardware. How you configure it matters more.
Frequently asked questions
Does the none scheduler work on all NVMe drives?
Yes. It disables scheduling entirely, letting requests go straight to the drive's internal controller. Every NVMe device handles queueing natively, so there's no downside.
How do I know if my VPS actually has NVMe?
Run lsblk -d -o name,rota. If ROTA shows 0, it's SSD or NVMe. Check nvme list or look for /dev/nvme* devices. Some providers label SATA SSDs as "NVMe" in marketing—verify with the vendor.
Will raising queue depth hurt latency?
Not under normal load. Deep queues help when many processes compete for I/O. If you run a single-threaded app with light I/O, you won't notice a difference. The risk is memory overhead—each queued request consumes a small amount of RAM.
Should I disable swap entirely?
No. Set swappiness low, but keep a swap partition or file. The kernel uses it for emergency memory pressure and certain process management tasks. Disabling it can cause OOM kills under load spikes.
What if fio benchmarks look fine but real-world performance is slow?
Synthetic benchmarks test the drive, not the full stack. Check application logs, query performance, and I/O wait. The bottleneck might be in your app's write patterns, lock contention, or network latency—not the disk.
Start with the scheduler and measure everything
I've walked through nine mistakes. You don't need to fix all of them at once. Change the I/O scheduler first—it's the lowest-hanging fruit. Then measure with iostat and iotop to see where the actual bottleneck lives. Every workload is different. Your database might need deeper queues. Your web server might need better cache tuning.
The point is this: NVMe gives you the hardware. Configuration unlocks it. Make the changes, test under load, and watch your metrics. Speed follows.
