Skip to content
Back to Blog
Performance11 min read

9 Bare Metal Server Performance Mistakes That Kill Speed

Bare metal servers promise raw power, but common configuration mistakes can cripple performance. Learn the nine errors that slow down your dedicated hardware and how to fix them.

Written by Abdul AbrorTechnical Hosting Support Engineer
9 Bare Metal Server Performance Mistakes That Kill Speed
On this page

Bare metal servers give you dedicated hardware with no hypervisor overhead. But that power comes with responsibility — configure it wrong and you'll get worse performance than a well-tuned VM.

I've seen support tickets where clients paid for premium bare metal and got half the throughput they expected. The hardware wasn't the problem. Their setup was.

Mistake 1: Running the default kernel I/O scheduler

Most distributions ship with the BFQ or CFQ scheduler. These work fine for desktop workloads, but they're terrible for high-throughput database servers or storage arrays.

The default scheduler tries to be fair to all processes. That sounds good until you realize your PostgreSQL instance is competing for disk time with log rotation scripts.

The fix: Switch to mq-deadline for SSDs or NVMe, none for NVMe with very low latency requirements. Check your current scheduler:

cat /sys/block/nvme0n1/queue/scheduler

Change it permanently by adding this to your kernel boot parameters:

elevator=mq-deadline

For NVMe under heavy random I/O, none often wins because the drive's internal parallelism handles scheduling better than the kernel.

Mistake 2: Ignoring NUMA topology

Dual-socket servers have two separate memory controllers. Access memory attached to the local CPU and you get fast reads. Cross the socket boundary and latency doubles.

Applications that aren't NUMA-aware will allocate memory wherever the kernel decides. Your database might run on socket 0 while its buffers sit in socket 1's RAM.

The fix: Pin critical processes to specific NUMA nodes with numactl:

numactl --cpunodebind=0 --membind=0 /usr/bin/mysqld

Check your topology first:

numactl --hardware

For multi-threaded applications like Redis or Memcached, consider running separate instances per NUMA node rather than one big process.

Mistake 3: Choosing RAID 5 for write-heavy workloads

RAID 5 has a write penalty. Every write operation requires reading the old data, reading the old parity, writing new data, and writing new parity. Four operations for every write request.

That's fine for file servers with mostly reads. It's death for databases or mail servers that write constantly.

The fix: Use RAID 10 for write-intensive workloads. Yes, you lose half your capacity. That's the trade-off for avoiding the parity calculation overhead.

For pure speed with some risk tolerance, RAID 0 striping gives you the best throughput. For critical data with heavy writes, RAID 10 is the standard.

RAID 6 is even worse than RAID 5 for writes because it calculates two parity blocks. Reserve it for archive storage where write performance doesn't matter.

Mistake 4: Not tuning the network ring buffers

The default ring buffer sizes are tiny — often 256 or 512 descriptors. That's enough for a desktop. It's not enough for a 10Gbps or 25Gbps link pushing sustained traffic.

When the ring buffer fills, packets get dropped. You'll see this in ethtool -S as rx_dropped or rx_missed_errors, and your application logs will show timeouts that look like network problems.

The fix: Check current ring buffer size:

ethtool -g eth0

Increase it to the maximum supported by your NIC:

ethtool -G eth0 rx 4096 tx 4096

Make it permanent by adding the command to a startup script or using your distribution's network config. Also bump the kernel's receive buffer:

sysctl -w net.core.rmem_max=134217728
sysctl -w net.core.wmem_max=134217728

Those settings go in /etc/sysctl.conf to survive reboots.

Mistake 5: Running swap on the same disk as your database

Swap is emergency overflow. If your system is swapping actively, you've already lost the performance battle. But putting swap on your primary storage device makes it worse.

Every swap-in competes with your application's disk I/O. The database tries to read an index while the kernel swaps out some old process's memory, and both operations slow down.

The fix: Disable swap entirely on database servers with enough RAM:

swapoff -a

Remove the swap entry from /etc/fstab. If you must have swap as a safety net, put it on a separate physical disk — never on an LVM volume that shares the same underlying disk.

Better yet, configure the OOM killer to terminate runaway processes before the system starts swapping. Set vm.swappiness=1 to make swap a true last resort.

So what about CPU frequency scaling?

Servers ship with CPU frequency governors set to powersave or ondemand. These governors reduce clock speed when the system looks idle to save power and heat.

The problem is that they react too slowly. A request comes in, the CPU is still at low frequency, the application waits, then the governor notices and ramps up. By the time the CPU is at full speed, your response time has already blown your SLA.

The fix: Force the performance governor on latency-sensitive workloads:

for cpu in /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor; do
    echo performance > $cpu
done

Install cpufrequtils or tuned and set the profile to throughput-performance to make it permanent. Yes, power consumption goes up. That's the cost of consistent low latency.

For batch processing or background tasks, ondemand is fine. For web servers, databases, or real-time applications, run at full speed.

Mistake 7: Forgetting to disable transparent huge pages

Transparent Huge Pages sounds great — the kernel automatically promotes 4KB pages to 2MB pages for better TLB efficiency. In practice, the compaction and defragmentation cause stalls.

Databases like MongoDB and Redis explicitly recommend disabling THP because the kernel's defragmentation runs at the worst possible time, introducing latency spikes that look like I/O hangs.

The fix: Disable it at boot by adding this to your kernel command line:

transparent_hugepage=never

Or disable it at runtime:

echo never > /sys/kernel/mm/transparent_hugepage/enabled
echo never > /sys/kernel/mm/transparent_hugepage/defrag

Check if it's causing problems by monitoring /proc/vmstat for thp_fault_alloc, thp_collapse_alloc, and thp_split. High numbers during latency spikes confirm the issue.

Mistake 8: Using default TCP congestion control

The default TCP congestion control algorithm (usually Cubic) optimizes for general internet use. It's not tuned for low-latency datacenter networks or high-bandwidth long-distance links.

For connections within the same datacenter, BBR or Reno often performs better. For long-haul replication or backups, BBR significantly improves throughput.

The fix: Change the congestion control algorithm:

sysctl -w net.ipv4.tcp_congestion_control=bbr

Add it to /etc/sysctl.conf and ensure the tcp_bbr kernel module is loaded:

modprobe tcp_bbr
echo tcp_bbr >> /etc/modules-load.d/bbr.conf

BBR is particularly good for connections with packet loss or variable latency. It measures bottleneck bandwidth and round-trip time instead of relying on packet loss as a congestion signal.

Mistake 9: Not monitoring hardware RAID controller cache

Hardware RAID controllers have battery-backed or flash-backed cache. When that cache fails or the battery dies, the controller disables write-back caching and falls back to write-through mode.

Write performance drops by 10x or more. Your applications suddenly slow down and you have no idea why because the RAID array still shows healthy.

The fix: Install the vendor's monitoring tools — megacli for LSI/Broadcom, ssacli for HPE, storcli for newer LSI cards. Check cache status:

megacli -AdpBbuCmd -GetBbuStatus -a0

Set up monitoring alerts for cache status changes. If the battery-backup unit fails, replace it immediately. Some controllers let you force write-back mode without a working BBU, but you risk data loss on power failure.

Also verify your controller's cache policy:

megacli -LDInfo -Lall -aAll | grep "Current Cache Policy"

You want WriteBack for performance. WriteThrough is safe but slow.

What to check first

When you deploy a bare metal server, run through this checklist before putting it into production:

  1. Confirm the I/O scheduler matches your workload (mq-deadline or none for NVMe)
  2. Pin processes to NUMA nodes if you have multiple sockets
  3. Verify RAID configuration matches your read/write pattern
  4. Increase network ring buffers and kernel socket buffers
  5. Disable or relocate swap
  6. Set CPU governor to performance for latency-sensitive apps
  7. Disable transparent huge pages if running databases
  8. Switch to BBR congestion control for datacenter or WAN links
  9. Monitor hardware RAID cache status

Each mistake on this list will cost you throughput, latency, or both. The defaults are designed for general-purpose use, not for extracting maximum performance from dedicated hardware.

The good news? Once you fix these, they stay fixed. Make these changes part of your standard provisioning process and every new bare metal server starts optimized.

Common questions

Do these fixes apply to virtual machines too?

Some do, some don't. NUMA awareness, TCP tuning, and THP settings carry over to VMs. RAID configuration and hardware controller tuning don't — those are the host's problem. I/O scheduler choice matters less in VMs because the hypervisor adds its own scheduling layer.

Will performance governor increase my power bill significantly?

Yes, but the increase is usually 10-20% depending on your CPU model and typical load. For revenue-generating production systems, the cost is negligible compared to the value of consistent performance. Save power on dev/staging servers instead.

How do I know which mistake is affecting me?

Start with monitoring. Install sysstat and collect data with sar over several days. Look for I/O wait, network drops, NUMA misses, and swap activity. Each metric points to a specific problem on this list.

Can I just use a tuning profile like tuned-adm?

Tuned profiles are a good starting point and they'll catch some of these issues. But they're generic. You still need to verify RAID choice, NUMA pinning for your specific application, and hardware cache status manually. Tuned won't know your workload pattern.

Start with the I/O path

If you only fix one thing from this list, fix your I/O scheduler and RAID configuration. Storage bottlenecks hurt every application — databases, web servers, mail systems, everything touches disk.

The other fixes matter, but bad storage configuration will tank performance no matter how well you tune the network and CPU. Check your scheduler, confirm your RAID level makes sense, and verify the hardware controller cache is working. Those three alone will solve most bare metal performance complaints I see in support tickets.