Skip to content
Back to Blog
Performance11 min read

9 Bare Metal Server Performance Mistakes [Solved]

Discover the most common bare metal server hosting performance mistakes that slow down your infrastructure and learn the correct approach to fix each one.

Written by Abdul AbrorTechnical Hosting Support Engineer
9 Bare Metal Server Performance Mistakes [Solved]
On this page

Bare metal servers offer raw performance that virtualized environments can't match. But I've seen countless deployments that never reach their potential because of easily avoidable mistakes.

The gap between theoretical and actual performance often comes down to configuration choices made during setup or left at defaults that made sense for a different workload. Here are the nine mistakes that show up most often in performance audits, along with what you should do instead.

Running default kernel parameters

Most Linux distributions ship with kernel parameters tuned for general desktop or light server use. These defaults assume modest network traffic, small file operations, and limited concurrent connections.

Your bare metal server handling thousands of requests per second needs different settings. The net.core.somaxconn parameter typically defaults to 128, meaning your server will only queue 128 pending connections before dropping new ones. Under load, legitimate traffic gets rejected.

The correct approach: Tune kernel parameters for your workload. For web servers handling high traffic:

# /etc/sysctl.conf or /etc/sysctl.d/99-custom.conf
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 8192
net.ipv4.ip_local_port_range = 1024 65535
net.ipv4.tcp_tw_reuse = 1
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216

Apply with sysctl -p. Monitor connection drops in /proc/net/netstat and adjust based on your actual traffic patterns. Database servers need different tuning focused on shared memory and file handles instead.

Ignoring NUMA architecture

Modern multi-socket servers use Non-Uniform Memory Access. CPU socket zero has faster access to RAM bank zero than to RAM bank one. When your application spawns threads across sockets but allocates memory from a distant bank, every memory access crosses the interconnect.

I've diagnosed production databases running 40% slower than expected because MySQL was hitting remote NUMA nodes for every query. The CPUs spent cycles waiting for memory instead of processing data.

The correct approach: Pin processes to NUMA nodes that match their memory allocation.

Check your topology first:

numactl --hardware
lscpu | grep NUMA

For services like databases, bind both CPU and memory to the same node:

numactl --cpunodebind=0 --membind=0 mysqld

Containerized workloads should set CPU sets that respect NUMA boundaries. Your scheduler (systemd, Docker, Kubernetes) should know about the topology. Some applications have their own NUMA awareness settings you can enable in config files.

Using the wrong filesystem for your workload

Ext4 is reliable and well-tested. It's also the wrong choice for many bare metal workloads.

If you're running a database with millions of small files, ext4's directory indexing becomes a bottleneck. For large sequential I/O like video encoding or log aggregation, XFS provides better throughput. Copy-on-write filesystems like Btrfs offer snapshot capabilities that can eliminate backup windows entirely.

The correct approach: Match filesystem features to your I/O patterns. Database servers with lots of small random writes perform better on XFS with specific mount options:

# /etc/fstab
/dev/sda1 /var/lib/mysql xfs noatime,nodiratime,logbufs=8,logbsize=256k 0 0

The noatime flag alone can improve performance by 10-15% for read-heavy workloads because the kernel stops writing access timestamps on every file read. Test with realistic data before committing to production.

Leaving CPU governor on powersave

Server distributions often ship with the CPU frequency governor set to powersave or ondemand to reduce power consumption. These governors scale CPU frequency based on load.

The problem? Frequency scaling introduces latency. When a request arrives, the CPU is running at minimum frequency. By the time it ramps up to full speed, you've already added milliseconds to response time. Over thousands of requests, this compounds into noticeably slower performance.

The correct approach: Set the governor to performance mode for consistent low latency.

# Check current governor
cat /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor

# Set to performance
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor

Make it permanent by installing cpufrequtils and setting GOVERNOR="performance" in /etc/default/cpufrequtils. Yes, power consumption increases. That's the trade-off you make for bare metal performance.

Running without I/O scheduling tuning

The Linux I/O scheduler decides the order in which disk requests are processed. The default scheduler varies by distribution and might not suit your storage hardware.

For NVMe drives, the none scheduler (also called noop) performs best because the drive's internal parallelism exceeds what software scheduling can optimize. For rotational drives, deadline provides better latency than the older cfq scheduler. Using the wrong scheduler leaves performance on the table.

The correct approach: Match the I/O scheduler to your storage type.

Check current schedulers:

cat /sys/block/sda/queue/scheduler
cat /sys/block/nvme0n1/queue/scheduler

For NVMe:

echo none > /sys/block/nvme0n1/queue/scheduler

For SSD:

echo deadline > /sys/block/sda/queue/scheduler

Persist the setting through udev rules or kernel boot parameters. Test with fio benchmarks before and after to measure the actual impact on your workload.

Not configuring IRQ affinity

Network cards with multiple queues distribute incoming packets across CPU cores through interrupt requests. Without proper configuration, all IRQs might fire on CPU zero, creating a bottleneck while other cores sit idle.

This is especially problematic on 10GbE and faster networks where a single core can't process packets quickly enough. You'll see high interrupt counts on one CPU while network throughput remains limited despite available bandwidth.

The correct approach: Spread network IRQs across cores.

First, identify your network interface's IRQ numbers:

grep eth0 /proc/interrupts

Then set IRQ affinity to specific cores:

echo 1 > /proc/irq/85/smp_affinity  # Core 0
echo 2 > /proc/irq/86/smp_affinity  # Core 1
echo 4 > /proc/irq/87/smp_affinity  # Core 2

The values are bitmasks in hex. Most distributions include irqbalance, but it makes generic choices. Manual tuning based on your application's CPU usage gives better results. Pin network IRQs to CPUs that aren't handling your application's main threads.

Allocating insufficient hugepages

Applications that allocate large amounts of memory (databases, in-memory caches, scientific computing) spend CPU cycles managing page tables. With standard 4KB pages, a 64GB allocation requires millions of page table entries.

Hugepages (2MB or 1GB each) reduce this overhead dramatically. A database that needs 32GB of buffer pool can use 16,000 hugepages instead of 8 million standard pages. Page table walks complete faster and TLB cache misses drop.

The correct approach: Pre-allocate hugepages for memory-intensive applications.

# Check current allocation
grep Huge /proc/meminfo

# Allocate 2MB hugepages (number = memory needed / 2MB)
echo 16000 > /proc/sys/vm/nr_hugepages

# Make permanent in /etc/sysctl.conf
vm.nr_hugepages = 16000

Configure your application to use them. For MySQL:

[mysqld]
large-pages = ON

For Redis:

vm.overcommit_memory = 1
transparent_hugepage=never

Transparent Hugepages sound convenient but introduce latency spikes during page allocation. Disable them for latency-sensitive services.

Over-provisioning swap on servers with plenty of RAM

The old rule of thumb was to create swap space equal to or double your RAM. On a bare metal server with 256GB of memory running a specific workload, that's pointless.

Worse, it's dangerous. If your application starts swapping with that much RAM available, something is seriously wrong. Swap thrashing will make the problem worse, not better. The server becomes unresponsive while the kernel moves gigabytes between RAM and disk.

The correct approach: Size swap based on your actual needs or eliminate it entirely.

For servers running services that should never swap (databases, caches, real-time processing), disable swap completely:

swapoff -a
# Remove swap entries from /etc/fstab

If you need swap for hibernate or specific use cases, limit it:

vm.swappiness = 1

This tells the kernel to avoid swapping unless absolutely necessary. Monitor memory pressure with sar -r or vmstat. If you're actually running out of RAM regularly, you need more RAM or fewer services, not more swap.

Missing storage firmware and driver updates

Controller firmware and driver versions directly affect I/O performance and stability. A two-year-old RAID controller firmware might have bugs that cause periodic stalls or operate at reduced queue depth.

I've seen storage performance double after updating NVMe firmware that fixed a bug causing excessive write amplification. The drives were working, so nobody thought to check for updates. They were also performing at half their rated speed.

The correct approach: Maintain current firmware for all storage components.

For NVMe drives:

# Check current firmware
nvme list
nvme id-ctrl /dev/nvme0 | grep fr

# Update (vendor-specific tool required)
# Example for Intel:
nvme fw-download /dev/nvme0 --fw=firmware.bin
nvme fw-commit /dev/nvme0 --slot=0 --action=1

For RAID controllers, use the vendor's management utility. LSI/Broadcom controllers use StorCLI, Dell uses PERCCLI, HPE has ssacli. Check for updates quarterly and review release notes for performance improvements or critical fixes.

Driver updates matter too. In-kernel drivers are usually fine, but vendor-supplied drivers sometimes offer better performance or newer features. Test updates in staging first because bad firmware can render devices unusable.

What to check first

When performance isn't meeting expectations, start with the low-hanging fruit. Check your CPU governor setting and kernel network parameters because these are often wrong and easy to fix.

Run basic monitoring to identify your actual bottleneck before making changes. Tuning the I/O scheduler won't help if you're CPU-bound. Adding hugepages won't matter if your application doesn't allocate enough memory to benefit.

Make one change at a time and measure the result. Keep notes on what you changed and why. Some of these tunings interact with each other in non-obvious ways, and you'll want a way to backtrack if something makes performance worse instead of better.

The best bare metal performance comes from understanding your specific workload and configuring the hardware to support it efficiently. These nine mistakes are a starting point, not the entire picture.

FAQ

How do I know which mistakes are affecting my server?

Start with top, iostat, and sar to identify whether you're CPU-bound, I/O-bound, or memory-bound. High iowait points to storage issues. High softirq suggests network IRQ problems. CPU frequency stuck below maximum indicates governor misconfiguration.

Can I fix all of these on a live production server?

Most kernel parameter changes take effect immediately without restart. CPU governor, IRQ affinity, and hugepages can be changed live. Filesystem changes require unmounting. Firmware updates often need a reboot. Test each change on staging first.

Will these tunings help a virtualized environment?

Some will, some won't. You can't tune hardware firmware or IRQ affinity on a VM. Kernel parameters, filesystem choices, and CPU governor settings still apply. Hugepages help but are usually pre-configured by the hypervisor.

How much performance gain should I expect?

It depends entirely on which bottleneck you're addressing and how severe it was. Fixing a severe bottleneck (like all network IRQs on one CPU at 10GbE speeds) can double throughput. Tuning an already well-configured system might gain 5-10%. Measure before and after.

Should I hire someone or learn this myself?

If you're managing bare metal infrastructure long-term, invest time learning the fundamentals. One-time optimizations or complex NUMA tuning might justify hiring an expert. Start with the easy wins (CPU governor, basic sysctl tuning) and work up to the harder stuff.