Running Docker containers in production VPS environments demands more than disabling privileged mode and running rootless containers. This guide targets experienced operators who have already implemented baseline hardening and need to push their container security posture to the next level. We'll cover syscall filtering, advanced capability management, runtime monitoring, and orchestration-level controls that separate hardened deployments from merely configured ones.
Kernel-Level Isolation with Seccomp Profiles
Seccomp (secure computing mode) restricts the system calls a container can make. Docker ships with a default seccomp profile that blocks roughly 44 of the 300+ available syscalls, but custom profiles let you lock down containers to their exact runtime requirements.
Building Custom Seccomp Profiles
Start by auditing what your application actually needs. Run your container with seccomp disabled and trace syscalls:
strace -c -f -S name docker run --rm --security-opt seccomp=unconfined your-image:tag
Analyze the output and build a whitelist-based profile. A production-grade seccomp profile for a typical Node.js API might look like this:
{
"defaultAction": "SCMP_ACT_ERRNO",
"architectures": [
"SCMP_ARCH_X86_64",
"SCMP_ARCH_X86",
"SCMP_ARCH_X32"
],
"syscalls": [
{
"names": [
"accept4",
"bind",
"brk",
"clone",
"close",
"connect",
"epoll_create1",
"epoll_ctl",
"epoll_wait",
"exit_group",
"fcntl",
"fstat",
"futex",
"getpid",
"getsockname",
"listen",
"mmap",
"munmap",
"openat",
"poll",
"read",
"recvfrom",
"rt_sigaction",
"sendto",
"socket",
"write"
],
"action": "SCMP_ACT_ALLOW"
}
]
}
Apply with --security-opt seccomp=/path/to/profile.json. Test exhaustively in staging. One missing syscall will crash your application in ways that are difficult to debug.
Edge Cases and Compatibility
Different base images and runtimes need different syscalls. Alpine-based images require fewer syscalls than Debian. Interpreted languages need more runtime syscalls than compiled binaries. Always version your seccomp profiles alongside your Dockerfiles and test across kernel versions.
Capability Dropping Beyond Defaults
Linux capabilities divide root privileges into 40+ distinct units. Docker drops several by default, but production workloads should drop everything not explicitly required.
Auditing and Minimal Capability Sets
Most web applications need zero capabilities. Start with --cap-drop=ALL:
docker run --cap-drop=ALL \
--security-opt=no-new-privileges:true \
your-image:tag
If your application binds to privileged ports (below 1024), add only CAP_NET_BIND_SERVICE:
docker run --cap-drop=ALL \
--cap-add=NET_BIND_SERVICE \
--security-opt=no-new-privileges:true \
-p 443:443 your-image:tag
For applications that need raw sockets (ping utilities, network diagnostics), add CAP_NET_RAW. Never add CAP_SYS_ADMIN or CAP_SYS_MODULE in production.
No-New-Privileges Flag
The no-new-privileges security option prevents privilege escalation via setuid binaries or file capabilities inside the container. This blocks a common container escape vector. Always combine it with capability dropping.
Mandatory Access Control with AppArmor and SELinux
Seccomp filters syscalls; MAC systems like AppArmor and SELinux define what processes can access. Use them together for defense in depth.
AppArmor Profile Strategy
Docker generates a default AppArmor profile (docker-default) that's adequate for generic workloads. Custom profiles let you restrict file access, network operations, and process capabilities at a granular level.
Create an AppArmor profile for a read-only application container:
#include <tunables/global>
profile docker-readonly-app flags=(attach_disconnected,mediate_deleted) {
#include <abstractions/base>
network inet tcp,
network inet udp,
deny /bin/** wl,
deny /boot/** rwlx,
deny /dev/mem rwlx,
deny /sys/** wl,
/app/** r,
/usr/lib/** r,
/lib/** r,
/etc/ssl/** r,
deny /tmp/** wl,
owner /tmp/** rw,
}
Load and apply:
apparmor_parser -r -W /etc/apparmor.d/docker-readonly-app
docker run --security-opt apparmor=docker-readonly-app your-image:tag
SELinux Type Enforcement
On RHEL-based systems, use SELinux contexts. Run containers in svirt_sandbox_file_t context and mount volumes with appropriate labels:
docker run -v /host/data:/data:z \
--security-opt label=type:svirt_apache_t \
your-image:tag
The :z flag relabels the volume to allow the container to access it. The label=type option assigns a custom SELinux type. Verify with ps -eZ | grep docker.
Runtime Resource Constraints and Isolation
Resource limits prevent noisy neighbor problems and limit blast radius during compromise.
Cgroup v2 and Memory Management
Modern kernels use cgroup v2. Set hard memory limits, swap limits, and OOM score adjustments:
docker run \
--memory="512m" \
--memory-swap="512m" \
--memory-reservation="256m" \
--oom-score-adj=500 \
your-image:tag
The memory-swap equal to memory disables swap usage entirely. The memory-reservation sets a soft limit that allows burstable workloads. The oom-score-adj makes the container more likely to be killed than the host processes during OOM conditions.
CPU and I/O Quotas
Pin containers to specific CPU sets and limit I/O throughput:
docker run \
--cpuset-cpus="0-1" \
--cpu-shares=512 \
--blkio-weight=300 \
--device-read-bps /dev/sda:10mb \
your-image:tag
CPU shares are relative weights. I/O weight ranges from 10 to 1000. Device-specific I/O limits prevent one container from saturating disk bandwidth.
PID Limits
Limit the number of processes a container can spawn to prevent fork bombs:
docker run --pids-limit=100 your-image:tag
Choose values based on your application's actual concurrency model. Web servers with worker pools need higher limits than single-threaded applications.
User Namespace Remapping
User namespaces remap container UIDs to unprivileged host UIDs. A process running as root (UID 0) inside the container maps to a non-root UID on the host.
Enable daemon-wide user namespace remapping in /etc/docker/daemon.json:
{
"userns-remap": "default"
}
Restart Docker. This creates a dockremap user and remaps all container UIDs. A container process with UID 0 runs as UID 65536 on the host.
For per-container control, specify custom mappings:
docker run --userns=host your-image:tag
Note that user namespace remapping breaks volume permissions and some capabilities. Test thoroughly.
Read-Only Root Filesystems and tmpfs
Run containers with immutable root filesystems and explicit temporary storage:
docker run --read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=100m \
--tmpfs /var/run:rw,noexec,nosuid,size=50m \
your-image:tag
The --read-only flag prevents writing anywhere except explicitly mounted volumes and tmpfs. Mount tmpfs with noexec to prevent running binaries from temporary storage.
For applications that need persistent writable directories, mount named volumes:
docker run --read-only \
--tmpfs /tmp:noexec,nosuid,size=100m \
-v app-cache:/var/cache/app:rw \
your-image:tag
Network Segmentation and Firewall Rules
Default bridge networks expose containers to each other. Use custom networks with explicit links or overlay networks with encryption.
Custom Bridge Networks
Create isolated networks per application stack:
docker network create \
--driver bridge \
--subnet=172.25.0.0/16 \
--ip-range=172.25.1.0/24 \
--gateway=172.25.1.1 \
--opt "com.docker.network.bridge.name"="br-prod" \
prod-network
Attach containers selectively:
docker run --network=prod-network your-image:tag
Host Firewall Integration
Docker manipulates iptables directly. Lock down the DOCKER-USER chain to control inter-container and external traffic:
iptables -I DOCKER-USER -i br-prod ! -s 172.25.1.0/24 -j DROP
iptables -I DOCKER-USER -i br-prod -o br-prod -j ACCEPT
iptables -I DOCKER-USER -i br-prod -p tcp --dport 443 -j ACCEPT
Persist rules with iptables-persistent or your distribution's firewall management tool.
Runtime Security Monitoring
Hardening is incomplete without runtime visibility. Monitor container behavior for anomalies.
Auditd for Container Events
Configure auditd to log container syscalls:
auditctl -a exit,always -F arch=b64 -S execve -F key=docker-exec
auditctl -a exit,always -F arch=b64 -S open -S openat -F key=docker-file-access
Query logs with ausearch:
ausearch -k docker-exec -ts recent
File Integrity Monitoring
Run AIDE or Tripwire on the host to detect changes to Docker binaries, daemon configuration, and image storage:
aide --init
cp /var/lib/aide/aide.db.new.gz /var/lib/aide/aide.db.gz
aide --check
Schedule regular checks via cron.
Image Scanning and Supply Chain Security
Build-time hardening starts with trusted base images and regular vulnerability scanning.
Multi-Stage Builds for Minimal Attack Surface
Use multi-stage builds to exclude build tools from production images:
FROM golang:alpine AS builder
WORKDIR /build
COPY . .
RUN go build -o app
FROM alpine:latest
RUN apk --no-cache add ca-certificates
COPY --from=builder /build/app /usr/local/bin/app
USER 1000:1000
ENTRYPOINT ["/usr/local/bin/app"]
This pattern reduces image size and eliminates compiler toolchains from runtime.
Automated Scanning Pipelines
Integrate scanners into CI/CD pipelines. Run scans on every build and block deployment on high-severity findings. Store scan results and track remediation.
Orchestration-Level Controls
In orchestrated environments, enforce policies at the cluster level.
Pod Security Standards
For Kubernetes users, enforce restricted pod security standards:
apiVersion: v1
kind: Namespace
metadata:
name: production
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/audit: restricted
pod-security.kubernetes.io/warn: restricted
This blocks privileged containers, host namespaces, and insecure capabilities cluster-wide.
Policy Engines
Deploy OPA (Open Policy Agent) or Kyverno to enforce custom admission control policies. Example Kyverno policy requiring read-only root filesystems:
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-ro-rootfs
spec:
validationFailureAction: enforce
rules:
- name: check-readOnlyRootFilesystem
match:
resources:
kinds:
- Pod
validate:
message: "Root filesystem must be read-only"
pattern:
spec:
containers:
- securityContext:
readOnlyRootFilesystem: true
Conclusion
Advanced container hardening is a layered approach. Seccomp profiles constrain syscalls, capabilities limit privileges, MAC systems enforce access controls, resource limits contain blast radius, and runtime monitoring detects anomalies. Each layer addresses different attack vectors. Start with the controls that map to your actual threat model, test exhaustively in staging, and iterate. The goal is not perfect security but raising the cost of exploitation high enough that attackers move to softer targets. Combine these runtime protections with secure build practices, regular patching, and defense-in-depth networking for production-grade container security on your VPS infrastructure.
