Most Docker security guides stop at user namespaces and read-only filesystems. If you're running production containers on a VPS and already understand the fundamentals, you need strategies that address privilege escalation vectors, kernel attack surface, supply chain integrity, and the performance cost of defense-in-depth. This guide covers advanced hardening techniques that balance security posture with operational reality.
Runtime Security Beyond User Namespaces
User namespaces remap container root to an unprivileged host user, but they don't address all privilege escalation paths. Combine them with additional kernel-level controls.
Seccomp Profiles Tuned to Your Workload
The default Docker seccomp profile blocks around 44 of 300+ syscalls. For production workloads, audit which syscalls your application actually uses and build a custom profile that denies everything else.
Generate a baseline by running your container with strace in a test environment:
strace -c -f -S calls docker run --rm your-image:tag
Convert the output into a seccomp whitelist. Start with Docker's default profile as a template and remove allowed syscalls your app never invokes. Pay special attention to blocking:
ptraceandprocess_vm_readv(process inspection)perf_event_open(kernel profiling)bpf(eBPF program loading)keyctlandadd_key(kernel keyring)mount,umount2,pivot_root(filesystem manipulation)
Apply the profile at runtime:
docker run --security-opt seccomp=/path/to/custom-profile.json your-image:tag
For Kubernetes, specify the profile in the pod's securityContext.seccompProfile field.
AppArmor or SELinux Mandatory Access Control
Seccomp filters syscalls; MAC systems like AppArmor and SELinux control file access, network operations, and capabilities at a finer granularity.
Create an AppArmor profile that restricts container filesystem access to only the paths your application needs:
#include <tunables/global>
profile docker-restricted flags=(attach_disconnected,mediate_deleted) {
#include <abstractions/base>
network inet stream,
network inet6 stream,
deny /proc/sys/kernel/** wklx,
deny /sys/kernel/security/** rwklx,
deny @{PROC}/kcore r,
/app/** r,
/app/data/** rw,
/tmp/** rw,
deny /** w,
}
Load and enforce it:
apparmor_parser -r -W /etc/apparmor.d/docker-restricted
docker run --security-opt apparmor=docker-restricted your-image:tag
SELinux offers similar control through type enforcement. On RHEL-based systems, use udica to generate custom policies from running containers.
Rootless Docker Architecture
Rootless mode runs the Docker daemon itself as an unprivileged user, eliminating the largest privilege escalation surface. Even if an attacker breaks out of the container, they land in an unprivileged user session, not host root.
Installation and Caveats
Install rootless Docker using the upstream script:
curl -fsSL https://get.docker.com/rootless | sh
Add the daemon to your user session:
systemctl --user enable docker
systemctl --user start docker
export DOCKER_HOST=unix://$XDG_RUNTIME_DIR/docker.sock
Limitations you'll encounter:
- Cannot bind to privileged ports below 1024 without
sysctl net.ipv4.ip_unprivileged_port_start=80 overlay2storage requires kernel 5.11+ with unprivileged overlayfs, otherwise falls back tofuse-overlayfswith a performance penalty- cgroups v2 required for resource limits; cgroups v1 will not enforce CPU or memory constraints
- AppArmor and SELinux profiles must be adapted for the unprivileged daemon
For production VPS environments, verify your kernel supports unprivileged user namespaces:
sysctl kernel.unprivileged_userns_clone
If disabled, you'll need to enable it or work with your hosting provider.
Capability Dropping and Ambient Sets
Linux capabilities partition root privileges into discrete units. Docker drops several by default, but many containers still run with more capabilities than necessary.
Audit and Minimize Capability Sets
List the capabilities your container currently has:
docker run --rm your-image:tag sh -c 'apk add -q libcap; capsh --print'
Drop all capabilities and add back only what's required:
docker run --cap-drop=ALL \
--cap-add=NET_BIND_SERVICE \
--cap-add=CHOWN \
your-image:tag
Common capabilities to audit:
NET_ADMIN: required for network configuration; avoid unless running network toolsSYS_ADMIN: extraordinarily broad; almost never needed in application containersDAC_OVERRIDE: bypasses file permission checks; often unnecessary if file ownership is correctSETUID/SETGID: needed only if your app changes user context at runtime
For orchestration, specify in your compose file:
services:
app:
cap_drop:
- ALL
cap_add:
- NET_BIND_SERVICE
No New Privileges Flag
Prevent processes inside the container from gaining additional privileges via setuid binaries:
docker run --security-opt=no-new-privileges:true your-image:tag
This blocks privilege escalation through binaries like sudo, su, or custom setuid executables even if they exist in the container image.
Immutable Infrastructure Patterns
Treat containers as immutable units. Write operations should target ephemeral storage or external volumes, never the container's root filesystem.
Read-Only Root Filesystem with Targeted Tmpfs
Mount the container filesystem as read-only and provide writable tmpfs mounts only where needed:
docker run --read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=64m \
--tmpfs /var/run:rw,noexec,nosuid,size=16m \
your-image:tag
This prevents malware from persisting changes to the image or dropping executables in writable directories. The noexec flag on tmpfs mounts blocks execution of anything written there.
For applications that must write persistent data, use named volumes with restricted permissions:
docker run --read-only \
-v app-data:/app/data:rw \
--tmpfs /tmp:rw,noexec,nosuid \
your-image:tag
Resource Quotas and Fork Bombs
Resource exhaustion attacks can destabilize the host. Enforce hard limits on CPU, memory, and process count.
PIDs Limit to Prevent Fork Bombs
Without a PID limit, a compromised container can spawn processes until the host kernel's PID table is exhausted:
docker run --pids-limit=256 your-image:tag
Choose a limit based on your application's legitimate process count. A typical web application with a few worker processes needs far fewer than 100 PIDs.
Memory and CPU Constraints with Swap Disabled
Memory limits without swap prevent containers from using host swap space, which can cause severe performance degradation:
docker run -m 512m --memory-swap=512m \
--cpus=1.5 \
--cpu-shares=512 \
your-image:tag
Set --memory-swap equal to -m to disable swap for the container. Use --cpus for hard limits and --cpu-shares for relative weighting when contention occurs.
Network Segmentation and Egress Filtering
Default bridge networks allow unrestricted inter-container communication. Isolate workloads and control egress.
Custom Networks with Internal Flag
Create networks that cannot route to the host or external internet:
docker network create --internal backend-net
docker run --network=backend-net db-container:tag
Front-end services can connect to a separate network with external access, while backend services remain isolated.
Explicit Egress Rules with iptables
Even if your container doesn't need internet access for normal operation, vulnerabilities or supply chain compromises might attempt to establish outbound connections. Use iptables to whitelist only necessary egress:
iptables -A DOCKER-USER -s 172.18.0.0/16 -d 0.0.0.0/0 -j DROP
iptables -I DOCKER-USER -s 172.18.0.0/16 -d your-allowed-api.com -p tcp --dport 443 -j ACCEPT
The DOCKER-USER chain persists across Docker daemon restarts and takes precedence over Docker's own rules.
Image Supply Chain and Content Trust
Hardening the runtime is incomplete if you're pulling compromised images.
Content Trust and Signature Verification
Enable Docker Content Trust to require signed images:
export DOCKER_CONTENT_TRUST=1
docker pull your-registry.com/signed-image:tag
This enforces that images are signed with Notary keys. For private registries, configure your own Notary server and signing workflow.
Multi-Stage Builds and Distroless Base Images
Minimize attack surface by removing build tools, package managers, and shells from the final image:
FROM golang:1.21 AS builder
WORKDIR /build
COPY . .
RUN go build -o app .
FROM gcr.io/distroless/base-debian12
COPY --from=builder /build/app /app
USER nonroot:nonroot
ENTRYPOINT ["/app"]
Distroless images contain only your application and runtime dependencies, no shell, no package manager. This eliminates entire classes of privilege escalation techniques.
Monitoring and Runtime Detection
Hardening reduces risk but doesn't eliminate it. Detect anomalous behavior in real time.
Syscall Auditing with Falco
Falco monitors syscalls against a rule set and alerts on suspicious activity. Install on the VPS host and point it at your containers:
helm install falco falcosecurity/falco \
--set driver.kind=modern_ebpf \
--set falco.grpc.enabled=true
Create custom rules for your workload:
- rule: Unexpected Outbound Connection
desc: Detect outbound connections to unexpected IPs
condition: outbound and not fd.sip in (allowed_ips)
output: "Suspicious outbound connection (container=%container.name ip=%fd.sip)"
priority: WARNING
Integrate alerts with your monitoring stack to trigger incident response workflows.
Log Immutability and Centralization
Containers are ephemeral; logs must be shipped off-host immediately. Use a logging driver that forwards to a remote syslog or SIEM:
docker run --log-driver=syslog \
--log-opt syslog-address=tcp://log-aggregator.internal:514 \
--log-opt tag="{{.Name}}/{{.ID}}" \
your-image:tag
Centralized logs survive container destruction and can't be tampered with by a compromised container.
Kernel Hardening for Container Hosts
Your VPS kernel configuration directly impacts container security.
Sysctl Tuning
Harden the host kernel to limit information leakage and restrict unprivileged operations:
sysctl -w kernel.dmesg_restrict=1
sysctl -w kernel.kptr_restrict=2
sysctl -w kernel.unprivileged_bpf_disabled=1
sysctl -w net.ipv4.conf.all.rp_filter=1
sysctl -w net.ipv4.conf.default.rp_filter=1
Persist these in /etc/sysctl.d/99-container-hardening.conf.
Disable Unused Kernel Modules
If your containers don't need certain protocols or filesystems, blacklist the modules:
echo "install dccp /bin/true" >> /etc/modprobe.d/disable-uncommon.conf
echo "install sctp /bin/true" >> /etc/modprobe.d/disable-uncommon.conf
echo "install rds /bin/true" >> /etc/modprobe.d/disable-uncommon.conf
Reducing the loaded kernel modules shrinks the attack surface available to containers.
Conclusion
Advanced Docker hardening on a VPS requires layering kernel controls, runtime policies, network segmentation, and supply chain verification. No single technique provides complete security, but combining seccomp, AppArmor, rootless mode, capability dropping, and runtime monitoring creates defense-in-depth that raises the cost of exploitation. Audit your current posture, implement these controls incrementally, and validate them under load. The performance cost of proper hardening is minimal compared to the operational cost of a compromised container environment.
