Skip to content
Back to Blog
Linux & Server11 min read

Dokploy vs Coolify vs CapRover: Advanced Tuning & Edge Cases

Deep dive into production-grade optimization, security hardening, and performance tuning for Dokploy, Coolify, and CapRover deployments beyond the basic setup guides.

Written by Abdul AbrorTechnical Hosting Support Engineer
Dokploy vs Coolify vs CapRover: Advanced Tuning & Edge Cases
On this page

You've installed Dokploy, Coolify, or CapRover and deployed your first application. Now you need production-grade performance, security hardening, and strategies for edge cases that documentation rarely covers. This guide cuts past the setup tutorials and focuses on optimization, troubleshooting, and architectural decisions that separate hobby projects from resilient production deployments.

Architecture Patterns That Matter

Reverse Proxy Layer Configuration

All three platforms use Traefik (Dokploy, Coolify) or Nginx (CapRover) as the reverse proxy layer. The default configurations prioritize convenience over performance. For production workloads, tune these aggressively.

Traefik connection limits in Dokploy and Coolify should match your expected concurrent user load. Edit the Traefik static configuration:

entryPoints:
  web:
    address: ":80"
    transport:
      respondingTimeouts:
        readTimeout: 60s
        writeTimeout: 60s
        idleTimeout: 180s
  websecure:
    address: ":443"
    transport:
      respondingTimeouts:
        readTimeout: 60s
        writeTimeout: 60s
        idleTimeout: 180s
      lifeCycle:
        gracePeriod: 30s

For CapRover's Nginx, increase worker connections and enable HTTP/2 push:

events {
    worker_connections 4096;
    use epoll;
    multi_accept on;
}

http {
    http2_push_preload on;
    keepalive_timeout 65;
    keepalive_requests 100;
}

Access CapRover's Nginx config through the CLI or by mounting /etc/nginx/nginx.conf as a volume in the captain-nginx container.

Database Connection Pooling

Each platform handles service communication differently. Dokploy and Coolify use Docker networks with automatic service discovery. CapRover uses a captain-overlay network. Your database connection strings matter.

Wrong approach: letting application containers open individual connections.

Right approach: deploy PgBouncer or ProxySQL as an intermediate service.

Dokploy example for PostgreSQL with PgBouncer:

[databases]
app_db = host=postgres-service port=5432 dbname=production

[pgbouncer]
listen_addr = 0.0.0.0
listen_port = 6432
auth_type = md5
auth_file = /etc/pgbouncer/userlist.txt
pool_mode = transaction
max_client_conn = 1000
default_pool_size = 25
reserve_pool_size = 5
reserve_pool_timeout = 3
server_idle_timeout = 600

Deploy PgBouncer as a separate service in the same project, then point application containers to pgbouncer-service:6432 instead of the database directly. This pattern reduces PostgreSQL connection overhead and prevents exhaustion under burst traffic.

Storage and Volume Management

Persistent Volume Performance

All three platforms mount Docker volumes by default. For databases and high-I/O workloads, this creates a performance bottleneck on most cloud providers.

Bind mounts deliver better IOPS when the underlying host uses SSD storage:

volumes:
  - /mnt/fast-storage/postgres-data:/var/lib/postgresql/data

On cloud VPS instances, provision separate block storage volumes with provisioned IOPS, format them with ext4 or XFS, mount them at a dedicated path, and bind-mount that path into your containers.

Coolify-specific: use the volume override in the service configuration UI. Navigate to the service, expand Advanced, and replace the named volume with an absolute host path.

CapRover-specific: use persistent directories via the app definition:

{
  "volumes": [
    {
      "hostPath": "/mnt/fast-storage/app-data",
      "containerPath": "/app/data"
    }
  ]
}

Backup Strategy for Multi-Node Setups

Dokploy and Coolify are single-node by design. CapRover supports Docker Swarm, enabling multi-node deployments. In Swarm mode, volumes do not replicate automatically.

Implement a sidecar backup container that runs on the same node as your database:

services:
  postgres-backup:
    image: postgres:latest
    command: |
      bash -c 'while true; do
        pg_dump -h postgres-service -U user -d production > /backup/db-$$(date +%Y%m%d-%H%M%S).sql
        find /backup -name "db-*.sql" -mtime +7 -delete
        sleep 86400
      done'
    volumes:
      - /mnt/backups:/backup
    deploy:
      placement:
        constraints:
          - node.labels.database == true

Tag your database node with docker node update --label-add database=true <node-id> to ensure both services stay co-located.

SSL/TLS Hardening

Beyond Let's Encrypt Defaults

All three platforms automate Let's Encrypt certificate provisioning. The default cipher suites and TLS versions are permissive for compatibility. Tighten them for security-conscious deployments.

Traefik TLS options (Dokploy, Coolify):

tls:
  options:
    default:
      minVersion: VersionTLS12
      cipherSuites:
        - TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256
        - TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384
        - TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305
      curvePreferences:
        - CurveP521
        - CurveP384
      sniStrict: true

Apply this by creating a ConfigMap or modifying Traefik's dynamic configuration file, usually stored at /data/traefik/dynamic.yml.

CapRover Nginx SSL tuning:

ssl_protocols TLSv1.2 TLSv1.3;
ssl_ciphers 'ECDHE-RSA-AES128-GCM-SHA256:ECDHE-RSA-AES256-GCM-SHA384:ECDHE-RSA-CHACHA20-POLY1305';
ssl_prefer_server_ciphers on;
ssl_session_cache shared:SSL:10m;
ssl_session_timeout 10m;
ssl_stapling on;
ssl_stapling_verify on;

Wildcard Certificates and DNS Validation

For multi-tenant or subdomain-heavy deployments, wildcard certificates reduce renewal load. All three platforms support DNS-01 challenge, but configuration differs.

Dokploy and Coolify use Traefik's certificate resolvers. Add your DNS provider credentials as environment variables:

environment:
  - [email protected]
  - CF_API_KEY=your-cloudflare-api-key

Then configure the certificate resolver:

certificatesResolvers:
  cloudflare:
    acme:
      email: [email protected]
      storage: /data/acme.json
      dnsChallenge:
        provider: cloudflare
        delayBeforeCheck: 30

CapRover requires manual wildcard certificate upload via the web UI or CLI. Generate the certificate locally using certbot with DNS plugin:

certbot certonly --dns-cloudflare \
  --dns-cloudflare-credentials ~/.secrets/cloudflare.ini \
  -d '*.yourdomain.com' -d yourdomain.com

Upload the resulting certificate files through CapRover's SSL management interface.

Resource Limits and Autoscaling

Container Resource Constraints

Default deployments allow containers to consume unlimited host resources. This causes noisy neighbor problems in multi-tenant setups and OOM kills under load.

Set explicit limits for every service:

deploy:
  resources:
    limits:
      cpus: '2.0'
      memory: 2G
    reservations:
      cpus: '0.5'
      memory: 512M

Dokploy exposes resource limits in the UI under service settings. Coolify requires editing the docker-compose.yml override. CapRover supports resource limits in the app definition JSON.

Monitor actual usage with docker stats to calibrate these values. Overprovisioning wastes resources; underprovisioning triggers restarts.

Horizontal Pod Autoscaling Alternative

None of these platforms natively support Kubernetes-style autoscaling. Implement a custom solution using Docker Swarm's scaling API (CapRover) or external monitoring.

For CapRover in Swarm mode:

#!/bin/bash
SERVICE="captain-app-name"
MAX_REPLICAS=10
MIN_REPLICAS=2
CPU_THRESHOLD=70

while true; do
  CPU_USAGE=$(docker stats --no-stream --format "{{.CPUPerc}}" $SERVICE | sed 's/%//')
  CURRENT=$(docker service ls --filter name=$SERVICE --format "{{.Replicas}}" | cut -d'/' -f1)

  if (( $(echo "$CPU_USAGE > $CPU_THRESHOLD" | bc -l) )) && [ $CURRENT -lt $MAX_REPLICAS ]; then
    docker service scale $SERVICE=$((CURRENT + 1))
  elif (( $(echo "$CPU_USAGE < 30" | bc -l) )) && [ $CURRENT -gt $MIN_REPLICAS ]; then
    docker service scale $SERVICE=$((CURRENT - 1))
  fi

  sleep 60
done

Deploy this script as a systemd service or cron job on the manager node.

Monitoring and Observability

Centralized Logging

Default logging to Docker's JSON driver becomes unmanageable beyond a handful of services. Ship logs to a centralized system.

Promtail + Loki stack integrates cleanly with all three platforms:

services:
  promtail:
    image: grafana/promtail:latest
    volumes:
      - /var/log:/var/log
      - /var/lib/docker/containers:/var/lib/docker/containers:ro
      - ./promtail-config.yml:/etc/promtail/config.yml
    command: -config.file=/etc/promtail/config.yml

Promtail configuration:

server:
  http_listen_port: 9080

positions:
  filename: /tmp/positions.yaml

clients:
  - url: http://loki-service:3100/loki/api/v1/push

scrape_configs:
  - job_name: containers
    static_configs:
      - targets:
          - localhost
        labels:
          job: docker
          __path__: /var/lib/docker/containers/*/*.log

Deploy Loki and Grafana as separate services within your platform. Query logs by container name, timestamp, or custom labels.

Metrics Collection Without Overhead

cAdvisor collects container metrics with minimal footprint:

cadvisor:
  image: gcr.io/cadvisor/cadvisor:latest
  volumes:
    - /:/rootfs:ro
    - /var/run:/var/run:ro
    - /sys:/sys:ro
    - /var/lib/docker/:/var/lib/docker:ro
  ports:
    - "8080:8080"
  privileged: true

Scrape cAdvisor metrics with Prometheus or push them to a time-series database. Track CPU throttling, memory pressure, and network I/O per container.

Edge Cases and Troubleshooting

WebSocket Connection Drops

Traefik's default idle timeout closes WebSocket connections prematurely. Applications using real-time features (chat, live updates) break intermittently.

Increase timeouts in Traefik labels (Dokploy, Coolify):

labels:
  - "traefik.http.services.app.loadbalancer.server.transportserver.timeout=3600s"
  - "traefik.http.services.app.loadbalancer.server.transportserver.responseheadertimeout=600s"

For CapRover Nginx:

location /ws {
    proxy_pass http://backend;
    proxy_http_version 1.1;
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection "upgrade";
    proxy_read_timeout 3600s;
    proxy_send_timeout 3600s;
}

Docker Network MTU Mismatches

Cloud providers often use jumbo frames or encapsulation that reduces effective MTU. Symptom: large HTTP responses hang, small requests succeed.

Set explicit MTU on Docker networks:

docker network create --driver=overlay --opt com.docker.network.driver.mtu=1400 my-network

For existing networks in CapRover Swarm:

docker network rm captain-overlay
docker network create --driver=overlay --opt com.docker.network.driver.mtu=1400 --attachable captain-overlay

Restart all services after recreating the network.

Build Cache Exhaustion

Repeated deployments fill /var/lib/docker with dangling images and build cache. Platforms do not prune automatically.

Schedule aggressive pruning:

0 2 * * * docker system prune -af --volumes --filter "until=72h"

For production systems, exclude active volumes by tagging them:

docker system prune -af --filter "label!=production" --filter "until=72h"

Label production volumes in your compose files:

volumes:
  postgres-data:
    labels:
      - production=true

Migration and Portability

Exporting Configurations

Dokploy stores project definitions in SQLite at /data/dokploy.db. Export applications by dumping the database and backing up volumes.

Coolify uses PostgreSQL for configuration. Export with:

docker exec coolify-db pg_dump -U coolify > coolify-backup.sql

CapRover stores app definitions in the captain-captain container's /captain/data volume. Copy the entire directory:

docker cp captain-captain:/captain/data ./caprover-backup

None of these formats are interoperable. Migrating between platforms requires manual translation of docker-compose files and environment variables.

Zero-Downtime Platform Upgrades

Updating the platform itself risks downtime. Stage upgrades on a separate node when possible.

CapRover in Swarm mode allows this pattern:

  1. Add a new manager node to the cluster
  2. Drain the old manager: docker node update --availability drain <old-node>
  3. Promote the new node: docker node promote <new-node>
  4. Update CapRover on the new node
  5. Remove the old node after validation

Dokploy and Coolify require in-place upgrades. Schedule maintenance windows and test updates on staging environments first.

Conclusion

Production-grade deployments on Dokploy, Coolify, or CapRover require explicit configuration beyond the defaults. Tune reverse proxy timeouts, implement connection pooling, enforce resource limits, centralize logging, and prepare for edge cases like WebSocket handling and MTU mismatches. None of these platforms match the maturity of Kubernetes for large-scale operations, but with deliberate optimization they deliver reliable self-hosted PaaS for teams that value simplicity and control. Test every change in staging, monitor resource usage continuously, and maintain runbooks for common failure scenarios.

FAQ

Which platform handles multi-region deployments best?

None of these platforms are designed for true multi-region active-active deployments. CapRover's Swarm support enables multi-node clusters within a single region. For global distribution, use these platforms as regional control planes and route traffic with GeoDNS or a CDN.

How do I debug container networking issues?

Deploy a netshoot container in the same network:

Can I run these platforms behind Cloudflare?

Yes, but disable Cloudflare proxy for health check endpoints. Traefik and CapRover health checks fail when Cloudflare caches or rate-limits the health check path. Create a dedicated subdomain like health.yourdomain.com with DNS-only mode (gray cloud) for health endpoints.

What's the best backup strategy?

Scheduled volume snapshots at the block storage level combined with application-level exports (database dumps, file archives). Test restoration regularly. Volume snapshots alone do not guarantee consistent database state; dump databases before snapshotting.