Skip to content

Join the Seedly owners community →

Docker

Docker in Production

Health checks, restart policies, logging, security, and CI/CD for production containers

Written by 14 min read1 activity
Sprout, your presenter

Sprout presents

Health checks, restart policies and sensible logging. Three habits that keep containers running in production, and I have strong, well-reasoned opinions about each.Health checks, restart policies and sensible logging. Three habits that keep containers running in production, and I have strong, well-reasoned opinions about each.

Sprout listens to a shipping container through a stethoscope
Health checks confirm your app is really working

Running Docker on your laptop is one thing. Running it in production, where real people lean on your app around the clock, is a whole different animal. This lesson covers the stuff that keeps production containers alive... health checks, restart policies, logging, security and deployment pipelines.

Health Checks: Is Your App Really Working?#

A container can say "Up" in docker ps while your app is completely busted. Maybe the process is running but stuck in an infinite loop. Maybe the database connection quietly dropped. Health checks tell Docker how to confirm your app is actually responding.

# curl has to actually exist in your image for this to work
HEALTHCHECK --interval=30s --timeout=5s --retries=3 \
  CMD curl -f http://localhost:3000/health || exit 1

That instruction tells Docker three things.

  • Check every 30 seconds
  • Wait up to 5 seconds for an answer
  • After 3 failures in a row, mark the container unhealthy
# See health status
docker ps
 
# CONTAINER ID   IMAGE   STATUS                    NAMES
# a1b2c3d4e5f6   myapp   Up 5 minutes (healthy)     web
# f6e5d4c3b2a1   myapp   Up 5 minutes (unhealthy)   web-2

I care about this one a LOT. A rented software stack once broke a client flow I'd spent weeks tuning, and I only found out Monday when dozens of forms were dead. "It looked like it was running" is an expensive sentence.

Health Check in Your App#

# Express.js health endpoint example
# GET /health returns 200 if DB is connected, 500 otherwise
# In your Dockerfile (Alpine images ship wget but not curl)
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
  CMD wget --no-verbose --tries=1 --spider http://localhost:3000/health || exit 1

The --start-period=10s gives your app 10 seconds to warm up before failed health checks start counting against it.

Restart Policies: Self-Healing Containers#

So what happens when your container crashes at 3 AM? With no restart policy, it stays down until somebody wakes up and restarts it by hand. Restart policies handle the recovery for you.

# Always restart (even after Docker daemon restarts)
docker run -d --restart always myapp
 
# Restart unless manually stopped
docker run -d --restart unless-stopped myapp
 
# Restart on failure only (with max attempts)
docker run -d --restart on-failure:5 myapp

In Docker Compose it looks like this.

services:
  web:
    image: myapp
    restart: unless-stopped
    # or: restart: always
    # or: restart: on-failure

Logging: The Right Way#

In production you need to collect and search the logs from all your containers. The Docker best practice is dead simple. Write logs to stdout/stderr.

Why stdout/stderr?#

# BAD: Writing to a file inside the container
# - Lost when container is removed
# - Hard to access from outside
# - Can fill up container storage
CMD ["sh", "-c", "node server.js >> /app/logs/app.log 2>&1"]
 
# GOOD: Writing to stdout (Docker captures it automatically)
CMD ["node", "server.js"]

When your app writes to stdout/stderr, a bunch of good stuff happens.

  • docker logs shows the output instantly
  • Docker's logging drivers can ship logs off to a central service
  • There's no log file slowly filling up the container's storage
  • It works on every container orchestrator out there

Configuring Log Drivers#

Docker can send logs to a lot of different places.

# Send logs to a JSON file (default)
docker run --log-driver json-file myapp
 
# Limit log file size (prevent disk filling)
docker run --log-opt max-size=10m --log-opt max-file=3 myapp

And in Docker Compose.

services:
  web:
    image: myapp
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"

Security: Running as Non-Root#

By default, processes inside a container run as root. That's a security risk, because if an attacker finds a hole in your app, they've got root access inside the container. Run as a non-root user. Every time.

FROM node:24-alpine
WORKDIR /app
 
# Install dependencies
COPY package*.json ./
RUN npm ci --omit=dev
 
# Copy application code
COPY . .
 
# Create a non-root user and switch to it
RUN addgroup -S appgroup && adduser -S appuser -G appgroup
RUN chown -R appuser:appgroup /app
USER appuser
 
# Now the app runs as "appuser", not root
EXPOSE 3000
CMD ["node", "server.js"]

Additional Security Practices#

# Don't store secrets in the image
# BAD:
ENV API_KEY=sk-secret-12345
 
# GOOD: Pass secrets at runtime
# docker run -e API_KEY=sk-secret-12345 myapp
 
# Use specific image versions (not :latest)
# Pinned version, predictable
FROM node:24.21-alpine
 
# Scan images for vulnerabilities
# docker scout cve myapp:latest

CI/CD: Automated Build and Deploy#

In production you really shouldn't be building and deploying by hand. Let your CI/CD pipeline do it for you.

# Example GitHub Actions workflow
name: Build and Deploy
 
on:
  push:
    branches: [main]
 
jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
 
      - name: Build Docker image
        run: docker build -t myapp:${{ github.sha }} .
 
      - name: Push to registry
        run: |
          docker tag myapp:${{ github.sha }} registry.example.com/myapp:latest
          docker push registry.example.com/myapp:latest
 
      - name: Deploy
        run: |
          # Pull new image and restart on production server
          ssh production "docker pull registry.example.com/myapp:latest && docker compose up -d"

The usual CI/CD flow goes like this.

  1. A developer pushes code to GitHub
  2. The CI pipeline builds a new Docker image
  3. The image gets pushed to a private registry
  4. The production server pulls the new image
  5. Old containers get swapped out for new ones

A client once asked me what happens to their site if I die. Kinda morbid (fair question though), and it pushed me to rebuild my whole agency around real systems. An automated pipeline is that same idea for your app. The deploy steps live in a file instead of one person's head.

Resource Limits#

Sprout watches a wooden wobble toy rock back upright after being knocked over
Restart policies let a crashed container get back up

Don't let one greedy container hog the whole server.

# In compose.yaml
services:
  web:
    image: myapp
    deploy:
      resources:
        limits:
          cpus: "1.0"        # Max 1 CPU core
          memory: 512M       # Max 512MB RAM
        reservations:
          cpus: "0.25"       # Guarantee 0.25 cores
          memory: 128M       # Guarantee 128MB RAM

Production Dockerfile: Complete Example#

Here's a production-ready Dockerfile that pulls all of it together.

# Build stage
FROM node:24-alpine AS builder
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build
 
# Production stage
FROM node:24-alpine
WORKDIR /app
 
# Security: non-root user
RUN addgroup -S app && adduser -S app -G app
 
# Production dependencies only
COPY package.json package-lock.json ./
RUN npm ci --omit=dev && npm cache clean --force
 
# Copy built application
COPY --from=builder /app/dist ./dist
 
# Switch to non-root user
USER app
 
# Health check
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
  CMD wget --no-verbose --tries=1 --spider http://localhost:3000/health || exit 1
 
# Configure
EXPOSE 3000
ENV NODE_ENV=production
CMD ["node", "dist/index.js"]

Scaling Beyond One Server#

Once your traffic outgrows a single server, you need container orchestration.

  • Docker Swarm is Docker's built-in orchestrator (simpler, good for smaller setups)
  • Kubernetes is the industry standard for running containers at big scale

These tools take care of a lot.

  • Running containers across a bunch of machines
  • Automatic load balancing
  • Rolling updates with zero downtime
  • Self-healing (restarting failed containers on healthy machines)
  • Horizontal scaling (adding more containers as traffic grows)

Production Checklist#

Before you ship containers to production, make sure of all this.

  1. Health checks are set up and actually working
  2. A restart policy is set (unless-stopped or always)
  3. Logs go to stdout/stderr with size limits
  4. The app runs as a non-root user
  5. No secrets are baked into the image
  6. Specific image versions are pinned (not :latest)
  7. Resource limits are set so nothing runs away with the server
  8. A CI/CD pipeline handles builds and deploys
  9. Images get pushed to a private registry

TL;DR#

  • HEALTHCHECK confirms your app is actually responding, beyond just running
  • Restart policies bring containers back after they crash
  • Write logs to stdout/stderr and ship them to a central logging service
  • Run as a non-root user for security (with the USER instruction)
  • CI/CD pipelines handle building, pushing and deploying images for you
  • Set resource limits so one container can't eat the whole server
  • Reach for container orchestrators (Kubernetes, Swarm) when you need more than one server

What's Next?#

That's a wrap on the Docker module! You went from what containers even are, through Dockerfiles and multi-container apps, all the way to shipping lean images to production. Professional developers and DevOps folks use these exact skills every day... and now you've got 'em too.

This lesson ends with a short activity.