Production operations
What I want to know before I trust a Compose deployment
A practical operating note on artifacts, state, health checks, rollback, backups, and the questions a production Docker Compose file should answer.
A Compose file can start containers and still leave the product impossible to operate. Before I trust it, I want to be able to answer a smaller set of questions: which artifact is running, where state lives, how failure becomes visible, and how we get back to a known version.
This is the baseline I use when reading an inherited deployment. It is not a universal stack and it is not a substitute for testing the actual recovery path.
Start with the operating path
I sketch the release as a sequence before changing YAML:
source commit → immutable image → environment configuration → migration → health check → traffic → verification
Every arrow needs an owner. If a step exists only in one person’s shell history, the deployment is not repeatable yet.
The first pass is mostly questions:
- Can I identify the exact commit behind the running image?
- Is the same image promoted, or rebuilt differently for each environment?
- Which data survives a container replacement?
- Can the application report readiness without depending on a customer finding a problem?
- What happens when a migration succeeds halfway?
- Can I return to the previous image without guessing?
Build once, then name the artifact
Production should run a specific image, not whatever a mutable tag happens to point to. I prefer a tag or digest tied to the source revision.
services:
app:
image: registry.example.com/product:${APP_RELEASE:?set APP_RELEASE}
restart: unless-stoppedThe important part is not the registry vendor. It is that the release record can connect a commit, image, configuration change, and deployment time.
Multi-stage builds are useful when they separate build tools from runtime files. I still inspect the resulting image rather than assuming that “multi-stage” means small or safe.
FROM node:lts-alpine AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build
FROM node:lts-alpine AS runtime
WORKDIR /app
ENV NODE_ENV=production
COPY --from=build /app/dist ./dist
COPY --from=build /app/node_modules ./node_modules
USER node
CMD ["node", "dist/server.js"]That example still needs to be adapted to the application. Native modules, generated files, runtime writes, and framework-specific output can all change what belongs in the final image.
Treat state as a separate design problem
Containers are replaceable. Customer data is not.
I list every stateful path explicitly:
- the primary database;
- uploaded or generated media;
- queues and delayed work;
- caches that contain more than disposable acceleration;
- certificates and external credentials;
- backups and the information needed to restore them.
A named volume is persistence, not a backup. A backup is useful only after a restore has been exercised and the restored application has been checked.
Health is more than “the process exists”
A running process can be unable to serve a customer. The useful check depends on the product, but it should answer a real question.
I usually separate three concerns:
- Process health: is the service alive?
- Readiness: can it accept work without immediately failing?
- Journey verification: does a critical customer path still complete after release?
Compose can declare a container health check, but the endpoint and command need to match the image:
services:
app:
healthcheck:
test: ['CMD', './bin/healthcheck']
interval: 30s
timeout: 5s
retries: 3
start_period: 20sThe script should fail clearly. A check that always returns success only creates decorative confidence.
Make configuration inspectable without exposing secrets
I want to know which configuration keys a release expects and which system supplies them. I do not want secrets committed beside the Compose file or copied into an image layer.
The operating record should distinguish:
- non-secret configuration that can be reviewed;
- secret values supplied at deployment time;
- environment-specific endpoints;
- defaults that are safe locally but dangerous in production.
Missing required values should stop the release early. Quiet fallback is especially risky for storage, email, payment, and callback configuration.
Give migrations their own moment
Database migrations are a production event, not incidental container startup noise.
Before release I need to know:
- whether the migration is backward-compatible with the current application;
- how long it can lock or rewrite data;
- whether workers and web processes can run across the transition;
- what recovery looks like if application rollback follows schema change.
For higher-risk changes, expansion and cleanup belong in separate releases. The safest rollback is often an application rollback that does not require reversing data immediately.
Keep logs outside the disappearing container
docker compose logs is helpful during diagnosis, but it should not be the only history of a production failure.
At minimum I check:
- that application output reaches a retained destination;
- that request or job identifiers connect related events;
- that repeated failures cannot fill the host disk silently;
- that alerts point to an inspectable error rather than a generic “service down” message.
The goal is not maximal logging. It is enough context to reconstruct a failed customer journey.
Rollback should be a written command, not a memory test
Before a release, I record the current image reference and the command that restores it. If configuration or schema changed, those constraints belong beside the rollback step.
APP_RELEASE=previous-known-release docker compose up -d
docker compose psThat command is only the beginning. Verification after rollback matters just as much as verification after forward release.
The short release record
I keep a compact note for each meaningful production change:
| Field | What I record |
|---|---|
| Source | Commit or release reference |
| Artifact | Exact image tag or digest |
| Change | Configuration, schema, and service changes |
| Checks | Health and customer-path verification |
| Recovery | Previous artifact and rollback constraint |
| Result | What was observed after release |
This is intentionally unglamorous. When an incident happens months later, the useful thing is a trail that explains what changed and how the system was checked.
What “ready” means to me
I do not call a Compose deployment production-ready because the containers start. I call it operable when another authorized person can identify the running artifact, inspect failure, protect the state, execute the release, and recover without reconstructing the system from chat history.
That standard is small enough to use and strict enough to expose the gaps that matter.