Why I Left Docker Compose in Production

Title first, honesty second: I did not leave because Compose is bad. I left because the things I run for years outlive the assumptions Compose is built on, and the migration was cheaper than the recurring tax. If your service is restarted weekly, redeployed from CI, and staffed by more than one on-call rotation, much of this will not apply to you. If you run a box for five years and expect it to survive your own absence, read on.

What Compose is honest about

Compose describes a desired state of containers on one machine. For development that is exactly the right scope, and I use it daily for that. The trouble is that production on a long-lived host is not a state problem. It is a history problem: five years of upgrades, a kernel you did not choose, a certificate workflow that predates you, and 3 a.m. you, who did not write the file and cannot remember which of the three similar services matters.

The four taxes

1. The YAML is not the system

The compose file says what should run. It does not say how it got there, what depends on what, or in which order things must come back after a reboot. Mine accumulated a README, three “remember to” shell scripts, and finally a Makefile that ran the compose commands in the right order — at which point I had built a worse unit file with extra steps. The system's real definition lived in my head and in shell history, and shell history is not a runbook.

2. Restart semantics are shallower than they feel

restart: unless-stopped covers “the process died.” It does not cover dependencies coming up out of order after a host reboot, health checks gating readiness, or the difference between “running” and “serving.” I have watched a stack come up green on a Sunday morning and fail its first real request at 08:59 because a database was accepting TCP before it was accepting queries. systemd solved this in 2014 with After= and ExecStartPre=, and the units are three lines each.

3. Upgrades couple to the Compose version

Every few years the compose file format or the plugin's behavior shifts enough that the file needs reviewing. On a five-year box that review lands during an unrelated emergency, because that is when you finally run docker compose up after a long gap. Native units have survived three OS upgrades of mine with zero file changes. The distro's promise — stable interfaces for years — is exactly what a long-lived server needs.

4. The escape hatch is always host networking

The final straw for me: one service needed host networking for latency reasons, which silently broke the internal DNS assumptions of the other three, which forced more host networking, and so on. When half your services are in host mode, the network isolation you were running Compose for exists mostly in the file's feelings. Either commit to the container network or commit to processes; the middle ground is the worst of both and the hardest to debug.

What I run instead

For services that are one process each — which is most of them, forever, honestly: plain systemd units, generated from a template, with DynamicUser=, hardcoded Environment= lines, and a Restart=on-failure. The whole thing is reviewable in one terminal window. For the two services that genuinely want a container boundary, a container under systemd, not beside it.

The checklist I now apply before letting Compose near production, which is also the checklist that talked me out of it:

The part where I defend Compose anyway

Use it for development environments, for CI, for one-week hacks, and for services on machines whose lifecycle is shorter than the file's. Compose failing on a three-year horizon is not a design flaw; three years is simply not its design target. The mistake was mine: I used a development tool as an operations platform because the first month was free. The bill arrives in year two, always on a Sunday, always while you are traveling.