IONOS deploy.sh — automatic recovery from the docker compose CLI hang
The bug
deploy.sh's start_container_watchdog already documents a known, unresolved upstream docker/containerd race (moby/buildkit#748, #4438): mid-compose up, the CLI can stall for minutes with containers left fully created but never started. The watchdog mitigates that by polling and force-starting any container stuck in created once its own depends_on preconditions are satisfied.
What the watchdog does not cover: the parent docker compose up process itself can hang even after the watchdog has finished fixing every container underneath it. Observed live on 2026-08-28: the watchdog finished its work in under a minute (confirmed via docker ps — every container was Up and healthy), but the foreground compose up command deploy.sh was waiting on sat idle for 14+ minutes, never returning success or failure. deploy.sh had no way to notice — it just blocked forever on a command that had nothing left to do. A human had to SSH in, inspect the process tree directly (ps --forest) to confirm the process was genuinely idle rather than slow, and manually kill it.
The fix
Wrapped the foreground compose up -d --build ... call in timeout ${COMPOSE_UP_TIMEOUT_SECONDS} (default 480s / 8 min — generous over the observed normal range including a slow cold build + backfill (~7 min observed live), well short of the ~14 min point the hang was confirmed at).
If timeout fires (exit code 124), deploy.sh doesn't assume it's safe to continue — it re-checks every container under the sentinel-signal compose project via docker inspect and only proceeds if all of them reached a terminal state (running or exited). If anything is still stuck, that's treated as a real orchestration failure (not just the known CLI hang) and the script exits non-zero instead of silently pressing on. Any other non-timeout failure from compose up is still a hard failure, same as before this change.
This is the automatic version of the exact manual recovery that was done live: verify the containers, then kill the process that has nothing left to do.
Not tested by the app's own pytest suite
This lives in deploy/ionos/deploy.sh, a bash script the app's pytest suite doesn't execute — verified instead via bash -n (syntax), shellcheck (no new warnings beyond two pre-existing, unrelated ones at different line numbers), and python3 -m py_compile on the embedded Python check block. The real validation is empirical: this exact recovery sequence (verify containers are terminal, then kill the hung process) was performed manually tonight and worked cleanly both times it was needed.