Sentinel Signal

v1.0.669 — Sentinel Policy: S3 object storage provisioned + migrate logging bug fixed

Source: docs/sentinel-policy-s3-and-migrate-logging-fix-v1.0.669.md

Document Content

v1.0.669 — Sentinel Policy: S3 object storage provisioned + migrate logging bug fixed

Context

Follow-up to M1's launch (v1.0.665-667). The user added the DNS A record for policy.sentinelsignal.io (confirmed live, TLS cert issued after a Caddy restart forced an immediate retry — Caddy's first attempt during the original deploy predated the DNS record and wasn't scheduled to retry again on any short timescale). The user then asked whether IONOS offers S3-compatible object storage that could be reused for Sentinel Policy's raw-artifact archive — yes, confirmed via the existing Postgres backup script (deploy/ionos/backup-postgres.sh), which already uses https://s3.us-central-1.ionoscloud.com with real (non-placeholder) credentials already on the host. Sentinel Policy's S3ObjectStore was already built to default to that same endpoint.

What happened operationally

Set SENTINEL_POLICY_S3_BUCKET=sentinel-signal-policy-artifacts in /etc/sentinel-signal/prod.env (new bucket, consistent naming with the existing sentinel-signal-prod-db-backups bucket) and reused the existing IONOS S3 credentials. A docker compose restart policy-migrate appeared to succeed (exit 0) but never actually created the bucket — confirmed via a direct head_bucket check. Root-caused: docker compose restart reuses the existing container with whatever environment was baked in when it was originally created (during the M1 deploy, when the bucket var was still empty) — it does not re-read the env file. docker compose up -d policy-migrate policy-web policy-worker (which recreates the containers) was required to actually pick up the new setting. Verified end-to-end after that: registered a real source, triggered a real discovery run through the public API, and confirmed the fetched document landed at a real, sharded, content-addressed key (raw/cc/03/...) inside the actual IONOS S3 bucket — not the local-filesystem fallback.

A real bug found and fixed along the way

While diagnosing why the first policy-migrate restart looked like a silent no-op, workers/migrate.py's own log output turned out to be silently going dark after the Alembic call. Root cause: alembic/env.py calls fileConfig(config.config_file_name) (triggered inside command.upgrade()), and Python's logging.config.fileConfig() defaults to disable_existing_loggers=True — which disables every logger not explicitly declared in alembic.ini's [loggers] section, including sentinel_policy.workers.migrate's own logger. Every log line after command.upgrade() — including whether the object-store bucket actually got created — silently became a no-op, with no error of any kind. Confirmed directly: fileConfig("alembic.ini") flips logger.disabled to True. Fixed with an explicit logger.disabled = False; logger.setLevel(logging.INFO) re-assertion immediately after the command.upgrade() call, verified locally (log lines now appear correctly) before shipping.

Verification

  • PYTHONPATH=policy/src:policy/tests python -m pytest policy/tests -q — 30 passed (was 29; +1 new
  • regression test that directly exercises the exact fileConfig() disable-then-reassert sequence against the real alembic.ini, not a mock).

  • python -m pytest tests/ -q — 334 passed. verify/tests — 824 passed (no changes to either).
  • Live: confirmed the bucket exists (head_bucket succeeds), confirmed a real discovery run's
  • fetched document lands in it (list_objects_v2 shows the expected content-addressed key), confirmed policy-web's running container has SENTINEL_POLICY_S3_BUCKET correctly set in its environment.

  • https://policy.sentinelsignal.io/healthz and /v1/build confirmed reachable over HTTPS with a
  • valid Let's Encrypt certificate.

Operational notes for future deploys

docker compose restart <service> does NOT pick up .env/env-file changes — it must be docker compose up -d <service> (or a full deploy.sh run) to recreate the container with fresh environment variables. Worth remembering for any future manual/interim config change outside the normal deploy pipeline.