MCP Verify Production Performance Runbook
This runbook covers the IONOS Verify performance stabilization release. It does not change public HTTP contracts, scoring, trust semantics, or validation cadence. Database retention remains disabled until its separate restore-readiness gate is approved.
Release topology
Every Verify process must use the same image tagged with the full Git commit SHA:
verify-migrate: schema migration only;verify-web: one Uvicorn process;verify-scheduler: one replica, scheduling only;verify-worker: three replicas, all executable job types exceptdatabase_maintenance;verify-maintenance-worker: one replica,database_maintenanceonly.
The initial rollout uses:
MCP_VERIFY_DATABASE_MAINTENANCE_INTERVAL_SECONDS=1800
MCP_VERIFY_INCREMENTAL_ROLLUPS_ENABLED=false
MCP_VERIFY_DATABASE_ROLLUP_WORK_MEM_MB=256
MCP_VERIFY_OBSERVED_ATTENTION_SESSION_CACHE_SECONDS=1800
MCP_VERIFY_INDEX_PAGE_FAST_MISS_ENABLED=true
Do not restore the five-minute interval until incremental/full parity and the 24-hour acceptance window both pass.
Pre-deployment diagnosis
Run these on the IONOS host without printing /etc/sentinel-signal/prod.env:
cd /opt/sentinel-signal
docker compose --env-file /etc/sentinel-signal/prod.env -f deploy/ionos/docker-compose.yml ps
curl -fsS https://verify.sentinelsignal.io/healthz
curl -fsS https://verify.sentinelsignal.io/metrics | grep -E 'mcp_verify_database_(maintenance|rollup)|mcp_verify_job_oldest_pending'
docker compose --env-file /etc/sentinel-signal/prod.env -f deploy/ionos/docker-compose.yml exec -T postgres \
psql -U postgres -d mcp_verify -c "select application_name,state,count(*) from pg_stat_activity where datname='mcp_verify' group by 1,2 order by 1,2;"
docker compose --env-file /etc/sentinel-signal/prod.env -f deploy/ionos/docker-compose.yml exec -T postgres \
psql -U postgres -d mcp_verify -c "select job_type,status,count(*),min(scheduled_for) from jobs where status in ('pending','running') group by 1,2 order by 1,2;"
Capture the homepage and server-detail baseline from Prometheus before deploying. Also record the 15-minute temporary-byte increase, maintenance duration, rollup lag, and oldest runnable job age.
Immutable deployment
The deployed source and release input must identify the same full SHA. IONOS is populated by rsync rather than a Git checkout, so write the release marker from the already-clean local checkout after the transfer and before deployment:
cd /path/to/local/claim-risk-score
test -z "$(git status --porcelain)"
test "$(git rev-parse HEAD)" = "<FULL_RELEASE_SHA>"
rsync -az --delete \
--exclude '.git' --exclude '.venv' --exclude 'node_modules' \
--exclude '__pycache__' --exclude '*.pyc' --exclude '.pytest_cache' \
--exclude 'dist' --exclude 'artifacts' --exclude '.DS_Store' \
./ root@66.179.248.190:/opt/sentinel-signal/
printf '%s\n' '<FULL_RELEASE_SHA>' | \
ssh root@66.179.248.190 'install -m 0444 /dev/stdin /opt/sentinel-signal/.release-sha'
ssh root@66.179.248.190 \
'cd /opt/sentinel-signal && RELEASE_SHA=<FULL_RELEASE_SHA> RELEASE_VERSION=<VERSION> ./deploy/ionos/deploy.sh'
The deployment script verifies either Git HEAD or the required rsync marker, builds one SHA-tagged Verify image, force-recreates the topology, validates public route/build identity, and compares every Verify container's Docker image ID and MCP_VERIFY_BUILD_SHA. A parity failure triggers the immutable rollback script.
Verify the topology directly:
cd /opt/sentinel-signal
VERIFY_EXPECTED_BUILD_SHA=<FULL_RELEASE_SHA> ./scripts/verify_runtime_revision_parity.sh
docker compose --env-file /etc/sentinel-signal/prod.env -f deploy/ionos/docker-compose.yml ps verify-web verify-scheduler verify-worker verify-maintenance-worker
Incremental-rollup parity gate
Do not enable incremental rollups directly against the only production copy. Restore a current encrypted production backup into an isolated PostgreSQL database, migrate it to the release revision, and retain a snapshot of all four summary tables plus analytics_rollup_state.
Run one full refresh with incremental mode disabled, export ordered rows from:
analytics_daily_rollups;validation_daily_summaries;validation_monthly_summaries;job_daily_summaries.
Then restore the same starting snapshot, enable incremental mode, seed the maintenance state as an existing installation, run enough refreshes to cover the changed window, and export the same ordered rows. Compare exact canonical JSON/CSV output, including counts, status JSON, minimum/maximum scores, latest validation IDs, evidence digests, and materialization timestamps excluded from equality. Exercise:
- a record completed inside the one-hour overlap;
- a late record in the prior UTC day;
- a UTC day boundary;
- a UTC month boundary;
- a failed refresh followed by retry;
- two idempotent refreshes with no source changes.
Only after exact parity passes, update the production environment without revealing it:
sudo sed -i 's/^MCP_VERIFY_INCREMENTAL_ROLLUPS_ENABLED=.*/MCP_VERIFY_INCREMENTAL_ROLLUPS_ENABLED=true/' /etc/sentinel-signal/prod.env
cd /opt/sentinel-signal
docker compose --env-file /etc/sentinel-signal/prod.env -f deploy/ionos/docker-compose.yml up -d --no-deps --force-recreate verify-maintenance-worker
Confirm the next maintenance result reports incremental mode and bounded touched-period counts. Confirm the two concurrent indexes exist and are valid:
select indexrelid::regclass, indisvalid, indisready
from pg_index
where indexrelid::regclass::text in (
'ix_validation_runs_completed_at_global',
'ix_jobs_updated_at'
);
Prometheus verification
Prometheus configuration and rules must load successfully:
docker compose --env-file /etc/sentinel-signal/prod.env -f deploy/ionos/docker-compose.yml exec -T prometheus \
promtool check rules /etc/prometheus/alerts.yml
curl -fsS https://verify.sentinelsignal.io/metrics | grep -E \
'mcp_verify_database_maintenance_(last_duration|interval|running_age|touched_periods)|mcp_verify_database_rollup_lag|mcp_verify_job_oldest_pending_age'
The MCP Verify Performance Stabilization Grafana dashboard covers homepage/server-detail p50/p95/p99, route-cache outcomes, PostgreSQL temporary-byte rate, maintenance duration/running age, rollup lag, queue age, and touched periods.
24-hour acceptance window
Keep the 30-minute interval for at least 24 continuous hours after incremental mode is enabled. Acceptance requires:
- homepage p95 at or below 1 second and p99 at or below 2 seconds;
- server-detail p95 at or below 1.5 seconds and p99 at or below 3 seconds;
- maintenance p95 at or below 60 seconds with no run above 120 seconds;
- PostgreSQL temporary writes below 1 GiB per 15 minutes;
- no duplicate maintenance or validation execution;
- exact summary parity and rollup lag below two configured intervals;
- no Search, Pricing, health, route-version, or trust-surface coherence regression.
After 24 green hours, restore the five-minute interval and align the observed-attention session cache with it. Recreate the web and maintenance roles so both settings take effect:
sudo sed -i 's/^MCP_VERIFY_DATABASE_MAINTENANCE_INTERVAL_SECONDS=.*/MCP_VERIFY_DATABASE_MAINTENANCE_INTERVAL_SECONDS=300/' /etc/sentinel-signal/prod.env
sudo sed -i 's/^MCP_VERIFY_OBSERVED_ATTENTION_SESSION_CACHE_SECONDS=.*/MCP_VERIFY_OBSERVED_ATTENTION_SESSION_CACHE_SECONDS=300/' /etc/sentinel-signal/prod.env
cd /opt/sentinel-signal
docker compose --env-file /etc/sentinel-signal/prod.env -f deploy/ionos/docker-compose.yml up -d --no-deps --force-recreate verify-web verify-maintenance-worker
Start a new 24-hour observation period after restoring five-minute freshness.
Rollback
If any gate fails, keep Verify online, revert to full refresh at 30-minute cadence, and retain the harmless concurrent indexes:
cd /opt/sentinel-signal
./deploy/ionos/rollback-verify.sh
For an incremental-only rollback that keeps the release image:
sudo sed -i 's/^MCP_VERIFY_INCREMENTAL_ROLLUPS_ENABLED=.*/MCP_VERIFY_INCREMENTAL_ROLLUPS_ENABLED=false/' /etc/sentinel-signal/prod.env
sudo sed -i 's/^MCP_VERIFY_DATABASE_MAINTENANCE_INTERVAL_SECONDS=.*/MCP_VERIFY_DATABASE_MAINTENANCE_INTERVAL_SECONDS=1800/' /etc/sentinel-signal/prod.env
sudo sed -i 's/^MCP_VERIFY_OBSERVED_ATTENTION_SESSION_CACHE_SECONDS=.*/MCP_VERIFY_OBSERVED_ATTENTION_SESSION_CACHE_SECONDS=1800/' /etc/sentinel-signal/prod.env
cd /opt/sentinel-signal
docker compose --env-file /etc/sentinel-signal/prod.env -f deploy/ionos/docker-compose.yml up -d --no-deps --force-recreate verify-web verify-maintenance-worker
./scripts/verify_public_route_versions.sh
python3 scripts/verify_trust_surface_coherence.py --base-url https://verify.sentinelsignal.io --server awesome-malinoto/tracepass-mcp-server --iterations 2
Record the failed gate, relevant Prometheus range, release SHA, image ID, job IDs, and rollback time. Never paste environment files, credentials, authorization headers, request bodies, or session identifiers into the incident record.