Sentinel Signal

v1.0.639 — Server detail page fast-fallback caching

Source: docs/mcp-verify-v1.0.639-server-detail-fast-fallback.md

Document Content

v1.0.639 — Server detail page fast-fallback caching

Problem

A week-long traffic assessment (2026-08-15) found /servers/{namespace}/{name} — the public HTML server detail page, Verify's single most-visited surface — had no fast-fallback cache path. Every request, hit or miss, did a full synchronous build: build_server_insights, alias/claim/watch/dispute/rebuttal queries, agent-commerce assessment, then HTML render. It was the single largest driver (56.7%) of the site's slow-request tail (p95/p99 in the 4-15s range, up to a 30s ceiling), live-confirmed with a real 6-14s page load on awesome-stevejford/shiply-mcp.

The same investigation also found a still-open regression, independent of this fix: slow-request ratio rose from a baseline 11-15% to 25-33% and 5xx ratio from 0% to ~0.7-1.2%, starting almost exactly at the ecdd0cd "bound database growth and catalog queries" deploy (2026-08-14 20:07, v1.0.636), which tightened verify-web's DB pool (pool_size=10, max_overflow=5, pool_timeout=5s) and added a new, web-only idle_in_transaction_session_timeout=60000ms. The detail-page handler held its request-scoped transaction open for its entire duration — already observed at 4-30s — making it the most plausible single contributor to that regression, though not confirmed by direct log evidence (container log history for the incident window was gone by the time it was investigated, since deploy.sh's up -d --build recreates containers and Docker's default log driver doesn't persist through recreation). This change addresses the chronic tail directly and independently reduces pressure on that pool regardless of which exact setting was the regression's proximate trigger; root-causing the regression itself with durable logs/Loki is tracked separately, not part of this change.

What shipped

/servers/{namespace}/{name} now goes through the same _get_cached_route_value fast-fallback pattern already used by /report, /policy, and /manage, finally wiring up build_fast_server_page — previously a dangling reference in one docstring and one test comment that explicitly said it was "hardcoded to None."

  • New settings flag MCP_VERIFY_SERVER_DETAIL_FAST_MISS_ENABLED (default
  • False), distinct from server_page_fast_miss_enabled (which only gates /manage) so this much-higher-traffic route can be rolled out and rolled back independently.

  • New dedicated cache store (server_detail_page, its own lock/limit of
  • 4096 entries) rather than sharing /manage's server_page store, so the site's highest-traffic surface doesn't contend with low-traffic publisher-only pages for capacity.

  • build_fast_server_page constructs a ServerDetailResponse directly from
  • the already-fetched server row (no extra queries), applies apply_fast_fallback_readiness(..., validations=None) — the same in-memory-only correction /policy//report use to avoid the r19 bug class ("Evaluation only"/50.0 shown for a server that was genuinely never assessed) — then renders the real detail template with empty validations/observed-attention/percentile placeholders. evidence_confidence is left exactly as that readiness call sets it (not overwritten with a generic "warming" label), so a metadata-only server still shows a correct not-assessed verdict on the fast path, not a misleading placeholder. refresh_on_miss_fallback=True and store_miss_fallback=True, so a cold miss returns quickly while the real build runs in the background and replaces the cached entry.

  • stale_key_prefix is gated through _machine_stale_prefix (matching
  • /report//policy, unlike /manage's unconditional version): stale cross-revision serving only applies to crawler/bot traffic. A regular browser request always gets its own just-committed data, never a previous-fingerprint cached page — required so a validation completing between two requests is reflected immediately for real visitors, not just eventually.

  • Cache participates in the existing global write-invalidation sweep
  • (should_invalidate_public_route_caches) alongside every other route cache, and its own cache key already rotates automatically on any change to the server's mutable state (score, last_validated_at, tool count, etc.) via server_detail_snapshot_cache_key, independent of that sweep.

Not included in this change

  • Prewarming: _prewarm_default_route_caches/prewarm_server_page were not
  • extended to also populate the new HTML cache. The existing prewarm blocks for server_report/badge write directly into their stores using a raw (unscoped) key, and it wasn't established with confidence that this matches the scoped key _get_cached_route_value computes at actual request time for those routes either — extending prewarm the same way for this route risked quietly shipping a second, subtler version of that same question rather than fixing anything. Deferred; the fast-fallback path already makes a cold miss cheap regardless.

  • The Tuesday DB-tuning regression itself: durable Postgres/Loki log
  • evidence is needed to confirm or rule out the idle-transaction-timeout hypothesis before tuning MCP_VERIFY_DB_IDLE_TRANSACTION_TIMEOUT_MS or the web pool sizing. Re-measure the regression after this change ships before deciding whether further DB-side changes are still needed.

Rollout

Ships with MCP_VERIFY_SERVER_DETAIL_FAST_MISS_ENABLED unset (off) — no behavior change until explicitly enabled in deploy/ionos/prod.env.

Verification

  • PYTHONPATH=verify/src:verify/tests python -m pytest verify/tests -q — 821
  • passed. Two pre-existing tests (test_trust_snapshot_caches_follow_latest_server_state, test_trust_surface_revision_rotates_without_rendered_page_cache_clear) updated: they previously asserted the route's hardcoded, cache-free uncached_current status; they now assert the real miss_refresh/hit sequence every other cached route already exhibits, and continue to prove a data change between two requests is reflected immediately (via the rotating cache key, not an explicit clear).

  • python -m pytest -m unit -q — 194 passed.
  • Live, post-deploy: confirm no behavior change with the flag off; enable
  • for a canary window and re-check mcp_verify_slow_requests_total ratio and mcp_verify_http_request_duration_seconds p95/p99 for route="/servers/{namespace}/{name}" against last week's baseline, and load a previously-slow page fresh to confirm sub-second response.