Sentinel Signal

MCP Verify Round 4 implementation notes

Source: docs/mcp-verify-round-4-implementation.md

Document Content

MCP Verify Round 4 implementation notes

Date: 2026-08-04

This documents the Round 4 corrections for MCP Verify. Round 4's own spec opened by withdrawing Round 1 TASK-08 / Round 3 TASK-08b (root-path cache): the reviewer's fetch tool caches by exact URL, and the origin was serving correctly the whole time. That's a real, useful finding worth recording here even though it required no code change: **always mint a fresh ?cb=<nonce> on every review URL**, never reuse one.

TASK-31 (refined) / TASK-22 — Degraded-on-contact hypothesis, falsified; real root cause found

Round 3 found Degraded was rare (22/74,804 servers) and dominated by a real signal (OAuth metadata mismatch). Round 4's spec proposed a new hypothesis: revalidating a Healthy server flips it to Degraded, explaining why top scorers were all stuck at the same stale age.

**Step 1 (required before any other work): force-revalidated ai.dreamlit/mcp directly** via RemoteValidationService.validate_server inside the production verify-web container. Result: summary_status=healthy, score=82.05 (higher, not lower). Hypothesis falsified — revalidation does not flip Healthy to Degraded.

Real root cause, found by inspecting the job queue directly: ai.dreamlit/mcp had a validate_server job pending since 2026-08-02 (2+ days overdue) that had simply never been picked up. Production had 39,657 pending validation jobs, and compute_validation_priority() in verify/src/mcp_verify/workers/scheduler.py gave fresh_trusted_index candidates (the site's own top-ranked servers) only a -6 priority bonus versus -20 for failing-status servers — and failing servers outnumber fresh_trusted_index candidates roughly 500:1 in production (74,679 vs ~150). Confirmed via direct query: pending failing_refresh jobs averaged priority 23.1; pending fresh_trusted_index_refresh jobs averaged 29.2 (worse). This is a genuine queue-starvation bug, not a validation-check defect — it fully explains why every server that had recently been revalidated showed as Degraded (it's disproportionately the failing/degraded servers getting re-processed) while top scorers went days without a fresh run.

Fix:

  • Raised the is_fresh_trusted_index_candidate priority bonus from -6 to -25 (workers/scheduler.py, compute_validation_priority).
  • Added a reserved per-tick enqueue quota for fresh_trusted_index candidates in enqueue_due_validation_jobs, mirroring the existing never_validated_candidates quota pattern — a deterministic guarantee, not just a better priority number, since a sufficiently large volume of low-score failing servers can still out-rank on priority alone.
  • Remediated the existing production backlog directly: UPDATE jobs SET priority = priority - 19 WHERE job_type='validate_server' AND status='pending' AND payload->>'reason'='fresh_trusted_index_refresh' (152 rows), unblocking ai.dreamlit/mcp's stuck job immediately rather than waiting for a new enqueue cycle.
  • This is also Round 1/3's carried TASK-22 ("scheduler priority for top-ranked servers") — resolved as part of this diagnosis, not separately.
  • Per the spec's acceptance criteria: did not touch compute_summary_status (data showed no reason to) and did not adjust strong_live_candidates thresholds — if that count rises, it's because top servers are now actually getting revalidated within the 48h window, not because a threshold moved.

TASK-32 — False "Freshly Validated" / "No Critical Risk" badges

Both were real, confirmed logic bugs, not just window-mismatch labeling:

  • "Freshly Validated" was bound to build_freshness_profile's internal 7-day bucket (bucket in {"verified_last_24h", "verified_last_7d"}), not the public 24h freshness window used everywhere else. Fixed to bucket == "verified_last_24h" only.
  • "No Critical Risk" checked security_summary.get("critical_risk_tools") / .get("critical_tools") — neither key exists. summarize_tool_security_inventory() (validation/service.py) returns high_risk_tools (which merges high+critical) and a nested risk_distribution.critical count; the top-level keys the badge read were always absent, so has_critical was always False regardless of actual critical-risk tools. Fixed to read risk_distribution.critical.

Both badges, the on-page display, and the embeddable SVG/JSON badge endpoints all resolve through the single build_trust_badge_states() function, so this was one fix, not several — and the SVG/JSON surfaces can no longer diverge from the on-page badges by construction.

TASK-33 — Freshness windows

Five windows were in simultaneous use with the words "fresh"/"stale" leaking across several of them. Reserved "fresh"/"stale" for the single 24h public window; renamed everything else:

  • build_freshness_profile's 7d/30d/stale bucket labels → "Evidence retained: within 7d/30d" / "Evidence over 30 days old".
  • "Freshness band" → "Evidence age tier" (Current trust snapshot, MCP TrustOps).
  • "stale score suppressed" → "evidence too old to display".
  • Publisher checklist "Fresh validation" (bound to the 48h window) → "Recent validation (48h)".
  • A generic "fresh evidence + healthy protocol checks" fallback reason (reachable for evidence up to 30 days old) → "no blocking gaps found in current evidence".
  • Added a windows table to /methodology documenting all five and what each is actually for.

Not fully audited: a handful of internal reporting buckets (e.g. the schema-backfill page's "warming" tier) were checked and found already correctly scoped to the canonical setting; a fully exhaustive site-wide grep of every remaining "fresh"/"stale" occurrence was not completed given the size of this round.

TASK-34 — Generic playbook fallback

remediation_playbook_for() and remediation_action_for() (insights.py) had a shared fallback used whenever a check's code wasn't in their lookup tables. 8 codes were missing playbooks (protocol_version_probe, session_resume_probe, advanced_capabilities_probe, tool_snapshot_probe, connector_replay_probe, openid_configuration, request_association_probe, action_safety_probe) and 7 of those were also missing from the action lookup. Wrote specific, concrete playbooks/actions for all 8 (naming actual headers, methods, and schema fields per the spec's requirement). Changed both functions to return None instead of generic text when a code truly isn't mapped — render_remediations() already omitted its "Playbook" <details> block on falsy input, so this is enforced structurally, not just by having filled in the known cases. extract_next_actions() (the Tier-1 "Recommended actions" list) now dedupes by rendered text, not just by remediation code, and shrinks below the target count rather than padding with repeats.

TASK-04 (partial) — De-duplication

  • "Production trust decision" / "Readiness class" removed from the render_trust_snapshot() HTML (Current trust snapshot section) and from the dedicated "Production readiness class" section (renamed "Critical alerts", which is the one piece of that section not already shown in the Tier-1 decision summary). The Tier-1 decision summary (render_server_answer_block) is now the only copy.
  • Alias-graph table score column now uses format_score() (one decimal) instead of the raw stored value, so the same server's score can't read as two different numbers (78.8 vs 78.77) on one page.
  • Not done: duplicate install snippets (2x) and duplicate machine-links block (3x) — lower-value, more mechanical dedup not reached this round.

TASK-03 (partial) — Empty section suppression

Implemented the single highest-impact item the spec named explicitly: the "Agent Commerce & Payment Readiness" section (~30 fields) now renders a one-line placeholder instead of the full assessment when commerce_signal is below moderate or payment_capable is none — exactly the ai.dreamlit/mcp case (weak/none). The full assessment remains in the underlying JSON payload; only the HTML section is gated.

Not done: the ~15 other empty-state placeholder strings the spec listed, a generic render-guard/suppressed-section-summary mechanism, and gating history sections on run_count >= 2. Most of these are embedded inside table rows within larger, otherwise-meaningful sections (not whole empty <h2> blocks), so a correct fix needs per-section judgment rather than one generic wrapper — scoped out given the size of everything else in this round.

TASK-35 — Percentile floor

db/repository.py's score_percentile() now also returns rank_from_top (a second COUNT query: servers strictly ahead of this one, +1). format_score_percentile() shows #N of M scored public servers for rank ≤10, and floors the percentage at 1% (never 0.0%) below that.

TASK-36 — Confidence label mismatch

infer_evidence_confidence_label() read evidence_confidence.get("level") / .get("status") — build_evidence_confidence() actually returns the band under the key "label". Neither of the read keys ever existed, so the function silently fell through to a crude validation-count heuristic (>=3 runs → High, >=1 run → Medium, else Low) that ignored the real confidence score entirely. That's exactly how the decision summary said "Medium" (1 validation run → heuristic fires) while six other surfaces on the same page correctly read the real score (41.2) and said "Low". Fixed to read .get("label"). Also added an inline explanation of the separate "confidence-weighted score" figure (an internal TrustOps input, not the public score) rather than leaving it unexplained.

TASK-37 — Client-compatibility gate polarity

Two of six publishability gates (write_actions_present, admin_refresh_required) are risk flags where False is the passing state — the other four need True to pass. Two separate render paths checked all six gates uniformly assuming "True passes":

  • The remediation checklist (insights.py, build_client_remediation_modes) flagged admin_refresh_required=False (the good state) as an unsatisfied blocker, rendered as the raw field name converted to prose ("admin refresh required is not yet satisfied").
  • The Publishability Policy Profiles table (main.py, render_publishability_policy_profiles) rendered the same passing state with good_when_true=True unconditionally, producing a red "no" badge on Admin Refresh Required: No — the exact contradiction the spec reported.

Fixed both with a shared polarity map and human-readable messages (no raw snake_case field names in user-facing text). Renamed the section title from "Why compatibility is limited by client" to "Client compatibility gate details" since it no longer presupposes a limitation exists.

TASK-38 — Discovered/Scored stat re-merge

The Round 3 merge used exact string equality (discovered_count == scored_count), which will essentially never hold on a live database (a handful of servers are always mid-validation). Changed to a relative threshold: merge into one "Indexed servers" stat whenever the gap is under 0.5% of the total. A genuine funnel (a meaningfully large scored/discovered gap) still renders as two distinct stats.

Not done this round (explicit scope cuts)

  • TASK-09 (tier restructure: collapse ~52 flat <h2> sections into 5 named, collapsible tiers behind the tab bar that already shipped) and TASK-16 (/manage split: move ~9 publisher-facing sections off the public page) are both large, standalone frontend restructuring projects. Attempting either in the same session as nine other fixes plus a production deploy was judged too high-risk to rush. The tab bar/anchors (#tool-risk, #compatibility, #evidence, #fix-it, #policy) already exist and are unaffected by this round's changes.
  • TASK-06 (percentile in the results table), TASK-11 (tool-name search), TASK-12 (filter tiering), TASK-13 (Risk column), TASK-18/20/23/24/25 — unchanged from prior rounds' backlog.

Verification

PYTHONPATH=verify/src pytest verify/tests   # 242 passed
python3 -m py_compile verify/src/mcp_verify/main.py verify/src/mcp_verify/insights.py \
  verify/src/mcp_verify/db/repository.py verify/src/mcp_verify/workers/scheduler.py \
  verify/src/mcp_verify/validation/service.py

Production diagnostics used this round (revalidation experiment, job-queue inspection, priority-distribution query) were run via docker compose exec verify-web python3 -c ... and docker compose exec postgres psql, not scripted -- worth turning the priority-distribution query into a standing script (alongside Round 3's scripts/report_status_distribution.py) if this becomes a recurring check.