Sentinel Signal

MCP Verify — round 18 implementation (post-1.0.524, "Requirements Round 12")

Source: docs/mcp-verify-round-18-implementation.md

Document Content

MCP Verify — round 18 implementation (post-1.0.524, "Requirements Round 12")

Implements the plan in verify_round18_architecture_plan (memory), built against one fixture: pipeworx-io/mcp-wakatime (trustsnap_8b9b79a195c7510a, 43 tools, real write/destructive surface, server card present). 441 tests passing in verify/tests, up from 425 at the end of round 17. No new capability-classifier rules (§0 honored — R62 measures, R60 is a scoring dimension over existing classifier output).

R56 — the stale server-detail cache, not a validator or scheduler bug

The planning pass proved, via a direct production DB query, an isolated in-process function call, and a live service restart, that every one of the fixture's "Healthy at zero tools" runs genuinely observed all 43 tools — the persisted data was never wrong. The bug was a stale route_server_detail_cache entry that survived far longer than its 30-minute TTL should have allowed.

**Fix, applied in main.py:**

  • server_detail_snapshot_cache_key gained an optional
  • latest_validation parameter. When provided, the validation run's own id is appended to the cache key.

  • _get_cached_server_detail_response's write path now computes its
  • cache key from validations[0] (the actual data the response was just built from), not from the server object's attributes as they stood when the function started. The read check is unchanged (still keyed from server, for a fast pre-fetch cache-hit check). If server's attributes are ever stale relative to what's actually in the database — regardless of the exact mechanism, which the planning pass could not pin down with full certainty from a single fixture — the worst outcome is now an extra cache miss (a fresh rebuild), never a wrong hit, because the write key is always self-consistent with the payload it labels.

  • ServerRepository.list_validation_runs_for_server_id ordered its
  • queries by started_at DESC instead of completed_at DESC (its sibling list_validation_runs_for_server_ids already used the correct ordering) — fixed in all three sub-queries. This was a correctness dependency for the cache-key fix above (both the "top N full rows" and the "latest validation" peek must agree on what "latest" means), and incidentally closes round 17's still-open R45 finding.

  • extract_capability_counts (insights.py) now returns None (not
  • 0) for tool_count/prompt_count/resource_count when validation.checks == {} — the sentinel the repository's lightweight "summary" rows (beyond full_payload_limit) use, by design, to avoid fetching full check payloads for older history rows. This was rendering as a false zero-tool observation for every row past the fetch window, on every server with more than 2 validation runs. render_validation_timeline now shows "not fetched" for these rows instead of a bare 0.

  • Follow-through: compute_current_score already treated tool_count=None
  • as "don't apply the zero-tool ceiling" (no change needed there), but build_validation_diff's tool_count_delta/prompt_count_delta/ resource_count_delta did a bare subtraction that crashed on None — fixed to return None (not a fabricated delta) when either side wasn't fetched. Two downstream consumers (insights.py's tool-count-drop alert, notifications.py's tool-surface-changed message) needed matching None-safe guards.

Not fully resolved, flagged explicitly, not silently left as "done": the exact mechanism that let a stale entry survive ~15 TTL cycles is not identified with full certainty — the fix closes every plausible variant of the bug (any mismatch between the cache key and the cached payload now self-heals on the next request) without requiring that certainty, but if this exact symptom resurfaces, the next step is the instrumentation the planning pass recommended (log the cache key and a fingerprint of the computed payload at every write) rather than another guess.

R62 — flag prevalence audit

New scripts/generate_flag_prevalence_report.py (compute_flag_prevalence_report): for every capability and risk flag, computes corpus-wide tool-level and server-level share, plus a per-server distribution for flags. Writes verify/eval/reports/flag-prevalence-<date>.json. Root cause traced for the triggering finding (38 of 43 wakatime tools carrying arbitrary_network_egress): detect_tool_capabilities's openWorldHint: true → network capability path reads a broad MCP annotation ("this tool talks to an external system") as a narrow, accusatory claim ("the caller can direct outbound requests anywhere") — not fixed this round, per §0/R62.4.

DEFAULT_PREVALENCE_REVIEW_THRESHOLD = 0.25 (25% of tools corpus-wide), documented in the script rather than tuned to make any specific flag pass. Added a "fixed-endpoint API wrapper" class to the R24.3 labeled classifier set (test_capability_classifier.py) — three hand-built tools matching wakatime's real shape (no fixture JSON exists for it locally), deliberately not wired into the enforced accusatory-precision assertion since network/arbitrary_network_egress aren't accusatory classes and the classifier itself isn't changing this round.

R63 — per-tool evidence links + dispute deep-linking

build_tool_security_inventory (validation/service.py) now includes each tool's input_schema and annotations in its per-tool output — the schema was already being inspected to produce the capability/flag classification, just never carried through to the response. render_tool_security_inventory gained a per-row id="tool-{name}" anchor, a <details> expansion showing the schema/capabilities/flags/ annotations, and a "Dispute this classification" link to ?dispute_tool={name}#dispute. The dispute form (render_dispute_section) gained a small client-side script that reads dispute_tool from the query string and pre-selects "Risk flags / capability classification" plus a pre-filled explanation naming the tool — the link itself is a plain GET link and still works with JS disabled (jumps to #dispute; only the pre-fill is progressive enhancement).

R64 — summary/remediation string split + publishability divergence explained

build_client_profiles's criteria tuples (insights.py) changed from (passed, observation) to (passed, observation, action) across all four client profiles (~26 criteria total). _make_profile now returns both missing_requirements (observation text, unchanged usage) and a new remediation_actions (instruction-phrased text) field, plus criteria_total/criteria_passed. build_client_remediation_modes reads remediation_actions for its checklist instead of missing_requirements (falls back to the old field for any already-cached response predating this change).

The ChatGPT-vs-Claude compatibility divergence traced to two genuinely different (not buggy) criteria lists — 9 items for ChatGPT (including two OAuth-related checks Claude's list doesn't have), 6 for Claude — so the same failing check (transport_compliance_probe in error) can cross one profile's 80% "compatible" threshold and not the other's. Not a one-question-two-computations bug; the six publishability_policy_profiles gates shared by both remain a single, correct, identical computation. render_client_profiles now shows "N of M requirements met" next to each compatibility label so the divergence is self-explanatory instead of looking like a contradiction.

R60 — Personal Data Exposure, kept as an experimental candidate

Round 13 post-1.0.524 verification superseded the earlier promotion: personal_data_exposure_score remains a visible, weight-0.0 candidate until the empirical report/gate exists. It is computed from export/bulk-access classification and schema-declared batch bounds, never prose inference, but is not persisted in production current_score_components and does not affect totals, verdicts, gates, percentiles, rankings, or company-knowledge policy.

Historical 1.0.526 rows may still carry the key in stored score components. The server renderer now suppresses that stale production rendering and shows the value only in the "Experimental candidate components" section.

R65 — one candidate list, not two identical ones

build_action_controls_diff's manual_review_candidates was defined as "every item whose name is in disabled_by_default_candidates" — the same filter applied twice, structurally guaranteed identical for every server, not a coincidence of the fixture. Removed; disabled_by_default_candidates is now the one list, and admin_refresh_required reads it directly instead of the removed field. build_tool_snapshot_diff's added_tools (capped at 25 with no indication) gained an added_tools_omitted_count field, matching the omitted_count pattern already used elsewhere.

R57 — schema_divergence_probe surfaces what it already compared

build_schema_divergence_probe's "no server card tools" early return computed server_version_mismatch/card_server_version/ live_server_version/auth_scheme_mismatch (same as server_name_mismatch) but only ever included server_name_mismatch in its response — compared_dimensions claimed all four were checked; three were silently dropped. All four now included. No status semantics change.

R7 — taxonomy: fourth request, "query"/"message" removed

TAXONOMY_RULES["database"] no longer includes "query"; ["communication"] no longer includes "message" — both ordinary API-documentation vocabulary ("query parameter", "commit message") for a coding-activity tracker with no database or messaging functionality. Synced across all three copies (insights.py, agent.py, db/repository.py — the latter never had "query" in its own database list, only "message" needed removing there, another data point that these three copies have never been fully in sync, a pre-existing condition this round didn't attempt to fix structurally).

Test suite

441 tests passing, up from 425. New/extended coverage: 4 tests for R56 (cache-key rotation, tri-state tool counts, score-ceiling exemption, an end-to-end timeline check), 2 for R62's prevalence report, 1 for R63's evidence links/dispute deep-link, 2 for R64's string split and checklist wiring, 3 for R60 (scoring function behavior, the company-knowledge gate in isolation, and a regenerated score-component-registry/status-matrix pair), 1 for R65, 1 for R57, and 3 for R7. Root tests/ suite (app/ token_service, unrelated to this round's changes) unaffected: 293 passing.

Deferred / carried forward

R56's exact stale-write mechanism (the fix closes every plausible variant without needing to identify which one occurred; instrumentation is the recommended next step if it recurs). R11 (still blocked on human reviewers). R64's remediation_actions action-text wording could use a second pass once real user feedback exists on whether it's actionable enough. The three taxonomy-rule copies' broader sync problem (not attempted structurally this round, same as every prior round that's touched this).