MCP Verify — Round 16 Implementation Report (post-1.0.518 verification)
Implements "MCP Verify — Requirements, Round 10" in full: R54 (owner opt-out enforcement, five sub-items), R55 (capability-classification union, three sub-items), R56 (duplicate dispatch, diagnosed then fixed at its real root cause), R47 (the remaining benchmark-citation contradiction), R39 (the sixth request, closed by documenting a genuine non-gap), and R57 (schema-divergence card-tools extraction). Found and verified against two fixtures at v1.0.518: sudomichael/gizmoanalytics-mcp (owner opt-out via robots.txt, 32 tools live-enumerated anyway) and awesome-malinoto/tracepass-mcp-server (5 write tools misread as 0, real server card whose tool list was never parsed). 418 tests passing, up from 403 at the start of this round.
Standing order (§0): classifier freeze honored
R55's fix is a wiring change (union an existing classification signal into existing consumers), not a new classifier rule — same category as R48 in round 15.
R54 — Enforce owner opt-out at the request layer, not the renderer
Root cause
RemoteValidationService.validate_server_record() fetched robots.txt after five other live requests had already fired, and the resulting disallow signal only gated 2 of 26 downstream probes via ad hoc per-call-site conditionals. Confirmed live: gizmo's page showed 32 tools enumerated, full risk-claim text for action_safety_probe, and multiple rendering contradictions on a server that declined validation.
R54.1/R54.2 — restructured validate_server_record: the robots.txt fetch (probe_noise_resilience) now happens first, before any other outbound request. When it disallows validation, every one of the other 26 probes (OWNER_OPT_OUT_GATED_PROBE_NAMES, a new module-level constant) is stamped not_assessed/owner_opt_out without ever reaching its own call site — robots.txt itself is the only request made that run. compute_summary_status gained an early-return: initialize_status == "not_assessed" → "unknown" (reusing round 13's R28 status, not a new value — the epistemic state is the same: genuinely don't know).
R54.3 — a second, independent bug: the Checks table's risk-claim text (render_validation_evidence/summarize_check_outcome, main.py) reads ValidationRun.checks directly, a data path redact_opted_out_insights() never touches. Fixed at the shared choke point: summarize_check_outcome now returns a neutral "Not assessed" string for any check whose status is not_assessed/owner_opt_out, before any per-check-name branch runs.
R54.4 — four specific rendering contradictions, each a different stale or wrongly-sourced field, fixed individually: header "Status" (was raw server.current_status, not an insights key), "Confidence" (a fallback heuristic read raw validations/last_validated_at, bypassing the already-correctly-redacted evidence_confidence dict), "Percentile pending" (reused the existing block-for-production suppression pattern), and "Live capability counts" tool count (fell back to raw server.tool_count instead of the already-redacted tool-name list).
R54.5 — scripts/backfill_owner_opt_out_evidence.py (new, dry-run-first per this engagement's standing backfill practice): purges stale, pre-fix probe evidence from the latest ValidationRun for servers already owner_opted_out, recomputing summary_status via the same compute_summary_status the live path uses. Not run against production in this pass — flagged for an explicit go-ahead, same as every prior round's backfill.
R55 — Capability-classification union across write-governance consumers
Root cause
Confirmed live and traced to exact code: capability_distribution: {"read": 4, "write": 5} on tracepass, declared_non_read_only_tools: 0 (readOnlyHint simply wasn't set on the 5 write tools). A bare "write" classification produces no risk flag at all (detect_tool_risk_flags only maps delete→destructive, exec→command_execution) and only "medium" risk level — so high_risk_tools/destructive_tools/exec_tools/ declared_non_read_only_tools all read 0, and every consumer answered "safe."
R55.1 — new has_non_read_capability boolean, computed once in summarize_tool_security_inventory ({write, delete, exec, financial} capability presence), returned alongside the existing aggregates so every consumer reads the same precomputed value.
R55.2 — audited and fixed every consumer: build_action_safety_probe's own status gate, build_write_action_governance's risky_action_surface and blast_radius, search_fetch_only/write_actions_present/ safe_for_company_knowledge/safe_for_messages_api_remote_mcp (one found dead-code correct in _is_search_fetch_only_surface, never wired in — left as-is, out of scope), write_confirmation_required_but_absent (insights.py's benchmark task and main.py's write-safe badge — the third independent site found reading declared_non_read_only_tools in isolation), and the hosted-runtime gate (build_runtime_readiness, which previously read neither signal at all). The hosted-runtime gate reuses blast_radius, not safe_to_publish — safe_to_publish requires positive action_safety_probe evidence to exist at all (the right bar for marketplace publishing), which would have wrongly blocked activation for a legitimate, evidence-free, zero-tool server; caught by a test failure during implementation and corrected.
A precedence bug found and fixed during implementation: unioning has_non_read_capability into write_confirmation_required_but_absent caused it to newly override an action_safety_status == "error" case with a milder "degraded" benchmark-task status — previously masked because the annotation-only signal rarely co-occurred with an already- error probe. Fixed by checking action_safety_status == "error" first, unconditionally, in build_benchmark_tasks's status expression.
R55.3 — new parametrized CI test (test_insights_goldens.py) proving no tool classified write/delete/exec/financial, with zero annotation corroboration, can satisfy search_fetch_only or safe_for_company_knowledge.
R56 — Duplicate dispatch: the doc's citation is stale; a different, live bug was found and fixed
The doc's specific tracepass evidence (three near-second runs, tools: 0) is real but is pre-fix production data — all three job rows were created 7-12 hours before round 14's job-dedup guard actually deployed, and all three runs actually show tools_list: ok with a full 6-tool result, not zero. Direct database querying instead found a different, currently-live bug: 437 near-duplicate validation pairs in ~24 hours, all strictly after round 14's fix deployed, concentrated on fast-failing servers.
Root cause, confirmed in code: iter_validation_queue_candidates takes a single unfiltered ~75k-row streamed snapshot at the top of enqueue_due_validation_jobs. Production runs 3 worker replicas polling every 15s with no cross-replica coordination beyond the DB-level partial unique index — which only protects coexisting pending/running job rows. A failing server's validation completes in ~10ms; if one replica completes a validation for server X while another replica's 75k-row scan is still mid-flight, the second replica's stale view still shows X as "due, not active," and by the time its loop reaches X, the first replica's job is already completed — outside the unique index's protected window, so a genuinely new job is inserted.
R56.1 — a fresh, single-row re-check of should_validate_server immediately before each insert in enqueue_from (scheduler.py), closing the staleness window without a full re-scan or a new locking primitive. Verified both directions: a stale candidate whose real row was already revalidated is correctly skipped; a genuinely-still-stale candidate is still enqueued exactly as before.
R56.2 — confirmed no change needed: validate_server_record's ValidationRun creation and final Server mutation commit together in one transaction per call; the bug was scheduling, not partial persistence (mirrors round 15's R50.2 conclusion).
R56.3 — compute_summary_status's tools_status != "ok" collapsed every reason into "failing", conflating "genuinely broken" with "reachable, but this specific call needed auth." New "unverified" status for tools_status == "auth_required" when initialize itself succeeded. Producing side: _jsonrpc_check previously only special-cased 401 for name == "initialize" — extended to tools_list (401/403, matching _optional_jsonrpc_probe's existing threshold for prompts_list/resources_list; initialize's own condition left unchanged). Wired into SUMMARY_SCORE_MAP, runtime_hosting.py's blocklist-style "not passing" check, and main.py's status filter/label/class maps — per the established R28 rollout precedent (update score map, DB filter surfaces, and blocklist-style checks explicitly; allowlist-style checks default new values to the conservative branch and don't need touching).
R56.4 — build_evidence_confidence's history_depth exclusion (round 15's R50.4, which excluded "unknown") extended to also exclude "unverified" runs.
R47 — Benchmark status must derive from cited evidence
build_benchmark_tasks's "Safe write flow with confirmation" task: already correctly unified (round 15) for the fresh-vs-stale action_safety_probe split; the remaining tracepass contradiction resolves as a side effect of R55 (both action_safety_probe's status and write_confirmation_required_but_absent were downstream of the same broken signal), confirmed via the precedence-bug fix above.
build_compatibility_fixtures's "Anthropic remote MCP fixture": a genuinely separate, structural gap — confirmed in code that safe_write_review (the displayed assumption) was computed independently and never fed into the fixture's own status expression at all (not a fresh-vs-stale duplication — a signal simply missing an and clause). Fixed by adding write_action_governance.safe_to_publish to the gate.
R39 — Sixth request: mostly already shipped; the remaining gap documented as a non-gap
Round 15 already delivered the generated, CI-enforced component × status matrix this item keeps asking for. This round's specific ask — extend it to cover official_registry_probe/score_official_registry_presence — was investigated by reading the actual function signatures: score_official_registry_presence(session, server) and score_cross_source_trust(session, server) take **no checks parameter at all**. There is no check-status axis to vary for them; their score is a pure function of DB peer lookups and static server columns, completely independent of official_registry_probe's own status. Executing them once per status would produce a column where every value is identical — not a meaningful table entry. Documented explicitly in the generator's module docstring rather than silently, so a future round doesn't re-diagnose the same non-gap a sixth time.
R57 — Schema-divergence: comparable card dimensions were being skipped
Root cause, confirmed in code: card_tools extraction read a top-level "tools" key; tracepass's card nests its tool list under capabilities.tools as 6 bare name strings, exactly matching all 6 live tools. server_card_payload.get("tools") returned None, so the probe fell into no_server_card_tools/compared_tool_count: 0 despite the card genuinely declaring a complete, accurate manifest. Fixed with a fallback to capabilities.tools when the top-level key is absent, and bare-string entries are now recorded as name-only membership (compared for tool_membership) routed through the same tools_with_omitted_schema treatment round 15's R49.1 established for dict-shaped tools missing inputSchema, rather than silently dropped.
Test suite
418 tests passing (PYTHONPATH=verify/src pytest verify/tests), up from 403 at the start of this round.
Deferred / carried forward
None new this round — every item in the requirements doc was addressed (implemented, or, for R39's specific extension ask, investigated and documented as not meaningfully implementable the way the doc's framing implied). R56.1's second, complementary optimization from the architecture plan (filtering iter_validation_queue_candidates's SQL query itself to shrink the scan window) was deliberately not built this round: the per-insert re-check already closes the correctness bug, and a SQL-level filter carries real risk of wrongly excluding a legitimately-due server if its WHERE clause doesn't exactly mirror should_validate_server's several threshold branches — a worse failure mode (silent starvation) than the bug it would optimize away. Flagged as an available follow-up, not required.