MCP Verify — Round 14 post-1.0.530 implementation
Initial target release: 1.0.533. Same-day R73 follow-up: 1.0.534.
Source: "MCP Verify — Requirements, Round 14 (post-1.0.530 verification)", verified against one fixture: dan-rozhkov/webharvest (trustsnap_8404e4709c79147d) — the first Failing fixture in eight rounds. Every claim in the architecture plan was checked against the live production API response for this fixture before any code changed.
Implemented changes
- R70: per-client remediation checklists render the human remediation
- R71.1: a server whose latest validation run was attempted and failed
- R71.2: publishability profiles (
chatgpt_custom_connector, - R72: six security-posture score components (Destructive Operation
- R73: an unreachable or non-MCP registry endpoint (e.g. a directory
action text again. render_client_remediation_modes checked item.get("check") (an internal identifier such as openai_connectors:2 or a gate key such as search_fetch_only) before item.get("label"), and every current producer always sets check — so the fallback to human text was structurally unreachable. Now reads action / label / observation / check, in that order.
is reported as Failing, not Metadata only. infer_server_verdict previously returned metadata_only whenever tool_count <= 0, regardless of whether a real run had happened — conflating "never validated" with "validated, and the run failed before enumerating anything." last_validated_at is None is now the only signal for metadata_only; current_status == "failing" with a real last_validated_at now returns a new failing verdict code, propagated through build_production_readiness, the write-safe-publishing gate, runtime_hosting's blocker set, the executive-verdict blocking check, and the ?verdict= filter dropdown.
claude_remote_mcp) now correctly report blocked status when the underlying client-compatibility profile is genuinely blocked. The status ternary previously reached "blocked" only when the profile dict itself was empty; a real, non-empty profile with compatibility == "blocked" always fell through to "caution" ("Compatible with review"), contradicting the client-compatibility verdict computed from the same data.
Safety, Egress/SSRF Resilience, Execution/Sandbox Safety, Least Privilege Scope, Secret Handling Hygiene, Abuse/Noise Resilience) are now excluded from scoring — not credited a fixed partial value — when the underlying tool surface was never confirmed observed (not is_core_success_from_check_results(checks)), matching the same not_assessed condition action_safety_probe already uses. A confirmed-healthy server with a genuinely empty tool surface is unchanged (still scores the prior fixed value) — no fixture evidence this round said that case needed to move. render_score_breakdown now shows an explicit "N component(s) not assessed for this run" line instead of silently omitting excluded rows.
listing page returned as text/html) now produces a medium-severity, registry-directed registry_endpoint_invalid finding instead of a critical, publisher-directed server_failing alert. parse_mcp_response raises a distinct NonMCPContentTypeError for a non-JSON, non-SSE content type rather than falling through to a blind json.loads() that produced a generic, indistinguishable JSONDecodeError. Publisher- directed remediation text ("Retest against {remote_url}...") is suppressed in favor of registry-directed text when this classification applies.
Live post-deploy re-check found the underlying evidence had shifted. Deployed to v1.0.533 and manually re-triggered a validation run against dan-rozhkov/webharvest inside the production container (bypassing the scheduler queue for an immediate result). R72's fix confirmed correctly end-to-end. R73's did not fire: the registry's remote_url had moved from returning 200 text/html to a 301 redirect (Glama restructured their listing URLs between when this round's plan was written and when it was deployed), which raises httpx.HTTPStatusError at response.raise_for_status() — before parse_mcp_response ever runs, so NonMCPContentTypeError never fires. Separately, the classifier originally required initialize/oauth_protected_resource/server_card to unanimously carry the marker (matching the doc's "every probe fails identically" framing); live data showed oauth_protected_resource/ server_card now fail with unrelated, legitimate 404s instead. Fixed same-day, same deploy cycle: both _fetch_json_check and _jsonrpc_check's HTTPStatusError handlers now also tag non_mcp_content_type as "redirect:{status_code}" when response.is_redirect; _non_mcp_endpoint_content_type now keys off initialize alone (authoritative on its own — the one call that must return real MCP JSON-RPC output — rather than requiring agreement from two optional .well-known probes that legitimately 404 on many real servers). Re-verified against the same live re-validation: registry_endpoint_invalid now fires correctly, server_failing does not.
- R74: removed
"running"from thefitnesstaxonomy marker list — a - R39 (seventh request):
score_prompt_contractand
genuine whole-word match (not a substring bug; word-boundary matching was already correct) that fires on ordinary tech prose ("running locally") as readily as on genuine fitness content. The requirements doc's second data point (ai.com.mcp/strava tagged other) did not reproduce against the live corpus at v1.0.530 — confirmed tagged fitness correctly, and an existing regression test already pins this; no code change made for that claim. "cycling" carries a similar risk but was left in place — no corpus-wide prevalence evidence this round said it's actually misfiring.
score_resource_contract now zero-anchor a skipped check status the same as missing. _skip_check's own docstring confirms skipped is used exclusively for "never attempted because a different check already failed" — the same evidentiary meaning as missing — but these two functions previously let skipped fall through to a generic mid-credit default, scoring strictly better than a check that had even less chance to demonstrate anything. The existing ml/registry/generate_score_status_matrix.py (built in round 15, round 14 confused for it) already correctly modeled three of the four cells the requirements doc cited; the actual gap was this specific skipped mishandling plus a genuinely mixed-status case (prompt_get=missing, prompts_list=skipped simultaneously) outside that generator's documented single-status-uniform scope, covered by a direct unit test instead.
Registry/generator changes
ml/registry/generate_score_component_registry.py's AST-based not_assessed_only_exclusion check previously recognized only a literal comparison against the string "not_assessed" as a valid not-assessed guard. R72's six components express the same genuine not-assessed reasoning through a call to is_core_success_from_check_results(checks) instead (the exact predicate action_safety_probe itself uses to decide its own literal "not_assessed" status) — the generator's guard-detection now also recognizes that call. test_r39_2_only_genuine_not_assessed_components_may_exclude's hardcoded exclusion_capable allowlist was extended from 3 to 9 components. Both score_component_registry.json and score_component_status_matrix.json were regenerated.
Not changed, deliberately
score_data_exfiltration_resiliencehas the identical `if not"cycling"in the fitness taxonomy marker list (see R74 above)._exclude_metadata_only_filters(the SQL-level default-browse
tool_inventory:` shape as the six R72 components but was not named by this round's fixture evidence — left unchanged, consistent with the standing "fix what's evidenced" discipline.
exclusion) — its job (hide tool_count <= 0 servers from default listings) is orthogonal to the verdict label fixed by R71.1; a failing, zero-tool server should still stay out of default browse results either way.
Validation
PYTHONPATH=verify/src pytest verify/tests -q— 467 passed (up frompytest(root repo,app/token_service) — 293 passed, unaffected- `python3 scripts/validate_test_quality_gates.py --coverage-json
python3 scripts/export_openapi.py --check— fresh after regeneratingPYTHONPATH=verify/src python3 ml/registry/generate_score_component_registry.py
453).
(no files outside verify/, ml/registry/, docs/, CHANGELOG.md, and version-artifact files touched).
artifacts/test-reports/coverage.json --min-line 70 --min-branch 55` — line 75.71%, branch 60.85%, both above gate.
for the version bump.
and generate_score_status_matrix.py — both regenerated and re-pinned.
Live re-verification (post-deploy) — confirmed
Deployed to v1.0.533, then v1.0.534 after the R73 follow-up fix. dan-rozhkov/webharvest's stale, pre-fix validation run was manually re-triggered inside the production container both times (bypassing the scheduler queue for an immediate result, rather than waiting on its natural revalidation cycle), and every claim below was independently re-fetched from the live public API afterward, not assumed from a successful deploy alone:
- R70:
client_remediation_modes[].why_not_ready[].labelrenders as - R71.1: `production_readiness == {"code": "failing", "label":
- R71.2:
publishability_policy_profiles[].status == "blocked"for - R72: all six components (
destructive_operation_safety_score, - R73: `active_alerts == [{"code": "registry_endpoint_invalid",
- R74:
taxonomy_tags == ["search", "web"]—fitnessno longer - R39:
prompt_contract_score == 0.0, `resource_contract_score ==
prose ("Add OAuth-based authentication for remote connector auth.") — no openai_connectors:2-style keys anywhere.
"Failing", "reason": "The latest validation run was attempted but did not complete successfully.", ...} — no longer metadata_only`, no longer contradicts its own evidence.
both chatgpt_custom_connector and claude_remote_mcp, matching client_profiles[].compatibility == "blocked".
egress_ssrf_resilience_score, execution_sandbox_safety_score, least_privilege_scope_score, secret_handling_hygiene_score, abuse_noise_ratio_score) confirmed absent from current_score_components on the fresh validation run.
"severity": "medium", "message": "The glama_registry entry's remote_url returns non-MCP content (redirect:301), not a validation finding about this server."}] — server_failing no longer present. Every remediation entry that would have shown "Retest against https://glama.ai/mcp/servers/jppq78k6mg from a clean client session" (fix_initialize_flow, publish_oauth_protected_resource, publish_server_card, respond_registry_endpoint_invalid`) now shows the registry-directed playbook instead.
present.
0.0 on the fresh run, where prompts_list/resources_list are both genuinely skipped`.
The R73 detection mechanism itself needed a same-day follow-up (see the R73 section above) after this first live re-check found the registry's actual failure mode had shifted from a 200 text/html body to a 301 redirect since the plan was written, and that the original unanimous-three-check classifier under-fired against real, independently-caused 404s on the two auxiliary .well-known probes. Fixed, redeployed, and re-confirmed against the same fixture in the same session.