MCP Verify — Consistency Pass Implementation Summary
Companion to docs/mcp-verify-consistency-pass-architecture-plan.md. Shipped 2026-08-11 at v1.0.549, baselined on v1.0.548. 525 tests passing (up from 518).
What shipped
Freshness (§A of the plan). Confirmed via code read that the homepage and Trust Index pages already shared one query (build_public_coverage_stats), one timestamp field (Server.last_validated_at), and one canonical window setting — the "Trust Index Fresh reads as almost the whole inventory" symptom was a label collision, not a data-source bug: the homepage's card showed a stricter, status-gated stat (healthy_and_fresh) and the Trust Index card showed a looser, any-outcome stat (fresh_validations), both bare-labeled "Fresh." Fixed by relabeling the homepage's card "Healthy & Fresh" and adding an explicit "any outcome included" note (with a cross-reference) to the Trust Index card. Also found and fixed a genuinely triplicated magic number: the 48h "strong live" window was independently hardcoded in main.py and db/repository.py; it's now one constant (STRONG_LIVE_CANDIDATE_FRESHNESS_HOURS, db/repository.py) that main.py imports.
Plan SLA consistency (§B). Found three independent, mutually contradicting tier taxonomies (TRUSTOPS_TIERS, VERIFY_PLAN_DEFINITIONS, trustops.py's TRUSTOPS_TIER_ORDER), plus /pricing's own hardcoded literal tuple that agreed with none of them (it showed Pro=24h/Enterprise="4-12 hours" — numbers that appeared nowhere else in the codebase). TRUSTOPS_TIERS is now the single source /pricing, /trustops's tier table, /trustops's plan cards, and /docs/scoring-specification's "Windows in use" table all render from via one new helper, plan_entitlement_freshness_label(). Since the real scheduler (workers/scheduler.py) has zero plan-tier awareness, every figure is now explicitly labeled "target (best effort)" or "SLA (guaranteed)" rather than implying uniform enforcement — Community/Pro are targets, Enterprise is the one guaranteed tier. Server-detail's "Policy SLA" badge now defaults from the server's own active billing tier (server_billing_tier_freshness_sla_hours) instead of a blanket 168h shown for every server regardless of plan.
Readiness vs. policy (§C). Production readiness and executive verdict were already reconciled in an earlier round (2026-08-10 "Track 1" fix) — not the contradiction this document's §5 assumed. The genuinely distinct, unreconciled concept was build_policy_export_payload (the /policy export endpoint's allow/requires_human_approval computation), which wasn't rendered on the server page at all before this pass. Added a labeled "Recommended runtime policy" panel to the Risks tier (render_recommended_runtime_policy) plus a new Methodology section explicitly distinguishing the two questions. Fixed the one real field-naming collision found: production_trust_decision is reused for two different shapes across build_trust_snapshot (full dict) and trust_index_item_from_server (bare code string) — added an unambiguous production_readiness_code sibling field to the latter, keeping the old field for backward compatibility.
Evidence semantics (§D). One real gap survived the prior round's fixes: a never-validated server's write-action governance panel could show "No high-risk tools were detected" (the "not_assessed" special case didn't cover build_write_action_governance's actual default of "missing" for that case). One-line fix extending the condition to both statuses.
Guardrails (§E/§13). Added coverage_stat_invariants() — four subset-relationship checks over real coverage data (e.g. healthy_and_fresh <= healthy_servers) — wired into /v1/build alongside the existing regression checks, so an impossible cross-stat state is observable without a dashboard. Added 8 new tests: freshness label consistency, plan SLA figure agreement across /pricing//trustops, recommended-policy panel distinctness, the production_readiness_code field, the write-action-governance fix, and both the unit and HTTP-level invariant checks.
§18. Removed one unsubstantiated "the first MCP Trust Index" claim from the llms.txt-style crawler growth-prompt copy.
What was found already satisfied (no new work needed)
- PRICE1 (pricing scaling factors):
/pricing's Pro tier already explains what drives its $99–$499/month range (API call volume, historical evidence packs, webhook access) from an earlier "Track 2" round. - S3/S4 (progressive disclosure polish): no concrete over-expanded section was identified during this pass; per the plan's own guidance, skipped rather than restructuring speculatively.
Deliberately not done this round
Real, plan-tier-differentiated validation scheduling (i.e. actually making the scheduler read BillingAccount.tier) — flagged as Open Question 1 in the architecture plan and left as a product decision. The fix here makes the display honest about what's guaranteed vs. targeted; it does not change backend enforcement.
Verification
PYTHONPATH=verify/src python3 -m pytest verify/tests -q --no-cov: 525 passed.python3 scripts/export_openapi.py --check: fresh (no route signature changes affect the schema;/v1/buildgained a DB dependency and two response fields, but itsresponse_model=None).test_score_component_registry.py: unaffected (no scoring-dimension or zero-point-justification text changed this round).- Live verification against the deployed site to follow this document (see CHANGELOG
[1.0.549]).