Sentinel Signal

MCP Verify — Round 11 Implementation Report (post-1.0.511 verification)

Source: docs/mcp-verify-round-11-implementation.md

Document Content

MCP Verify — Round 11 Implementation Report (post-1.0.511 verification)

Implements the P0/P1 items from mcp-verify-requirements-round5.md — a third external verification round, found against ai.rams/rams (live v1.0.511) cross-referenced against ai.pyrimid/pyrimid (v1.0.510). All P0 items shipped (R29, R30, R23 restated, R25.3), plus P1 (R31), plus one P2 item (R26) that turned out cheap to fully fix once located. Remaining P2 items (R27, R28, R12/R7/R9/R10 restated, R17/R19/R20/R14c/R15) are unchanged from round 10 and explicitly not started this round — see "Deferred" below.

R29 — Percentile was contradictory across servers

Root cause (R29.1): ServerRepository.score_percentile() ranked the caller's freshly-computed public_display_score() against every other cohort member's raw, possibly-stale Server.current_score database column. public_display_score()'s own docstring already documented that the persisted column "can drift from what a fresh recompute of the exact same components produces" — but score_percentile was the one caller still reading the stale column for the comparison population, even though every public surface (including the same page's own "Score" display) shows the fresh value. A server revalidated before vs. after a scoring-logic change (e.g. round 10's R23 zero-anchoring fix) carries a current_score computed under different rules than one revalidated after — comparing one server's fresh number against a corpus of columns computed under a mix of rule versions produced exactly the observed contradiction (pyrimid: 74.3, top 17.2%; rams: 73.0, a lower score, top 6.2%, a better percentile).

Fix (R29.2): score_percentile now loads the cohort's Server rows (same filters as before — live-validated, non-Failing, tool_count>0) and recomputes public_display_score() fresh for each one, ranking on that instead of the raw column. The result dict gained a cohort_definition field, rendered as a trailing sentence in format_score_percentile()'s output so a percentile always states what it's a percentile of, not just its size.

R29.3 — cross-server consistency test: new verify/tests/test_percentile_integrity.py. One test reproduces the exact live contradiction shape (two servers with deliberately swapped/inverted persisted current_score columns relative to what their own components compute fresh) and asserts the higher-displayed score never lands in a worse percentile. A second test checks pairwise monotonicity across a 12-server synthetic cohort with scrambled stale columns — not a 2-server sample, per the doc's own instruction.

R30 — MCP tool annotations as primary capability evidence

The live bug: ai.rams/rams's review_files tool carries {"readOnlyHint": false, ...} but classified read anyway — a stray "get" inside ordinary prose ("get their OK before starting one," in the tool's own token-cost warning) matched the text-hint fallback, because the old gate (if annotations.get("readOnlyHint") is not True) treated an explicit false the same as an absent annotation.

R30.1: detect_tool_capabilities() in validation/service.py now splits the text-hint fallback: "write" hints apply whenever readOnlyHint isn't True; "read" hints only apply when readOnlyHint is genuinely absent (None). A final hard guard discards read outright whenever readOnlyHint is False, regardless of what produced it. openWorldHint: true now also adds a network capability (previously unused). destructiveHint: true already correctly added delete/destructive_operation — confirmed, not changed.

R30.2: new detect_tool_annotation_conflicts() — when readOnlyHint: true is asserted but the schema shows positive structural evidence to the contrary (a command/script-shaped parameter, an admin-scoped parameter, a credential-shaped parameter, or destructiveHint: true on the same tool), it's surfaced as a finding on the tool's tool_security_inventory entry (annotation_conflicts), not silently resolved either direction. Only schema-shape evidence, same hint sets the classifier itself uses — no free-text keyword matching for accusatory classes, per R5's standing rule.

R30.3 — broke the propagation chain. The reported chain: readOnlyHint: false ignored → classified read → Write Actions Present: No → Search Fetch Only: Yes → Safe For Company Knowledge: Yes — "a tool that ingests arbitrary source files and consumes metered quota is certified as a search/fetch-only surface safe for company knowledge." Fixed at two levels: R30.1 above closes the root cause, and as defense in depth, build_tool_security_inventory() now records each tool's raw read_only_hint (tri-state), and build_publishability_policy_profiles() in insights.py vetoes search_fetch_only/safe_for_company_knowledge whenever declared_non_read_only_tools > 0 — independent of whatever the capability classifier concluded, so a future classifier bug can't resurrect this exact failure mode. Verified end to end against the real ai.rams/rams fixture.

R30.4 — corpus re-run: NOT executed. Production-scale write, same category as R6/R24.4; not run without explicit confirmation.

R23 (restated) — the mapping table, not another spot fix

OAuth Interop, the worst instance yet: score_oauth_interop()'s if not protected_ok and initialize_status == "ok": return 7.0 branch gave 3/4 points on a server with no OAuth metadata at all (oauth_protected_resource error/404, oauth_authorization_server/openid_configuration never even ran) — self-contradicting the same page's OAuth: ✕, DCR/CIMD: ✕, and two High-severity OAuth remediations. This metric's own description is "depth ... beyond the minimal protected-resource check" — it has nothing to measure depth of when that check itself never succeeded, regardless of whether the overall handshake did. Fixed: zero-anchored whenever oauth_protected_resource isn't ok, full stop. Verified against the real fixture: raw 0.0, 0 points.

Interactive Flow Safety / Request Association: confirmed, by direct fresh recomputation against the exact same fixture checks, that round 10's existing fix already returns None (excluded) for both on this server — the 3/4 shown in the live evidence table is persisted staleness (ai.rams/rams hasn't been revalidated since round 10's fix deployed), not a live gap. No code change needed; documented here so a future round doesn't re-investigate the same non-bug.

**auth_required semantics**: confirmed already correct without changes. score_prompt_contract/score_resource_contract already treat live_status == "auth_required" as a distinct branch (5.5/6.0) from missing (0/None after round 10's fix) — auth_required scores better than missing, matching this round's explicit instruction ("evidence that auth exists, not evidence the capability is absent").

R23.1 — mapping table regenerated. ml/registry/generate_score_component_registry.py (built round 10) re-run after the OAuth Interop fix; its own drift-detection test (test_registry_is_generated_and_up_to_date) caught the exact source-check change (the fix dropped initialize as a dependency) before this could be committed stale — direct confirmation the mechanism works as designed.

R23.2 — CI assertion: the pre-existing whole-corpus enforced-ceiling test already covers the "everything's broken" case for the now-corrected oauth_interop_score zero point. New targeted test (test_r23_oauth_interop_zero_anchored_when_protected_resource_check_failed) pins the exact live fixture shape.

R25.3 (restated) — absent sources cannot produce agreement

The bug: ai.rams/rams's server_card 404s (no card evidence at all), but its registry payload is present. Every drift check in build_provenance_divergence_probe() requires both sides non-empty to even run — round 8's earlier fix only guarded the zero-readable-sources case, so one readable source fell through to "no drift fields found" → status: ok → "no source disagreements detected," on a comparison that only ever had one side to compare.

Fix: the probe now counts readable sources (registry, server_card) and returns not_assessed whenever fewer than two are readable — not just when both are absent. score_provenance_divergence() returns None (excluded, same mechanism as R17/R18) for not_assessed. The probe's details now also states readable_sources explicitly, alongside the existing compared_fields, so a not_assessed verdict is self-explanatory.

Pattern-class grep sweep (per the standing instruction — this is the fourth+ instance of "unknown reads as pass" after R16/R18/R20, so the instruction was to grep for the class, not just fix the reported case): a background agent swept validation/service.py/insights.py for the same shape. Most divergence/consistency probes were already correctly guarded. One genuine, reachable instance found and fixed: score_registry_consistency()'s five comparison helpers (text_similarity, host_match_ratio, transport_match_ratio, numeric_consistency_ratio, bool_consistency_ratio) each returned 0.6 — a passing-adjacent default — whenever both sides were absent, folded directly into an fmean() average with no exclusion. A sparse server (no registry payload, no homepage, no metadata documents) silently accumulated several unearned 0.6 credits. Fixed the same way: the helpers now return None for "nothing to compare," and score_registry_consistency excludes None values from the average rather than defaulting to 0.6 (or, if nothing at all was comparable, to 0.0). One borderline candidate (insights.py's tool_change_documented, a different flavor — "no snapshot change" read as "documented" — vacuous-true on a no-op rather than on missing data) was noted but not fixed this round; flagged for a future pass.

R31 — a third fixture for R25.1

Neither existing fixture could keep testing schema divergence going forward: pyrimid's finding is now a fixed, cache-pinned case; rams's server_card 404s outright (that absence is what R25.3 tests instead). Reused the existing ai.dreamlit/mcp fixture (frozen in round 9, healthy, server_card status ok, 11 real tools) — a genuinely different divergence shape than pyrimid's renamed/recased parameters: the card lists every tool by name/description but never declares a per-tool inputSchema at all, so every live-declared required parameter reads as "only in live." New test: test_r31_schema_divergence_probe_against_a_third_independent_fixture.

R26 — validation-diff rendering, fully fixed (not just confirmed)

The round-5 doc marked this "Unconfirmed" since it rendered cleanly on the ai.rams/rams fixture. Traced the code directly instead of relying on fixture luck and found two real, still-live bugs:

  1. render_validation_diff() used value or "" for the previous/latest/
  2. delta table cells — 0.0 is falsy in Python, so a component that genuinely dropped to zero rendered as a blank cell, indistinguishable from "no data." This is exactly the originally-reported "Previous 1.0, Latest blank, Delta -1.0." Fixed: only an actually-None value renders blank now.

  3. build_validation_diff() used .get(key, 0.0) when comparing the two
  4. validations' score_components dicts — a component whose key is simply absent (because it now legitimately returns None/excluded, per this and last round's R23/R25.3 fixes) was compared against a fabricated 0.0, producing a misleading "dropped to zero" delta for a dimension that was never actually failing, just no longer assessed. Fixed: components present in only one side are tracked separately (newly_assessed_components/newly_excluded_components, both rendered on the page) rather than folded into component_deltas.

Both reproduced directly (not fixture-dependent) in verify/tests/test_score_integrity.py.

Test suite

355 tests passing (PYTHONPATH=verify/src pytest verify/tests), up from 334 at the start of this round.

Deferred to a follow-up round

Unchanged from round 10's own deferral list, still not started — lower priority than the P0/P1/R26 work above, each substantial enough on its own that rushing them alongside this round's work would have meant shipping them half-verified:

  • R27 — duplicate validation runs, now one second apart rather than
  • identical. Needs root cause (still-unresolved double-dispatch) before a display-only dedupe would be safe.

  • R28 — historical Healthy + tools: 0 rows persist. Forward-going
  • gate already covered by R18; backfill of existing bad rows not attempted (needs a dry-run count first, same discipline as any corpus-scale write).

  • R12/R7/R9/R10 restated — full funnel on the methodology page, 50-tag
  • taxonomy spot-check, remaining conclusory verdict-label lint, disclosure language. Not started.

  • R17, R19, R20, R14c, R15 — unchanged, carried forward.
  • R11 — ground-truth rubric, still blocking R24's full closure; still
  • needs human reviewers this session cannot recruit.

  • One borderline pattern-class candidate noted but not fixed:
  • insights.py's tool_change_documented (vacuous-true on a no-op, not on missing data — a different flavor of the same family, lower confidence).