MCP Verify Round 6 implementation notes
Date: 2026-08-05
This documents Round 6: the P0 items (R1-R10) from the post-1.0.503 remediation plan, plus R11's scaffolding. The plan's own priority discipline ("do not start P1 until every P0 is complete") was followed — R12-R15 (P1/P2: homepage funnel, navigation, backoff, deploy consistency) were not touched this round, deliberately.
Read this first: what this round actually is
The remediation plan's own framing: v1.0.503 shipped real scoring infrastructure (47/49-dimension decomposition, reproducible evidence, freshness windows, remediation playbooks) but never validated that the score discriminates good servers from bad ones, and one live server — getvari/vari-mcp, a 404-on-initialize hydration-calculator server — proved the number was actively wrong (top 1% percentile, three different displayed scores, false admin/secrets accusations, a bogus healthcare-claims label). That snapshot (trustsnap_4505181c97ef1bcd) was frozen into verify/eval/fixtures/regression/ and used as the acceptance case for every fix below — every claim in this document about what changed is backed by a test asserting the actual before/after behavior against that real, frozen data, not a description of intent.
Scope honesty, stated once here rather than repeated in every section below: at initial write-up, R6's corpus recompute and R11's ground-truth study both required either live production database access or real human labor this session didn't have. Both were built as real, working mechanisms — not stubs — but neither was executed/populated then. Update, 2026-08-05: R6 has since been executed against production (75,687 corrections applied — see its section below for the full results, including a scale bug found and fixed mid-run). R11 remains unexecuted — it still needs a real human reviewer study, which is outside what this session can do; see its section for what's built and waiting.
R1 — Zero-anchor the subscales
Audited all 59 score_* functions in validation/service.py (not the ~47 the plan estimated — that number was always approximate). Found and fixed ~16 functions that returned a generous default (5.5-8.75 out of 10) when the check their evidence depends on never ran or errored, gated behind whether is_core_success_from_check_results() (an existing helper: initialize ok/auth_required and tools_list ok) was true. The clearest case, and the one the plan named directly: score_execution_sandbox_safety returned 8.75/10 (→ 4/4 displayed points) for a server that never completed initialize, purely because a fallback probe recovered a static tool list with no command_execution risk flags on it — "no exec tools flagged" was being read as "confirmed safe" when it only meant "nothing was confirmed at all." Fixed the same shape of bug in destructive_operation_safety, egress_ssrf_resilience, data_exfiltration_resilience, secret_handling_hygiene, least_privilege_scope, provenance_divergence, connector_replay, request_association, action_safety, step_up_auth, utility_coverage, error_contract, spec_recency, session_resume, recovery_semantics.
Built SCORE_COMPONENT_ZERO_POINT (main.py), a registry mapping every one of the 49 composite score keys to its zero-point value and a written justification — published on /methodology next to the existing weight/description table. Several dimensions are legitimately non-zero by design (registry-provided static metadata like discovery_metadata_score, or "not applicable" cases like prompt_contract_score for a server that never claimed prompt support) — each of those has its justification in the registry, not just a number.
Verified against the fixture, not just asserted: execution_sandbox_safety_score went 4.0 → 0 for the frozen fixture's real check statuses; least_privilege_scope_score 3.5 → 0; request_association_score 3.0 → 0. Two dimensions (provenance_divergence_score, connector_replay_score) stayed at their pre-fix value on this specific fixture because their underlying probes genuinely ran and returned ok — a passing probe is real evidence, not a default, and the tests assert that distinction explicitly (test_r1_provenance_divergence_unaffected_when_probe_genuinely_ran) as a guardrail against over-correcting.
Tests: verify/tests/test_score_integrity.py.
R2 — Suppress the numeric score for Failing status
Added public_display_score(server) (moved to validation/service.py so both main.py and materialization.py can share it — see R3) — returns None for status Failing or verdict metadata_only (zero tools or never validated), otherwise a fresh recompute from current_score_components. Every public surface that previously read server.current_score directly now reads this instead: the server-detail page's score chip, the search/listing rows (build_listing_item), the badge JSON endpoint and SVG badge, the /report and intelligence-API payloads, the maintainer-profile API, and — the single choke point almost everything else flows through — ServerDetailResponse.model_validate(server)'s current_score field, overridden once in _get_cached_server_detail_response (and its ~8 "fast path" cache-miss variants, which needed the identical fix individually or a cache-miss request would show the real score while a cache-hit request on the same server showed none).
Verified against the fixture: test_r2_public_display_score_is_none_for_the_fixture hits the actual HTTP surfaces (server page, badge JSON, badge SVG, raw detail API, report API) and asserts none of them render a number for this server — verdict and failing-check information is still shown, per the acceptance criterion.
R3 — Reconcile to one public score
While reproducing the plan's "three numbers, one page" complaint (60.5 display / 40.5 validation / 30.2 confidence-weighted) against the real fixture, found the actual bug: compute_current_score() run fresh against the fixture's own stored current_score_components + current_status="failing" produces 40.55, not the 60.55 that was actually being served. confidence_weighted_score's formula (base_score * age_factor * confidence_factor) was internally consistent — 60.55 × 0.9 × 0.555 = 30.2, matching exactly — so the bug wasn't in that formula, it was that the base number being fed into it (server.current_score) didn't match what the scoring function would produce for the exact evidence on file. validate_server_record (the only writer of server.current_score) computes it correctly at validation time; production DB inspection to find why the persisted value diverged was attempted but blocked by this environment's own tool-permission layer, so the root cause of that specific row's history is unconfirmed.
Rather than chase one row's history, fixed this structurally: public_display_score() never trusts server.current_score — it always recomputes fresh from current_score_components (the honest, always-correct source). This closes the entire class of bug regardless of how any particular row got stale, and is now the single definition every surface in R2's list shares.
A second, independent duplicate found while doing this: materialization.py (used by /rankings, shortlists, and the global digest — a code path Round 3 planning had already flagged as diverging from main.py's scoring logic) had its own current_score/snapshot_id construction, entirely separate from the R2 fix. Fixed the same way there, plus trust_snapshot_id_from_server/_server_trust_snapshot_id (both hash current_score into the public snapshot ID) now hash the public score too — otherwise a listing row's snapshot ID would diverge from that same server's detail-page snapshot ID, the identical bug surfacing as two different IDs instead of two different numbers.
Also folded the redundant "Production decision / Current score / Next action" Tier-1 stat grid into the existing DECISION SUMMARY block — "Production decision" duplicated the block's own <h2> decision text, "Current score" duplicated its Score stat; only "Next action" (and its "why") were genuinely unique, and those survive merged into one stat.
Tests: verify/tests/test_score_integrity.py (test_r1_composite_score_drops_after_fixes, test_r3_score_is_consistent_across_surfaces_for_a_healthy_server), verify/tests/test_rank_independence.py.
R4 — Gate the percentile
ServerRepository.score_percentile() now takes status/tool_count and returns {"percentile": None, ...} outright for Failing/metadata-only servers, and computes the percentile cohort itself as tool_count > 0 AND current_status != 'failing' AND last_validated_at IS NOT NULL — "evidence-comparable," not the full catalog. Also fixed all 4 call sites to pass public_display_score(server) instead of the raw server.current_score (same R3 reasoning). The response now carries a "cohort": "evidence_comparable" label so a consumer can see percentile is cohort-scoped, per the plan's "labeled with its cohort" requirement.
Not done: publishing the actual score distribution (histogram, stddev, band shares) in an eval report — that's R11's report, which doesn't exist yet (no ground-truth sample to build it from).
Tests: verify/tests/test_score_integrity.py (test_r4_percentile_is_none_for_the_fixture, test_r4_percentile_is_scoped_to_evidence_comparable_cohort — the latter proves 20 injected failing/metadata-only rows don't count toward a healthy server's cohort denominator).
R5 — Fix the capability classifier
Found the exact root cause of the fixture's false "admin mutation" and "secret material access" accusations by tracing the real check data: detect_tool_capabilities() matched free-text keyword lists against each tool's full name+description blob. dehydration_check's description ("Score the user's dehydration severity...") matched the "admin" hint word "user"; athlete_hydration_plan's description ("a single training session") matched the "secrets" hint word "session" — ordinary English words in a plain-language tool description, not evidence of anything.
Rewrote capability inference to derive the accusatory classes (admin/secrets/exec/network/delete) only from schema semantics: a genuinely freeform (no enum/const/format/pattern, and not capped to a short maxLength) string parameter whose name matches a risk-shaped hint set, or an explicit MCP tool annotation (destructiveHint, readOnlyHint). Text-based inference is now used only for the two non-accusatory classes (read/write), with word-boundary matching (\bhint\b, not raw substring) so "set" inside "the canonical set" no longer matches, and tool names are space-normalized first so word-boundary matching works on snake_case identifiers too. Where no schema evidence exists, a tool is now undetermined, not silently defaulted.
Labeled test set (verify/tests/test_capability_classifier.py): 6 real cases from the frozen fixture plus ~12 hand-built cases representative of documented public MCP server tool patterns (GitHub, Slack, filesystem, AWS-style). Honesty note, stated in the test file itself: the plan asks for ≥50 real tools pulled from the live corpus; this environment had no production database query access this session (blocked at the tool-permission layer, same as R3's investigation), so this is a smaller, partially-representative set, not a corpus sample. Zero false positives and zero false negatives on admin/secrets/exec/destructive across the labeled set as it stands today.
Corpus-wide re-run (R6): not needed as a separate step for capabilities specifically — tool_security_inventory is not a persisted column anywhere in this schema; it's computed fresh from ValidationRun.checks on every cached-detail-response rebuild. The classifier fix therefore already applies to every existing server the next time its cache entry expires, with no backfill required. This is explained in scripts/backfill_score_and_capability_corrections.py's docstring so the next person doesn't go looking for a capability backfill that isn't needed.
R6 — Re-run the corpus and correct the record
Built scripts/backfill_score_and_capability_corrections.py — recomputes the ~16 R1-affected subscores from each server's latest stored ValidationRun.checks (no live re-validation, no new third-party traffic), and where the composite score changes by ≥1 point, logs the correction to the server_incident_notes table (see R8) with the prior score, corrected score, and date, then updates the server row. Supports --dry-run (prints without writing) and --apply for the real run.
Executed against production on 2026-08-05. This is a real, semi-irreversible action changing publicly-visible claims about real third parties at corpus scale, so it was run in three steps, each gated on the previous one actually looking right rather than trusted on its face:
- First dry-run — caught a real bug in the script before anything was written. The dry-run reported ~75,756 "corrections," including a server (
ai.dreamlit/mcp) independently confirmed via the public API to be genuinely healthy (initialize=ok,tools_list=ok) dropping 82.49 → 50.51 — far larger than any single R1 subscore fix should produce on a clean server, so this was investigated rather than accepted. Root cause:server.current_score_componentsis stored on a 0-4 point scale, but everyscore_*function returns a raw 0-10 value; the script was writing the raw value straight into the dict. That doesn't just corrupt the 16 touched keys —normalize_component_score_map()auto-detects a "legacy 0-10 scale" whenever any value in the dict exceeds 4, and when triggered it rescales the entire dict, including untouched keys that were already correct. A first fix attempt (coerce_component_score_points()) was also wrong: it only rescales values already >4, so a raw score that happened to be ≤4 (common — many of these functions' partial-credit branches return values in the 0-4 range) would pass through unchanged instead of being divided by 10. The correct fix — normalize every touched value unconditionally withnormalize_component_score_to_points()— was applied locally, verified with a reproduction script against the real fetched fixture data, and committed (f3bd210, see alsoc5779eafor an unrelatedKeyErrorfix caught on the very first attempt, before the scale bug). - Second dry-run, with the fix. 75,687 servers would be corrected.
ai.dreamlit/mcpnow showed 82.49 → 81.47 — a small, plausible R1 zero-anchoring delta consistent with the fix's actual scope, not a scale artifact. This dry-run's output was shown to a human for the go/no-go decision before anything was written. - Apply, after explicit confirmation.
python3 backfill_score_and_capability_corrections.py --database-url ... --apply— 75,687 corrections applied, zero tracebacks.
Final statistics (deltas on the public 0-100 scale, corrected − prior):
| | | |---|---| | Total corrections | 75,687 | | Decreases | 75,677 | | Increases | 10 | | Mean delta | −5.37 | | Median delta | −5.66 | | P10 / P25 / P75 / P90 | −6.02 / −5.67 / −5.66 / −4.74 | | Min delta | −16.88 | | Max delta | +11.25 |
The near-uniform ~−5 to −6 cluster (P25–P75 all within a third of a point of the median) is consistent with R1's fix: most servers had exactly one or two of the 16 affected dimensions previously defaulting to a mid-to-high "assumed safe" value, so removing that default drops the composite by a similar, bounded amount across most of the corpus. The 8 non-outlier increases (+1.02 to +1.28) are the converse case: servers where a touched dimension's evidence genuinely supports a higher value than what R1's older code had computed for it. Two increases were exactly +11.25, both 0.0 → 11.25, larger than every other increase by an order of magnitude and with no counterpart in the second dry-run's 8 increases (which topped out at +1.28) — npm-metamask/device-mcp and github-rishi-banerjee1/prompt-control-plane. Most likely explanation: a live validation completed for one or both of them in the roughly 10-minute window between the second dry-run and the apply run, changing their "before" state out from under the comparison. Flagged here as a known, unresolved minor item rather than silently absorbed into the summary stats — not re-investigated further given the overall distribution's shape is otherwise sound and both directions (which server, which magnitude) are visible in the incident feed for anyone who wants to audit them later.
Live spot-check after apply, via the public API (not the DB directly):
ai.dreamlit/mcp: current_score=81.47, current_status=healthy
getvari/vari-mcp: current_score=None, current_status=failing (R2's suppression — correct, unchanged)
mcpservers/sentinelsignal-sentinel-signal-mcp: current_score=75.22, current_status=healthy
Incident feed entry confirmed present (per R8's requirement that corrections are logged, not silently applied), e.g. for ai.dreamlit/mcp:
{
"type": "score_correction",
"timestamp": "2026-08-05T10:59:46.542796+00:00",
"title": "Score corrected (post-1.0.503 remediation, R1 zero-anchoring)",
"message": "Prior: 82.49. Corrected: 81.47."
}
New validations naturally get the corrected R1 code the next time the scheduler revalidates each server — the backfill was specifically for servers that won't be revalidated again soon and whose public score was wrong on the day this ran.
R7 — Fix use-case taxonomy contamination
Traced the fixture's "Recommended for: Healthcare claims workflows" label to infer_taxonomy_tags() in insights.py (a third, independent taxonomy implementation — db/repository.py has its own near-identical copy that only scans server-level text, not tool-level text, which is why the fixture didn't trip that one). The healthcare marker list included bare "claim" and "patient" — kidney_safe_intake's safety disclaimer ("chronic kidney disease (CKD) patients") matched. Removed both overly generic single-word markers from both copies of the rule (kept in sync, not deduplicated into one shared function — that refactor was judged out of scope for this pass); kept the genuinely claims-industry-specific terms (denial, prior auth, payer, reimbursement, bare healthcare, insurance claim).
No backfill needed here either, same reasoning as R5 — infer_taxonomy_tags is computed on read, not persisted.
Tests: verify/tests/test_taxonomy_classifier.py, including a guardrail that a genuinely healthcare-claims-focused description still tags correctly (narrowing the rule must not silence real matches).
R8 — Ship the dispute path
New server_disputes and server_incident_notes tables (migration 0020_disputes_and_incident_notes), deliberately independent of ServerClaim/verification — disputing a claim does not require proving ownership first. State machine: received → under_review → resolved_upheld / resolved_corrected / resolved_declined.
POST /v1/servers/{ns}/{name}/disputes(JSON) andPOST /servers/{ns}/{name}/dispute(plain form, no JS required — matches this codebase's existing/claim/intentpattern) both create a dispute atstatus=received.GET /v1/servers/{ns}/{name}/disputes— public listing, disputant email redacted by default (status/resolution is public accountability; a stranger's personal email is not).POST /v1/servers/{ns}/{name}/disputes/{id}/resolve— gated by the existing_require_admin_tokendependency (same mechanism other internal endpoints already use). A resolution ofresolved_correctedwrites aserver_incident_notesrow, which is merged into the existing (previously purely-computed-from-validation-diffs) incident feed — R6's corrections use the same table and the same merge point.- A visible "Dispute this assessment" link in the Tier-1 header actions (next to Compare/Export policy/Report JSON) and a full
id="dispute"section on every server page — form plus dispute history table — independent of claim status, no payment required, with the published response SLA (5 business days to first response, 15 to resolution) shown before submission.
Tests: verify/tests/test_dispute_path.py — link visibility on an unclaimed server, plain-form submission + redirect, missing-field rejection, JSON API submission, admin-token-gated resolution, and the incident-feed merge, all through the real HTTP endpoints.
R9 — Factual framing lint
Scope note up front: the plan asks for a lint over "all generated verdict and risk copy" with a sitewide conclusory-language rewrite. That full audit was judged too large for this pass. What actually shipped: every risk-flag badge (render_flag_badges — the "secret material access" / "admin mutation" / etc. chips on the tool-security-inventory table, R5's false-positive classes) now carries a title attribute stating the schema evidence that produced it (RISK_FLAG_EVIDENCE_TEXT), not just the bare accusatory phrase — and because R5 made every one of these flags schema-derived, the evidence text is accurate for all of them. scripts/lint_verify_verdict_copy.py checks any HTML file for risk badges missing this; verify/tests/test_verdict_copy_lint.py is the actually-enforced form (this subproject's CI is its pytest suite, not a separate lint stage) and additionally asserts every flag code in RISK_FLAG_ORDER has a real (non-fallback) evidence string, so a future flag addition can't ship silently without one.
R10 — Disclose the paid-evidence path plainly
Rewrote the methodology's "never directly improves the objective score" line — a qualifier doing too much work — into two separate, true claims: (1) billing/claim state is never a scoring input at all (compute_current_score/public_display_score take only evidence, status, and tool count — no billing parameter exists anywhere in that call chain), verified by a rank-independence test; (2) authenticated validation, a paid feature, produces stronger evidence that can move a score, which is a real and disclosed evidence-path effect, not billing state itself affecting the score.
verify/tests/test_rank_independence.py proves (1) end-to-end: two identical servers differing only in claim/billing state produce byte-identical scores and identical search-ranking order, even when the lower-scoring server is the claimed/paid one (the scenario that would expose rank-buying if it existed).
For (2)'s actual number: built scripts/measure_evidence_path_effect.py (paired authenticated-vs-unauthenticated score delta per server) — its arithmetic is unit-tested against synthetic paired data, but not run against production (no database access this session). The methodology page states the mechanism and the disclosure commitment; it does not publish a number that hasn't actually been measured.
R11 — Score validation (scaffolding only)
Cannot be completed by an autonomous coding pass — it requires a real 150-300 server stratified sample rated by two independent human reviewers over real time, and fabricating that data would be exactly the kind of appearance-of-rigor-without-substance this whole plan exists to eliminate. What's built: verify/eval/ground-truth/ (rubric, disagreement process, the exact JSON-Lines schema a labeled sample must follow) and verify/eval/reports/ (report schema, scripts/generate_eval_report.py — Spearman rank correlation, inter-rater agreement via median-of-raters aggregation, discrimination, per-subscore correlation with an explicit cut threshold — and scripts/check_eval_report_freshness.py, a 90-day staleness gate). The freshness gate is not wired into the blocking pytest suite yet — doing that before a report exists would just permanently fail CI over a known, already-documented gap. Wire it in once the first real report lands.
verify/tests/test_eval_report_generation.py unit-tests the Spearman implementation and the reviewer-aggregation logic against synthetic data, so whoever runs the real study can trust the arithmetic on day one.
Verification
PYTHONPATH=verify/src pytest verify/tests # 283 passed
python3 -m py_compile verify/src/mcp_verify/main.py verify/src/mcp_verify/db/models.py \
verify/src/mcp_verify/db/repository.py verify/src/mcp_verify/validation/service.py \
verify/src/mcp_verify/materialization.py verify/src/mcp_verify/insights.py \
verify/src/mcp_verify/api/schemas.py
End-to-end check against the live canonical fixture (seeded into the test DB from the frozen JSON, then fetched through the real HTTP surfaces): no "Top 1.0%", no "Healthcare claims workflows", no "admin mutation", no "secret material access", a null badge score, and a visible "Dispute this assessment" section — all simultaneously true on one render of the one page this whole round was scoped against.
Not done this round (explicit scope cuts, per the plan's own priority discipline)
- R6's actual corpus execution and R11's actual ground-truth study — both need resources (production DB access, human reviewer time) this session didn't have. Both have real, tested mechanisms ready; neither has been run.
- R12-R15 (P1/P2: homepage funnel relabel, navigation cap, validation backoff, deploy consistency) — untouched, per "do not start P1 until every P0 is complete." P0 (R1-R10) is complete; R11 is scaffolded but not populated, which the plan's own sequencing note anticipated ("R11 has the longest lead time... begin recruiting reviewers... while P0 is in progress, even though its results land later").
- A full sitewide conclusory-copy audit (R9's larger ask) and consolidating the now-three independent taxonomy-rule copies into one shared function (R7) — both flagged as scope cuts in their own sections above.