MCP Verify — Round 14 Implementation Report (post-1.0.514 verification)
Implements the full "MCP Verify — Requirements, Round 8" doc, found against three fixtures at v1.0.514: awesome-hemmabo-se/hemmabo-mcp-server (healthy, 13 tools, real write/delete surface), ai.mitosislabs/mitosis (owner-opted-out via robots.txt), anananaTERRA/mcp-indesign-agent (never validated, metadata only). Preceded by a planning-only pass (see verify_round14_architecture_plan in memory) that diagnosed every P0/P1 item's root cause against current code before any implementation started. All items shipped (R41, R42, R43, R44, R27, R28-confirmed, R36, R45, R46- investigated). 388 tests passing, up from 377 at the start of this round.
Standing order (§0): classifier freeze
No new capability-classifier rules landed this round.
A unifying fix, found while investigating R41: one stale field, three consumers
Round 12's R33 built refresh_action_safety_probe_summary() specifically because action_safety_probe.details.summary is computed once at validation time and persisted verbatim, while every other "current" derived field recomputes fresh per request. R33 fixed the probe's .summary sub-field. It never extended to .status — confirmed still read directly from checks["action_safety_probe"] in three independent places (build_write_action_governance, the unsafe_for_write_actions client verdict, the confirmation benchmark's status and evidence citation).
Fixed by computing one genuinely fresh action_safety_probe per request (build_action_safety_probe(tool_inventory=tool_security_inventory, checks=deserialize_checks(latest_checks)), in build_server_insights) and threading it into all three consumers via a new fresh_action_safety_probe parameter, defaulting to None so historical-run rendering (build_validation_timeline, which deliberately shows honest per-run history, not today's re-scored opinion of a past run) is unaffected. Added deserialize_checks() (validation/service.py) — a production-grade version of a helper the test suite already had privately — to reconstruct real CheckResult objects from a ValidationRun.checks JSON dict.
R41 — One write-publishing gate
Confirmed three genuinely different threshold computations answering "is this server's write surface safe to publish": write_action_governance.safe_to_publish only treated action_status == "error" as disqualifying; the unsafe_for_write_actions client verdict had the same gap independently; client_remediation_modes.write_safe separately treated "warning" as blocking too. Confirmed live on hemmabo (post-deploy, fresh validation): safe_to_publish: true while write_safe mode showed blocked for the same "warning" status.
R41.1 — safe_to_publish now treats "warning" as disqualifying too (action_status not in {"error", "warning"}), matching write_safe mode's existing, stricter posture — a deliberate tightening, not a bug fix. The unsafe_for_write_actions verdict now reads write_action_governance["safe_to_publish"] directly instead of re-deriving its own threshold, so it can no longer disagree with either of the other two surfaces.
R41.2 — labels ("Write-action publishing" / "Write-safe publishing") left as two distinct API fields (removing/renaming a public JSON field is a bigger compatibility question than this round's scope), but they now always agree on the same underlying computation, which was the actual acceptance criterion.
R41.3 — confirmed mitosis's "blast radius Medium with 0 high-risk tools" is not a bug once the freshness fix lands — it's round 13's R37.1 declared_non_read_only_tools rule correctly firing. No code change needed beyond the freshness fix itself.
R41.4 — build_write_action_governance now returns declared_non_read_only_tool_names; client_remediation_modes's confirmation-blocker line names the actual tool(s) instead of a generic "risky actions" sentence when this specific reason fires.
R42 — Complete the opt-out publication policy
R17 (round 9) built the mechanism; the policy was only ever enforced on current_score/production_readiness. Confirmed live on mitosis: raw scores leaked through validation_timeline, incident_feed, recommended_for; a full risk assessment leaked through client_readiness_verdicts, tool_security_inventory, publishability_policy_profiles; every trust badge (Freshly Validated, OAuth Verified, No Critical Risk, etc.) was active with no owner_opted_out awareness at all.
R42.1/R42.2 — new redact_opted_out_insights() (insights.py): applied once, at the same choke point current_score is already overridden from (_get_cached_server_detail_response, main.py), over the fully-assembled insights dict (including hosted_runtime/ agent_commerce/disputes/the merged incident feed added after build_server_insights returns). Reduces every field to an empty value of its own type except production_readiness itself, which already correctly explains why everything else is empty. current_score_components (a base-model field, not part of insights) is separately nulled at the same call site.
R42.3 — build_trust_badge_states now returns all badges inactive, unconditionally, with a shared "opted out" reason, whenever production_readiness.code == "not_assessed_by_request".
R43 — Provenance divergence absorbs alias-consolidation disagreements
Ran the standing "unknown reads as pass" grep first (round 6's original ask, never previously completed) — five candidate instances found and recorded below as carried-forward findings, not fixed this round (out of this round's authorized scope).
Confirmed provenance_divergence_probe (registry-metadata-vs-server-card, one server) and build_alias_consolidation (cross-alias remote_url/ homepage/registry_source, request-time, always fresh) are two genuinely different, both-correct computations that nothing merged — a reader saw "Provenance Divergence: OK" directly above a real cross-alias endpoint split. Also confirmed the specific "ok" observed live was itself partly stale-data noise (hemmabo's fixture predated the round-13 deploy); on fresh validation data the probe correctly reads not_assessed on its own — the underlying merge gap is the real, still-live bug.
R43.1 — build_provenance_divergence_detail now accepts alias_consolidation and merges its source_disagreements into drift_fields, escalating status (to warning, or error for a remote_url split specifically) when the probe alone read ok/ not_assessed/missing. Deliberately a display-layer merge only — score_provenance_divergence (the persisted composite score) is unchanged this round; depressing it from alias data would need aliases threaded into the scoring pipeline itself, a bigger architectural change flagged as follow-up work rather than attempted under this round's timeline.
R43.2 — a remote_url disagreement now produces a named alias_endpoint_split finding in build_remediations, severity high, with its own playbook, independent of the generic score-decomposition-driven findings.
R43.3 — build_alias_graph now uses public_display_score() for both the canonical and alias rows instead of the raw Server.current_score column — an alias that was never independently validated now shows no score (None) instead of a misleading 0 next to the canonical server's real score.
R44 — Suppress install snippets for unvalidated endpoints
build_install_snippets formatted server.remote_url into four client-specific snippets unconditionally — confirmed live on the indesign fixture (a GitHub repo URL as remote_url, never validated). Now checks the latest initialize status; returns a single explanatory unavailable_reason (in place of all four snippets) when latest_validation is None or initialize didn't succeed (auth_required still counts as success, matching is_core_success_from_check_results elsewhere in this codebase).
R27 — Duplicate runs: the second, enqueue-time race
Round 13 shipped FOR UPDATE SKIP LOCKED, which stops two workers from claiming the same pending job row. New evidence — three runs in ~1-2 seconds, the third with a genuinely different (Failing) outcome — isn't fully explained by that fix: it requires three separate rows to have been enqueued for the same server, which SKIP LOCKED doesn't prevent.
Root cause: enqueue_due_validation_jobs's "does a pending job already exist for server X" check (has_pending_job/active_validation_servers) is a read followed by a separate write, not atomic — three verify-worker replicas (VERIFY_WORKER_SCALE=3), each running their own schedule_tick() independently, can each see "no pending job" and each insert one.
Fixed with a Postgres partial unique index (migration 0023, jobs(job_type, payload->>'server') WHERE job_type='validate_server' AND status IN ('pending','running')) and a new JobRepository.enqueue_validate_server_job_if_absent() that catches the resulting conflict via a SAVEPOINT and returns None instead of erroring — workers/scheduler.py's enqueue_from only counts a job as enqueued when it gets a real Job back. A genuine ordering subtlety was caught by this round's own test-writing, not assumed correct: begin_nested() must be called before session.add(job), not after — calling it after folds this job's own possibly-conflicting insert into begin_nested()'s internal snapshot-flush, raising the IntegrityError outside the intended try/except. Verified with a test that also proves unrelated, still-uncommitted work added earlier in the same session/batch survives a conflict (an incorrect first attempt that called a full session.rollback() on catch would have discarded it too — schedule_tick enqueues several job types into one session before a single commit). SQLite (the test suite) gets no index at all — matching the existing dialect-branch precedent in claim_due_jobs/ prune_duplicate_pending_jobs — verified against a plain (non-partial) SQLite index standing in for the real Postgres constraint, since the exact partial-index SQL syntax can't be exercised without a live Postgres server.
Deploy incident, self-caught and fixed same-session: the migration's first revision id (0023_validate_server_job_unique_active, 39 chars) exceeded alembic_version.version_num's VARCHAR(32) column — untestable locally (this environment has no alembic installed, only inside the verify-migrate container) and only surfaced on the actual IONOS deploy, which took the whole sentinel-signal.service stack down for ~15 minutes (verify-web/verify-worker both depend on verify-migrate completing). The failed transaction rolled back cleanly (confirmed via direct DB inspection — no partial state); fixed by shortening the revision id to 0023_validate_server_unique (27 chars), committed as a follow-up (33b05a5), redeployed, verified clean across all 6 smoke-checked routes. See verify_ionos_deploy_mechanics in memory for the durable lesson (revision ids must stay ≤32 chars; every prior one fit under that by coincidence, not enforced convention).
R28 — Confirmed still correct; new evidence is corpus confirmation, not a regression
Round 13's compute_summary_status(checks, tool_count=...) → "unknown" fix needed no changes. The new evidence (hemmabo 7/10, mitosis 5/7 historical zero-tool-Healthy runs) sharpens, but doesn't change, the already-deferred historical-backfill question — still not run without explicit confirmation.
R36 (reopened) — Tool list truncation
visible_tool_names = tool_names[:8] had no truncation indicator; closed in an earlier round only because that round's fixture happened to have exactly 8 tools. Now appends a "Showing 8 of N tools." line when the list is actually truncated.
R45 — Incident feed ordering
build_server_incident_feed already sorts its own output descending; the bug was at the merge point (main.py), which unconditionally prepended persisted correction/dispute notes ahead of the validation-derived feed without re-sorting the combined list. Fixed by merging then sorting, with the same [:8] cap the underlying function's own output already used.
R46 — Investigated, root cause confirmed, no fix designed (as specified)
Hemmabo's one ServerIncidentNote from round 6's backfill script shows an increase ("Prior: 75.98. Corrected: 77.01."), while mitosis and moonlings both correctly decreased from the same backfill run. Read the backfill script (scripts/backfill_score_and_capability_corrections.py) in full: confirmed the mechanism is architecturally capable of producing an increase, and this is not a bug in the script relative to what it actually promises. _set_or_exclude() recomputes each of the 16 R1_AFFECTED_KEYS components fully fresh using whatever score_* code is live at the moment the script runs — not a narrowly-scoped "only ever zero-anchor, never raise" patch. If any of those 16 functions had gained any other improvement between the server's original validation and the backfill run (not just the R1 zero-anchoring fix specifically), a net increase for a specific server's evidence shape is architecturally possible. Cannot be pinned to a specific component for hemmabo: ServerIncidentNote only persists the two aggregate totals (prior_value/ corrected_value as plain text), not a per-component diff — that data no longer exists anywhere. Recorded as a process gap for any future backfill script (persist the per-component diff, not just the aggregate), not fixed this round, per the plan's own "investigate first" instruction.
Standing "unknown reads as pass" grep (R43's own instruction)
Round 6 first asked for a full codebase sweep for this bug class; never previously done. Run this round via a dedicated research pass. Five candidates found, none fixed (out of this round's scope — the grep itself was the deliverable R43 asked for):
validation/service.py:4665—score_tool_snapshot_churnhardcodesinsights.py:912—build_compatibility_fixtures'srequest_associationinsights.py:1244— thesnapshot_churn_riskverdict falls through anvalidation/service.py:4151/~4213—prompts_list/resources_listvalidation/service.py:4405—score_installabilitycredits a
7.0 when there's no prior snapshot to diff, without checking whether the sibling probe already marked this "missing".
check treats status == "missing" as "passes".
unrun comparison to "low" risk with a "no churn detected" reason.
with status "skipped" (set when initialize is unreachable) isn't matched by the explicit "missing" zero-guard, falling to generic partial credit.
missing/None determinism_probe with 7.0; lower confidence since the live pipeline always populates this key.
Test suite
388 tests passing (PYTHONPATH=verify/src pytest verify/tests), up from 377 at the start of this round.
Deferred / carried forward
The five "unknown reads as pass" grep findings above; score_provenance_divergence's composite-score integration with alias data (R43.1's display-only scope limitation); R42.2's redaction policy (implemented as specified, per the doc's own explicit callout that this is Codex's job to implement, not decide — worth one confirmation given how sweeping the reduction is); R28's historical backfill; R46's process-improvement recommendation (persist per-component diffs in future backfill scripts).