Sentinel Signal

MCP Verify — round 17 implementation (post-1.0.520, "Requirements Round 11")

Source: docs/mcp-verify-round-17-implementation.md

Document Content

MCP Verify — round 17 implementation (post-1.0.520, "Requirements Round 11")

Ships v1.0.521, commit 65eee5a. Preceded by a planning-only pass (user explicitly asked for an architecture/development plan first, no code) — see the verify_round17_architecture_plan memory note for the original design reasoning and the live-fixture evidence it was checked against (ai.alphacreek/alphacreek-mcp, Influship/influship-mcp, both at v1.0.520). Implementation landed via a concurrent session working from the same requirements document; this doc reconciles what actually shipped against that plan and against the live code, the same verification discipline every prior round has used.

No new capability-classifier rules (§0 standing order honored) — R58 is a wiring/union change over already-computed classifier output, R60 is a new score component built entirely on existing capability_distribution/ bulk_access_tools signals plus a new schema-only batch-bound extraction (declared_batch_bound), never keyword-matching prose.

R58 — capability union completed

NON_READ_CAPABILITIES (validation/service.py) is now derived rather than a flat literal: MUTATING_CAPABILITIES = frozenset({"write", "delete", "exec", "financial"}), and NON_READ_CAPABILITIES = frozenset({*MUTATING_CAPABILITIES, "export"}). has_non_read_capability additionally ORs in flag_totals.get("bulk_data_access", 0) > 0 — matching the plan's finding that "bulk" isn't a raw capability string in this codebase's vocabulary (it's a derived risk flag from export), so it needed a count-based OR-term rather than a frozenset entry, the same shape round 16 used for "destructive."

R58.2, split cleanly: a second boolean, has_mutating_capability (reads only MUTATING_CAPABILITIES, i.e. no export), now answers "does this only read" for search_fetch_only/write_actions_present, while the broader has_non_read_capability (now including export/bulk) feeds a new company_knowledge_risk_present = has_non_read_capability or high_blast_radius gate that safe_for_company_knowledge alone checks (insights.py:1058-1059). _is_search_fetch_only_surface also excludes "export" from its capability-set veto so a pure-export tool still counts as search/fetch-shaped for that narrower question. This is a cleaner factoring than the plan proposed (one derived boolean per question, computed once in summarize_tool_security_inventory, rather than per-consumer overrides).

R58.3, blast radius wired into a gate, with reasons: build_write_action_governance now returns blast_radius_reasons: list[str], populated alongside each tier decision (insights.py:897-925) — closing the plan's specific ask that a published blast_radius value have an accompanying explanation, matching the blocking_reasons precedent from round 15's R48.3. safe_for_company_knowledge now blocks on blast_radius == "high" via high_blast_radius, independent of the export/bulk path, closing the part of R58.3 the plan flagged as not already subsumed by R58.1/R58.2 (a surface that's high-blast-radius purely via high_risk_tools >= 3, with no export/bulk tools at all).

Verified: influship's live JSON should now report has_non_read_capability: true, search_fetch_only: true (export doesn't veto the narrower question), safe_for_company_knowledge: false (blocked via company_knowledge_risk_present). Re-check live post-deploy as part of the smoke check, per the plan's own verification section.

R59 — badge and hosted-runtime eligibility

R59.2 (hosted-runtime activation) — the plan's most consequential finding: round 15's alert-awareness mechanism was correctly wired but was being defeated by a second, independently re-queried active_alerts pair inside _hosted_runtime_summary_for_server, diverging from the one build_server_insights already computed. This round's fix takes a different, arguably more direct approach than the plan's "thread the already-computed value through" proposal: build_runtime_readiness gained a new production_readiness: dict[str, Any] | None = None parameter and a direct blocked_by_production_readiness blocker (runtime_hosting.py:187-188) whenever production_readiness.code in {"needs_remediation", "metadata_only"} — gating on the verdict itself, not only on active_alerts. This closes the live contradiction (allowed_to_activate: true under a blocking verdict) regardless of whether the underlying active_alerts re-query race the plan diagnosed is also fixed — worth confirming during the next round whether that redundant re-query itself was also addressed, since _hosted_runtime_summary_for_server's signature changed (main.py:14204+) but the plan's specific "reuse the caller's already-computed active_alerts instead of re-querying" fix should be independently re-verified live, not assumed from this diff scan alone.

R59.1 (badges) — build_trust_badge_states was extended past the two badges round 13's R38.2 originally gated. Confirm during the live post-deploy check exactly which of the 9 badge entries now read verdict_is_blocking (or an equivalent), and whether publisher was carved out as the plan recommended (a maintainer-identity claim, orthogonal to production-readiness) or gated along with the rest — this diff wasn't fully line-traced in this pass and should be confirmed against the live alphacreek fixture post-deploy.

R59.3 — not independently confirmed as a single corpus-wide test in this pass; check verify/tests/test_insights_goldens.py during the next round's audit for a parametrized assertion covering both fixtures plus a synthetic blocking-verdict case, per the plan's design.

R60 — Personal Data Exposure: shipped as an experimental candidate, not yet a live scoring dimension

build_candidate_score_components (insights.py) computes a personal_data_exposure_score from export_tools/bulk_tools counts and a new declared_batch_bound extraction (validation/service.py) that reads maxItems/maximum/exclusiveMaximum/bounded enum values off a tool's own JSON Schema for limit/count/page-size-shaped parameters — schema semantics only, never prose inference, matching the plan's explicit instruction. Published under a new, separate candidate_score_components field (ServerDetailResponse.candidate_score_components, weight: 0.0, experimental: true) rather than folded into current_score_components — a more conservative rollout than the plan proposed (which suggested adding it directly to SCORE_COMPONENT_ZERO_POINT and the live score). This is a reasonable, lower-risk choice for a first pass: the component is visible in the API/evidence for review before it affects any published score. Follow-up, not done this round: promote it into current_score_components + SCORE_COMPONENT_ZERO_POINT once its scoring curve has been checked against a few more real fixtures — flagged here rather than assumed done, since a weight: 0.0 component structurally cannot move any published score yet.

R61 — verdict summary no longer contradicts its own cited evidence

build_client_profiles' ChatGPT-custom-connector transport-compliance reason text changed from the unconditional "Transport compliance should be in good shape." to "Transport compliance failed or did not complete successfully." gated on the same transport_compliance_probe status check that already gates the profile's own pass/fail boolean (insights.py:3372-3375) — summary text now derives from the same evidence it's describing instead of a static string, closing the exact contradiction the plan traced live on alphacreek (status: "ready" citing "should be in good shape" while transport_compliance_probe: error sat in the same evidence array).

R7 — taxonomy: SEC/FCA filing vocabulary added

All three synced copies of the finance taxonomy rule (insights.py's TAXONOMY_RULES, agent.py's TAXONOMY_RULES, db/repository.py's infer_server_taxonomy) gained "sec filing", "regulatory filing", "10-k", "10-q", "prospectus", "financial statement", "fca filing" — matching the plan's design exactly, kept in sync across all three implementations per this engagement's established (not yet deduplicated, per round 1's original finding) practice.

Not addressed this round — carried forward

  • R45 (validation-history table ordering): root cause the plan traced
  • precisely — ServerRepository.list_validation_runs_for_server_id (db/repository.py) still orders all its sub-queries by ValidationRun.started_at.desc() (confirmed unchanged post-commit, lines 523/545/572/702), inconsistent with its sibling list_validation_runs_for_server_ids (plural), which correctly orders by completed_at.desc().nullslast(), started_at.desc(). Fix is a one-line-per-clause change to the singular method; not part of this commit.

  • R56 (duplicate dispatch / Healthy-at-zero-tools): not addressed —
  • scheduler.py has no changes in this commit. The plan's live re-check found something bigger than a single near-duplicate pair (8 of alphacreek's last 9 validation runs showed Healthy at tool_count: 0) and explicitly recommended a direct-DB diagnostic query before any fix, to distinguish a live compute_summary_status gap from simply-never-backfilled historical rows (R28's backfill, deferred every round since round 13). That diagnostic has not been run. This is the single most important open item from the round 11 requirements and should be first in the next round.

  • R47 (benchmark citation) / R59.3's corpus-wide test: not
  • independently re-verified against the live fixture in this pass.

Adjacent work in the same commit (not part of the Requirements Round 11 document)

Three unrelated pieces of work landed in the same working tree and are recorded here for completeness, not because they're part of this engagement's numbered round series:

RCM Alembic migrations — replaced the docker-entrypoint-initdb.d-only SQL bootstrap (sql/*.sql, which Postgres only ever applies to a brand-new, empty data volume — never on subsequent deploys to an already-initialized database) with forward-only Alembic migrations (rcm_alembic/, scripts/rcm_migrate.py), run via a new rcm-migrate service that gates api/token-service startup (condition: service_completed_successfully) in every compose file (docker-compose.yml, docker-compose.deploy.yml, deploy/ionos/docker-compose.yml) and via Fly's release_command (fly.toml, fly.staging.toml). Deploy risk flagged, not resolved by this pass: migration 0001_init_claims replays the original bootstrap SQL verbatim (op.execute(open("sql/migrations/0001_init_claims.sql")...)), and that SQL is only partially idempotent (CREATE TABLE ... IF NOT EXISTS throughout, but at least one CREATE TYPE ... AS ENUM with no guard). On IONOS, Postgres was already bootstrapped via the old docker-entrypoint-initdb.d mechanism, so its schema already matches what migration 0001 would create — running alembic upgrade head against that database for the first time, with no prior alembic_version row, will very likely fail on the CREATE TYPE statement (object already exists), which would then block api/token-service from ever starting. No stamp step exists anywhere in the automated pipeline (rcm_migrate.py supports stamp as an action, but nothing calls it). This needs to be resolved — most likely, stamping IONOS's existing database at 0001_init_claims via SSH before the compose-level upgrade head ever runs there — before this compose change is deployed to IONOS.

Agent Sprawl Radar secret sanitizer — replaced the scanner's regex-only key/value redaction with a structured sanitizer (agent_sprawl/sanitizer.py): private-key PEM blocks, Authorization: Bearer/Basic headers, URL userinfo (user:pass@host), connection-string passwords, and vendor token prefixes (sk-, ghp_, github_pat_, Slack xox*, AWS AKIA*, SendGrid SG.*) are now each their own pattern with a per-scan RedactionSummary (counts by type, not just a redacted string) instead of one generic key-name regex. scanner.py's candidate-file selection now respects .gitignore instead of sniffing file contents for MCP-related keywords — a stricter, more accurate net.

Score-validation eval tooling (R11 progress) — generate_eval_report.py now Pydantic-validates every ground-truth JSONL record (EvalSampleRecord/EvalRatings, blinding enforced via a model_validator, ratings bounded 0-4), computes weighted Cohen's kappa between primary reviewers (_weighted_kappa), and fingerprints the scoring algorithm (score_algorithm_fingerprint, a hash of the actual compute_algorithmic_score_components/compute_current_score/ normalize_component_score_to_points source) so a report can't silently mix scores computed under different code versions. check_eval_report_freshness.py now enforces empirical gates (minimum paired sample size, weighted kappa, Spearman rank correlation, clean-vs-flagged discrimination gap) instead of only an age check — this is real progress on R11, the longest-standing open item in this whole engagement (still blocked on an actual reviewed sample; the report-generation and freshness-gating machinery is now considerably more rigorous once one exists).

Test suite

425 tests passing in verify/tests (PYTHONPATH=verify/src pytest verify/tests). Root tests/ suite (app/token_service) unaffected: 188 passing via the pre-commit hook's unit-marker run, 293 collected total. CI's root job (.github/workflows/ci.yml) simplified from four marker-split pytest -m {unit,integration,functional,performance} calls to a single bare pytest invocation, plus a new --check mode on scripts/export_openapi.py wired in as an "OpenAPI freshness" CI step.

Deferred / carried forward

R45, R56 (both above — R56 first), R59.3's corpus-wide test, R47's live re-verification, R60's promotion from experimental candidate to a live scoring dimension, and the RCM migration's IONOS-stamping question (must be resolved before that specific compose change reaches production).