Sentinel Signal

MCP Verify — Round 13 post-1.0.524 implementation

Source: docs/mcp-verify-round-13-post-1.0.524-implementation.md

Document Content

MCP Verify — Round 13 post-1.0.524 implementation

Initial target release: 1.0.529.

Implemented changes

  • R66: robots consent is now interpreted as a tri-state: true, false, or
  • "unknown". Non-200, timeout, parse failure, and fetch exceptions become consent unknown. Unknown and explicit opt-out stop downstream validation at the request layer with consent_unknown or owner_opt_out; unreadable robots.txt no longer penalizes Abuse / Noise Resilience.

  • R67: client-readiness, remediation, and publishability summaries now derive
  • from the same checklist/blocker evidence. Generic reassurance is emitted only when the profile is actually ready/compatible and has no blockers. Action safety evidence reads its own probe summary and reports write/export/egress counts instead of falling back to no-risk language.

  • R68: reachable but unenumerated tool surfaces are classified as unverified,
  • not Healthy. Unverified observations are excluded from healthy-ratio and evidence-confidence density calculations. A dry-run-first historical repair script reports affected rows before any mutation.

  • R60/R62: personal_data_exposure_score remains a zero-weight experimental
  • candidate until the empirical evaluation report/gate exists. It is rendered visibly in the candidate section and is suppressed from production component rendering for historical rows that persisted it during 1.0.526.

  • R69: added a taxonomy distribution reporting script that compares legacy
  • substring behavior with current word-boundary behavior and emits up to 50 spot-check candidates. No new capability-classifier rules were added.

Remaining blocked evidence gates

  • R62 empirical/flag-prevalence release gating remains blocked until a real
  • corpus report is generated under verify/eval/reports/.

  • R24.3 classifier precision work remains out of scope for this round; taxonomy
  • reporting is evidence-only.

  • Historical unverified repair remains dry-run-first and requires explicit
  • approval before applying mutations.

Validation

  • PYTHONPATH=verify/src:verify/tests pytest verify/tests -q
  • pytest
  • python3 scripts/validate_test_quality_gates.py --coverage-json artifacts/test-reports/coverage.json --min-line 70 --min-branch 55
  • python3 scripts/export_openapi.py --check
  • Compose config validation for local, deploy, and IONOS stacks

Feedback corrective release

Target release: 1.0.530.

This follow-up addresses the Strava / ai.com.mcp feedback without a database schema migration.

Alias identity policy

  • Alias consolidation now requires strong identity evidence: exact normalized
  • remote URL, exact normalized server-card URL, GitHub/GitLab repository slug from homepage/docs/support metadata, or an explicit registry cross-reference.

  • Shared provider namespaces, registry-source prefixes, provider homepages,
  • generic docs/support URLs, title matches, and host/title heuristics are not identity evidence.

  • Exact server-name matches are corroborating evidence only. They are not
  • sufficient to dedupe or render aliases.

  • Alias API and rendered alias surfaces include identity_strength: "strong"
  • plus identity_evidence values such as remote_url, server_card, repo_slug, or registry_cross_reference.

Client summary and write-safe wording

  • Readiness summaries now use observation text from checklist items.
  • why_not_ready continues to use remediation/action text.
  • Generic reassurance is limited to ready/compatible states with an empty
  • checklist.

  • Write-safe publishing and benchmark tasks consume the same action-safety and
  • write-governance result. Clean read-only evidence no longer produces confirmation/dry-run remediation or a failed safe-write benchmark.

  • Alert-only write-safe blockers lead with the active alert instead of implying
  • nonexistent risky write actions.

Taxonomy and history handling

  • Fitness/activity taxonomy markers cover Strava-style surfaces while retaining
  • word-boundary protections for prior substring false positives such as ChatGPT, database, and communication.

  • The taxonomy distribution report now includes before/after counts, mean tags
  • per server, percent other, top dropped tags, and deterministic spot-check candidates.

  • Recent healthy ratio and evidence-confidence history depth now require an
  • observed tool surface: confirmed fetched non-empty tools or an explicit empty tool list. unverified, unknown, auth-gated/unfetched, omitted lightweight payloads, and checks == {} do not count. If no recent observed runs exist, healthy ratio renders as n/a.

Personal-data score status

personal_data_exposure_score remains displayed/persisted where applicable as a zero-weight, non-gating candidate until empirical promotion gates pass. It does not affect totals, verdicts, percentiles, rankings, or release gates in this release.

R66 status

No functional R66 change is included in this feedback release. The Strava fixture returned readable robots.txt, so the existing tri-state consent behavior from 1.0.529 remains the intended implementation.