Sentinel Signal

Handoff: V1/V2 remediation-list bugs + compatibility fail-open defect

Source: docs/codex-handoff-remediation-and-compatibility-fixes.md

Document Content

Handoff: V1/V2 remediation-list bugs + compatibility fail-open defect

Status: NOT YET IMPLEMENTED. This document is a full handoff — a prior Claude Code session investigated and designed this fix set but ran out of budget before implementing. Everything below is ready to execute directly; no further design decisions should be needed except where explicitly flagged.

Repo / environment context (read this first)

  • Repo: claim-risk-score (this repo). The relevant subproject is verify/ — a FastAPI app at verify/src/mcp_verify/, deployed as verify.sentinelsignal.io.
  • Main files for this work: verify/src/mcp_verify/insights.py (remediation/alert/compatibility-profile building logic) and verify/src/mcp_verify/main.py (rendering, policy export, build_trust_snapshot).
  • Tests: verify/tests/ (pytest). Run with:
  PYTHONPATH=verify/src:verify/tests python -m pytest verify/tests -q

Root repo also has unrelated unit tests (python -m pytest -m unit -q from repo root) — not required for this change but harmless to run.

  • Git: a pre-commit hook auto-bumps VERSION across ~11 files via scripts/bump_version.py and runs pytest -m unit — this happens automatically on git commit, no manual action needed. Always create NEW commits, never amend, unless the user explicitly says otherwise.
  • Deploy mechanics (only do this if the user asks you to deploy, not automatically as part of "implement the plan"):
  • Production host: root@66.179.248.190, app lives at /opt/sentinel-signal (this is not a git checkout there — it's populated by rsync).
  • Deploy sequence: git commit → git push origin master → rsync -az --delete --exclude '.git' --exclude '.venv' --exclude 'node_modules' --exclude '__pycache__' --exclude '*.pyc' --exclude '.pytest_cache' --exclude 'dist' --exclude 'artifacts' --exclude '.DS_Store' ./ root@66.179.248.190:/opt/sentinel-signal/ → then on the remote: RELEASE_SHA=<local git rev-parse HEAD> RELEASE_VERSION=<local VERSION file contents> ./deploy/ionos/deploy.sh (run over SSH). **Compute RELEASE_SHA/RELEASE_VERSION locally and pass them as literal values** — the remote host has no .git dir, so $(git rev-parse HEAD) evaluated remotely fails.
  • The deploy restarts a systemd service (sentinel-signal.service, which itself runs docker compose ... up -d --build) — this takes over 60-90 seconds and always clears all in-memory route caches (there is no CDN in front of this app; a restart is the entire "cache invalidation" story).
  • After deploy, confirm success via curl https://verify.sentinelsignal.io/v1/build — check version, build_sha, and regression_checks_passing: true.
  • Two live servers are used as running examples throughout this doc and were the ones the original bug report was based on:
  • awesome-malinoto/tracepass-mcp-server — an assessed server (has real validation data).
  • ryanatindago/playwright-mcp-example — a never-successfully-contacted server (initialize never got a response; assessment_status is not_assessed).
  • Fetch either one's live policy/report via e.g. curl -s "https://verify.sentinelsignal.io/v1/servers/awesome-malinoto/tracepass-mcp-server/report".

Background: how this fix set was found

A user reviewed the live server-detail pages for the two servers above and reported 5 real defects (ranked most → least significant by them) plus 2 items that needed investigation to determine if they were real. Three research passes (by subagents) traced each to exact code. This document is the resulting, already-designed fix plan.

---

Fix 1 (highest priority): compatibility scoring silently treats unavailable evidence as a pass

File: verify/src/mcp_verify/insights.py, function build_client_profiles, the openai_connectors profile's criteria list, specifically the step_up_auth_probe criterion around line 4500-4504.

The bug: this one criterion hand-codes "missing" (meaning the check was never run / evidence unavailable) into its own passing set:

(
    str((latest_checks.get("step_up_auth_probe") or {}).get("status") or "missing") in {"ok", "warning", "missing"},
    "Step-up auth signals should be documented and connector-friendly.",
    "Document step-up auth requirements in a connector-friendly way.",
),

Every other criterion in this same file follows the opposite, correct rule — "unknown is never a pass" (a standing design principle used repeatedly elsewhere in this codebase: see main.py:18974-18985's policy-export evidence_unavailable_fail_closed blocker, the assessment_status tri-state, materialization.py:404-418's ranking gate). The transport_compliance_probe criterion three lines above this one, in the same profile, has a comment explaining exactly why "missing" must be excluded:

(
    # R18/standing constraint (round 9 requirements): "missing"
    # here can mean initialize already failed and this probe
    # was deliberately never attempted (see the R14 skip
    # logic in validation/service.py) -- not evidence
    # transport compliance is probably fine. Matches the
    # already-correct anthropic_remote_mcp profile below,
    # which never included "missing" in its passing set.
    str((latest_checks.get("transport_compliance_probe") or {}).get("status") or "missing") in {"ok", "warning"},
    "Transport compliance failed or did not complete successfully.",
    "Resolve the transport compliance failure -- see the transport compliance probe evidence for what specifically broke.",
),

Why this matters: the score computed from these criteria (_make_profile, insights.py:5061-5069, score = round((passed / len(criteria)) * 100, 1)) can render a full 100.0 / "N of N requirements met" verdict when a criterion's evidence was never actually gathered. This is only visible on tracepass today because a separately-computed "blocked" verdict (see build_client_readiness_verdicts, insights.py:1692-1812) happens to contradict the 100.0 score. On a server with nothing else blocking, this would produce a fully silent, unearned "100.0 compatible" result.

The fix: remove "missing" from the passing set, matching the sibling criterion above it exactly:

(
    str((latest_checks.get("step_up_auth_probe") or {}).get("status") or "missing") in {"ok", "warning"},
    "Step-up auth signals should be documented and connector-friendly.",
    "Document step-up auth requirements in a connector-friendly way.",
),

Important — this is a real scoring behavior change, flag it, don't just quietly ship it: any server whose step_up_auth_probe is currently "missing" will lose one passing criterion out of 9 for the OpenAI Connectors profile (100.0 → 88.9, or lower if it was already imperfect elsewhere). Most servers won't cross the 80% "compatible" threshold from this alone, but some marginal ones could. After deploying, spot-check a handful of live servers (tracepass plus 2-3 others whose step_up_auth_probe status is currently "missing") to confirm the new score/verdict combination looks coherent, not just that it changed.

Confirmed not a caching bug: build_client_profiles reads latest_validation.checks directly, as part of the full build_server_insights() computation pipeline — it never goes through the /policy route's fast/fallback builders (which a separate, already-fixed bug this session affected). This is a standalone logic bug in the scoring criteria themselves.

---

Fix 2: "sync tool schemas" remediation fires with zero actual divergences (V1)

File: verify/src/mcp_verify/insights.py, function build_remediations, the check-loop's handling of check_name == "schema_divergence_probe", around lines 4304-4370 (the check-loop starts at 4304; the schema_divergence_probe-specific branch is at 4355-4369; the row gets written into items[code] at 4370).

The bug: the loop's general absence-gate (line 4320) only skips a row when status in {"missing", "not_assessed"}:

if not fires_on_absence and status in {"missing", "not_assessed"}:
    continue

It never checks whether tool_divergences (the actual list of field-level mismatches) is non-empty. Meanwhile, build_active_alerts (a different function, same file, insights.py:3963-4013) already correctly distinguishes two cases for this exact probe:

# insights.py:3983-3990 (inside build_active_alerts)
severity_tier = schema_divergence_details.get("severity") or (
    "high" if schema_divergence_status == "error" else "medium"
)
if severity_tier == "low":
    alerts.append({
        "code": "card_omits_tool_schemas",
        "severity": "low",
        "title": "Server card does not publish tool schemas",
        "message": (
            f"{len(omitted_schema_tools)} tool(s) are listed on the published server card without an "
            f"inputSchema: ..."
            ". This is a completeness gap, not a contradiction -- clients must call tools/list "
            "for parameter details rather than trusting the card."
        ),
    })

severity_tier == "low" is exactly the "only an omission was found, tool_divergences is empty" case. build_remediations never checks this, so the same page snapshot can carry both the correct "completeness gap, not a contradiction" alert and a remediation row asserting "Sync the published server card's tool schemas with the live tool surface" / "The server card and live tools/list disagree on schema details" — a direct self-contradiction, confirmed live on tracepass with tool_divergences: [] in the same snapshot.

The fix: in the schema_divergence_probe branch (right after tool_divergences is extracted at line 4355, before the row is written to items[code] at line 4370), compute severity_tier the same way build_active_alerts does — don't re-derive the "is this a real contradiction" logic independently, reuse the identical formula so the two functions structurally cannot disagree again — and skip the row entirely when it's "low":

if check_name == "schema_divergence_probe":
    details = (payload or {}).get("details") or {}
    severity_tier = details.get("severity") or ("high" if status == "error" else "medium")
    if severity_tier == "low":
        continue
    why = _schema_divergence_remediation_why(tool_divergences)

(status at this point in the loop is already str((payload or {}).get("status") or "unknown") from line 4305 — same value build_active_alerts calls schema_divergence_status.)

Minor same-fix bonus — a typo: _schema_divergence_remediation_playbook (insights.py:3821-3830) builds its playbook text as:

dimension_text = ", ".join(dimension_labels) if dimension_labels else "the recorded schema details"
return [..., f"Fix the {dimension_text} mismatch(es) per tool."]

When dimension_labels is empty, the default string "the recorded schema details" already contains "the", so the composed sentence reads "Fix the the recorded schema details mismatch(es) per tool." (doubled "the"). This default is only reachable in exactly the same empty-divergence case the fix above now suppresses entirely, so after Fix 2 this text should never render in production — but fix the string anyway (defense in depth, in case another call path reaches it): change the default from "the recorded schema details" to "recorded schema details" (drop the leading "the", since the caller's f-string already supplies it).

---

Fix 3: publisher-facing remediation rows survive on a server that was never contacted (V2, first half)

File: verify/src/mcp_verify/insights.py, function build_remediations.

The bug, in three parts. server_unreachable (computed via _initialize_never_received_a_response, used at line 4298) currently only suppresses 3 checks:

# insights.py, near line 4191
NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS = {"initialize", "oauth_protected_resource", "server_card"}
# insights.py:4329-4330, inside the check-loop
if server_unreachable and check_name in NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS:
    continue
  1. tools_list has fires_on_absence=True by design in REMEDIATION_RULES (a deliberate "no news is itself worth flagging" case for the normal scenario) and is not in the suppression set above — so "Ensure tools/list succeeds consistently" fires unconditionally, even when the server was never reached at all (structurally, tools/list could never have even been attempted if initialize never got a response).
  2. prompts_list and determinism_probe have fires_on_absence=False and, when initialize never responds, get marked status="skipped" by service.py's _skip_check helper (service.py:73-79, called at e.g. service.py:809, 824, 868). But the general absence-gate at line 4320 only tests status in {"missing", "not_assessed"} — never "skipped" — so these rows slip through the exact gate meant to catch them. This produces "Stabilize repeated tools/list responses" (from determinism_probe) and "Repair prompts/list or stop advertising prompts" (from prompts_list, confirmed live on playwright — which never advertised prompts at all).
  3. The 5 "Raise X score" rows come from a separate loop, not the check-name loop above, around lines 4386-4403:
   for category in score_decomposition:
       if category["score"] <= max(4.0, category["max_score"] * 0.35):
           code = f"improve_{category['key']}"
           items.setdefault(code, {...})

This loop has no server_unreachable check at all — there's no scoring signal worth acting on for a server that was never reached, but the rows fire anyway.

The fix, three parts:

3a. Near NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS (~insights.py:4191), add a second, broader set. Do not just add items to the existing set — it's also used for addressee-routing (insights.py:4343-4345, deciding whether a row reads as "publisher" or "registry" addressed when non_mcp_endpoint is true, a related-but-distinct question from "was the server ever reached at all"). Keep that routing check on the original, narrower set, unchanged:

# A superset of NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS: every check
# that structurally cannot have run if initialize never got a response,
# not just the 3 checks that also need registry-addressed reframing.
SERVER_UNREACHABLE_SUPPRESSED_CHECKS = NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS | {
    "tools_list", "prompts_list", "resources_list", "determinism_probe",
}

Change the suppression check at line 4329 to use this broader set:

if server_unreachable and check_name in SERVER_UNREACHABLE_SUPPRESSED_CHECKS:
    continue

(Leave line 4343-4345's addressee-routing check — addressee = "registry" if non_mcp_endpoint and check_name in NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS else "publisher" — exactly as it is, referencing the original set.)

3b. Also add "skipped" to the general absence-gate at line 4320 (defense in depth — a check can be genuinely "missing"/"skipped" for reasons other than total server unreachability, and a fires_on_absence=False check should never fire on a probe that was never actually attempted, regardless of cause):

if not fires_on_absence and status in {"missing", "not_assessed", "skipped"}:
    continue

3c. Wrap the score-category loop (insights.py:4386-4403) so it's skipped entirely when server_unreachable:

if not server_unreachable:
    for category in score_decomposition:
        if category["score"] <= max(4.0, category["max_score"] * 0.35):
            code = f"improve_{category['key']}"
            items.setdefault(code, { ... })  # existing body unchanged

---

Fix 4: declared alert-count badge disagrees with the actual alert list (V2, second half)

Files: verify/src/mcp_verify/main.py (the badge) and verify/src/mcp_verify/insights.py (the severity table).

The bug: render_production_readiness (main.py:28410-28463) computes a displayed "N high/critical alerts" badge as:

# main.py:28447-28452
remediation_critical_count = sum(
    1 for item in (remediations or [])
    if str(item.get("severity") or "") in {"critical", "high"} and item.get("addressee") in (None, "publisher")
)
displayed_critical_count = max(int(readiness.get("critical_alerts") or 0), remediation_critical_count)

readiness["critical_alerts"] is correctly derived from the real active_alerts list. remediation_critical_count is a second, independent count built from remediations — and each remediation's severity only gets reconciled against what's actually in active_alerts for 2 check types, via:

# insights.py:4205-4208
CHECK_TO_ALERT_CODES = {
    "schema_divergence_probe": ("server_card_schema_drifted", "card_omits_tool_schemas"),
    "provenance_divergence_probe": ("official_registry_entry_drifted",),
}

...consumed by resolve_remediation_severities (insights.py:4211-4244). Every other check type keeps its static REMEDIATION_RULES severity (e.g. initialize → "critical", tools_list → "critical", server_card/oauth_* → "high") regardless of what actually survived build_active_alerts' own suppression/dedup logic into the final active_alerts list. Confirmed live: tracepass declares "1 high/critical alert" while its one actual active alert is severity "low"; playwright declares "4" while its one actual active alert is severity "medium".

**User's explicitly chosen fix direction (conservative — keep the existing max()-union design, just make the reconciliation complete)**: extend CHECK_TO_ALERT_CODES to cover every check type in REMEDIATION_RULES (insights.py:139+) that can produce a corresponding active_alerts entry, mapping each check name to the actual alert code(s) build_active_alerts emits for it. This requires reading both tables carefully — REMEDIATION_RULES (starting insights.py:139) for the full list of check names and their static severities, and build_active_alerts (search for alerts.append( inside it) for the actual code string(s) each check type can produce — and building an accurate name → code(s) mapping for each. Do this as a careful, complete pass, not a quick partial patch; an incomplete mapping just narrows the gap instead of closing it.

Once CHECK_TO_ALERT_CODES is complete, resolve_remediation_severities will cap every remediation's severity at what's real in active_alerts for every check type, so remediation_critical_count can never exceed what active_alerts itself supports — closing the gap without touching the max()-union structure in main.py at all.

---

Investigated and intentionally NOT included in this fix set

Two items from the original feedback were investigated and determined not to need a code change. Do not "fix" these — they were checked carefully and are either working as designed or unreproducible:

  1. Smithery shows "compatible" at score 80.0 with one blocker/missing-requirement present. Not a defect. _make_profile (insights.py:5061-5069) sets compatibility = "compatible" at score >= 80 — the same threshold used for every client profile in this file. "Compatible" has never meant zero missing requirements; Smithery's profile has 5 criteria, and 4-of-5 (80%) is exactly the threshold the design tolerates. Separately: Smithery is not in build_client_readiness_verdicts's broader blocker-reconciliation loop (only openai_connectors/claude_desktop are, insights.py:1720-1723), but Smithery has no corresponding publishability_policy_profile/OAuth-connector policy gate in this codebase for that broader reconciliation to meaningfully add — so the omission is inert, not a live gap.
  1. **evidence_confidence.score reads null while freshness.confidence_score reads 50.0 in the same tracepass snapshot.** Investigated live twice (~15 seconds apart) on tracepass's real (non-fallback) /policy response — current_snapshot.evidence_confidence.score and current_snapshot.freshness.confidence_score were consistently equal both times (both null in a fallback-served response, both a real non-null number once promoted to the real computed value). build_trust_snapshot (main.py:18266-18309) derives both fields from the identical payload.get("evidence_confidence") in a single pass, so they cannot structurally drift apart within the same current_snapshot. The split originally reported was very likely captured during the same /policy cold-cache fallback window that was already fixed elsewhere this session (the fallback path hand-rolls its own separate placeholder evidence_confidence dict, unrelated to the real computation) — this does not currently reproduce and no fix is proposed. If it's seen again, re-confirm live (fetch /policy twice, ~10-15s apart, checking for a partial: true/partial: false transition) before assuming a code bug.

---

Implementation order

Recommended order (independent fixes, but this order keeps related context together and de-risks the highest-priority item first):

  1. Fix 1 (compatibility scoring) — standalone, single file/function, most important.
  2. Fix 2 (V1 schema-sync suppression + typo) — standalone, single function.
  3. Fix 3 (V2 publisher-rows-on-unreachable-server) — standalone, single function, three sub-changes.
  4. Fix 4 (V2 alert-count reconciliation) — needs the most careful cross-referencing of two tables; do last so it isn't rushed.

Verification

  • After each fix: PYTHONPATH=verify/src:verify/tests python -m pytest verify/tests -q (must stay green — currently 770 passing before this work starts).
  • New tests to add (one or more per fix, in verify/tests/, following existing patterns in that directory — e.g. test_api.py or test_score_integrity.py):
  • Fix 1: a fixture with step_up_auth_probe status "missing" on an otherwise-perfect OpenAI Connectors profile — assert the score is no longer 100.0 and the criterion appears in missing_requirements.
  • Fix 2: a fixture with tool_divergences: [] and only an omitted-schema finding — assert the "sync tool schemas" remediation row is absent from remediations, while the existing low-severity "completeness gap" alert is still present in active_alerts.
  • Fix 3: a fixture matching playwright's shape (initialize never responded / server_unreachable true) — assert tools_list, prompts_list, determinism_probe, and all improve_* (score-category) rows are absent from remediations.
  • Fix 4: a fixture exercising the newly-covered part of CHECK_TO_ALERT_CODES — a check type outside the original 2 whose real active_alerts severity is lower than its static REMEDIATION_RULES severity — assert the displayed critical-alert count no longer exceeds what active_alerts supports.
  • python -m pytest -m unit -q from repo root (pre-commit will also run this automatically on git commit).
  • Live, after deploy (only if/when the user asks for a deploy):
  • curl -s "https://verify.sentinelsignal.io/v1/servers/awesome-malinoto/tracepass-mcp-server/report" — confirm the schema-sync remediation row is gone, the declared critical-alert count now matches its actual (low-severity) alert, and its OpenAI Connectors / Claude Desktop compatibility scores are no longer 100.0 from unavailable step_up_auth_probe evidence.
  • curl -s "https://verify.sentinelsignal.io/v1/servers/ryanatindago/playwright-mcp-example/report" — confirm the tools_list/prompts_list/determinism_probe/score-category publisher-addressed rows are gone from remediations.
  • Spot-check 2-3 other live servers whose step_up_auth_probe is currently "missing" to confirm Fix 1's score change looks coherent (not just "different").

Commit / push (only if the user asks for this — do not push automatically)

Standard flow: stage the changed files specifically (not git add -A), commit with a message explaining the "why" (this doc's context section has the substance), push to origin master. The pre-commit hook will auto-bump VERSION and run pytest -m unit — if that fails, fix the issue and create a new commit, don't amend.