Handoff: V1/V2 remediation-list bugs + compatibility fail-open defect
Status: NOT YET IMPLEMENTED. This document is a full handoff — a prior Claude Code session investigated and designed this fix set but ran out of budget before implementing. Everything below is ready to execute directly; no further design decisions should be needed except where explicitly flagged.
Repo / environment context (read this first)
- Repo:
claim-risk-score(this repo). The relevant subproject isverify/— a FastAPI app atverify/src/mcp_verify/, deployed asverify.sentinelsignal.io. - Main files for this work:
verify/src/mcp_verify/insights.py(remediation/alert/compatibility-profile building logic) andverify/src/mcp_verify/main.py(rendering, policy export,build_trust_snapshot). - Tests:
verify/tests/(pytest). Run with:
PYTHONPATH=verify/src:verify/tests python -m pytest verify/tests -q
Root repo also has unrelated unit tests (python -m pytest -m unit -q from repo root) — not required for this change but harmless to run.
- Git: a pre-commit hook auto-bumps
VERSIONacross ~11 files viascripts/bump_version.pyand runspytest -m unit— this happens automatically ongit commit, no manual action needed. Always create NEW commits, never amend, unless the user explicitly says otherwise. - Deploy mechanics (only do this if the user asks you to deploy, not automatically as part of "implement the plan"):
- Production host:
root@66.179.248.190, app lives at/opt/sentinel-signal(this is not a git checkout there — it's populated by rsync). - Deploy sequence:
git commit→git push origin master→rsync -az --delete --exclude '.git' --exclude '.venv' --exclude 'node_modules' --exclude '__pycache__' --exclude '*.pyc' --exclude '.pytest_cache' --exclude 'dist' --exclude 'artifacts' --exclude '.DS_Store' ./ root@66.179.248.190:/opt/sentinel-signal/→ then on the remote:RELEASE_SHA=<local git rev-parse HEAD> RELEASE_VERSION=<local VERSION file contents> ./deploy/ionos/deploy.sh(run over SSH). **ComputeRELEASE_SHA/RELEASE_VERSIONlocally and pass them as literal values** — the remote host has no.gitdir, so$(git rev-parse HEAD)evaluated remotely fails. - The deploy restarts a systemd service (
sentinel-signal.service, which itself runsdocker compose ... up -d --build) — this takes over 60-90 seconds and always clears all in-memory route caches (there is no CDN in front of this app; a restart is the entire "cache invalidation" story). - After deploy, confirm success via
curl https://verify.sentinelsignal.io/v1/build— checkversion,build_sha, andregression_checks_passing: true. - Two live servers are used as running examples throughout this doc and were the ones the original bug report was based on:
awesome-malinoto/tracepass-mcp-server— an assessed server (has real validation data).ryanatindago/playwright-mcp-example— a never-successfully-contacted server (initializenever got a response;assessment_statusisnot_assessed).- Fetch either one's live policy/report via e.g.
curl -s "https://verify.sentinelsignal.io/v1/servers/awesome-malinoto/tracepass-mcp-server/report".
Background: how this fix set was found
A user reviewed the live server-detail pages for the two servers above and reported 5 real defects (ranked most → least significant by them) plus 2 items that needed investigation to determine if they were real. Three research passes (by subagents) traced each to exact code. This document is the resulting, already-designed fix plan.
---
Fix 1 (highest priority): compatibility scoring silently treats unavailable evidence as a pass
File: verify/src/mcp_verify/insights.py, function build_client_profiles, the openai_connectors profile's criteria list, specifically the step_up_auth_probe criterion around line 4500-4504.
The bug: this one criterion hand-codes "missing" (meaning the check was never run / evidence unavailable) into its own passing set:
(
str((latest_checks.get("step_up_auth_probe") or {}).get("status") or "missing") in {"ok", "warning", "missing"},
"Step-up auth signals should be documented and connector-friendly.",
"Document step-up auth requirements in a connector-friendly way.",
),
Every other criterion in this same file follows the opposite, correct rule — "unknown is never a pass" (a standing design principle used repeatedly elsewhere in this codebase: see main.py:18974-18985's policy-export evidence_unavailable_fail_closed blocker, the assessment_status tri-state, materialization.py:404-418's ranking gate). The transport_compliance_probe criterion three lines above this one, in the same profile, has a comment explaining exactly why "missing" must be excluded:
(
# R18/standing constraint (round 9 requirements): "missing"
# here can mean initialize already failed and this probe
# was deliberately never attempted (see the R14 skip
# logic in validation/service.py) -- not evidence
# transport compliance is probably fine. Matches the
# already-correct anthropic_remote_mcp profile below,
# which never included "missing" in its passing set.
str((latest_checks.get("transport_compliance_probe") or {}).get("status") or "missing") in {"ok", "warning"},
"Transport compliance failed or did not complete successfully.",
"Resolve the transport compliance failure -- see the transport compliance probe evidence for what specifically broke.",
),
Why this matters: the score computed from these criteria (_make_profile, insights.py:5061-5069, score = round((passed / len(criteria)) * 100, 1)) can render a full 100.0 / "N of N requirements met" verdict when a criterion's evidence was never actually gathered. This is only visible on tracepass today because a separately-computed "blocked" verdict (see build_client_readiness_verdicts, insights.py:1692-1812) happens to contradict the 100.0 score. On a server with nothing else blocking, this would produce a fully silent, unearned "100.0 compatible" result.
The fix: remove "missing" from the passing set, matching the sibling criterion above it exactly:
(
str((latest_checks.get("step_up_auth_probe") or {}).get("status") or "missing") in {"ok", "warning"},
"Step-up auth signals should be documented and connector-friendly.",
"Document step-up auth requirements in a connector-friendly way.",
),
Important — this is a real scoring behavior change, flag it, don't just quietly ship it: any server whose step_up_auth_probe is currently "missing" will lose one passing criterion out of 9 for the OpenAI Connectors profile (100.0 → 88.9, or lower if it was already imperfect elsewhere). Most servers won't cross the 80% "compatible" threshold from this alone, but some marginal ones could. After deploying, spot-check a handful of live servers (tracepass plus 2-3 others whose step_up_auth_probe status is currently "missing") to confirm the new score/verdict combination looks coherent, not just that it changed.
Confirmed not a caching bug: build_client_profiles reads latest_validation.checks directly, as part of the full build_server_insights() computation pipeline — it never goes through the /policy route's fast/fallback builders (which a separate, already-fixed bug this session affected). This is a standalone logic bug in the scoring criteria themselves.
---
Fix 2: "sync tool schemas" remediation fires with zero actual divergences (V1)
File: verify/src/mcp_verify/insights.py, function build_remediations, the check-loop's handling of check_name == "schema_divergence_probe", around lines 4304-4370 (the check-loop starts at 4304; the schema_divergence_probe-specific branch is at 4355-4369; the row gets written into items[code] at 4370).
The bug: the loop's general absence-gate (line 4320) only skips a row when status in {"missing", "not_assessed"}:
if not fires_on_absence and status in {"missing", "not_assessed"}:
continue
It never checks whether tool_divergences (the actual list of field-level mismatches) is non-empty. Meanwhile, build_active_alerts (a different function, same file, insights.py:3963-4013) already correctly distinguishes two cases for this exact probe:
# insights.py:3983-3990 (inside build_active_alerts)
severity_tier = schema_divergence_details.get("severity") or (
"high" if schema_divergence_status == "error" else "medium"
)
if severity_tier == "low":
alerts.append({
"code": "card_omits_tool_schemas",
"severity": "low",
"title": "Server card does not publish tool schemas",
"message": (
f"{len(omitted_schema_tools)} tool(s) are listed on the published server card without an "
f"inputSchema: ..."
". This is a completeness gap, not a contradiction -- clients must call tools/list "
"for parameter details rather than trusting the card."
),
})
severity_tier == "low" is exactly the "only an omission was found, tool_divergences is empty" case. build_remediations never checks this, so the same page snapshot can carry both the correct "completeness gap, not a contradiction" alert and a remediation row asserting "Sync the published server card's tool schemas with the live tool surface" / "The server card and live tools/list disagree on schema details" — a direct self-contradiction, confirmed live on tracepass with tool_divergences: [] in the same snapshot.
The fix: in the schema_divergence_probe branch (right after tool_divergences is extracted at line 4355, before the row is written to items[code] at line 4370), compute severity_tier the same way build_active_alerts does — don't re-derive the "is this a real contradiction" logic independently, reuse the identical formula so the two functions structurally cannot disagree again — and skip the row entirely when it's "low":
if check_name == "schema_divergence_probe":
details = (payload or {}).get("details") or {}
severity_tier = details.get("severity") or ("high" if status == "error" else "medium")
if severity_tier == "low":
continue
why = _schema_divergence_remediation_why(tool_divergences)
(status at this point in the loop is already str((payload or {}).get("status") or "unknown") from line 4305 — same value build_active_alerts calls schema_divergence_status.)
Minor same-fix bonus — a typo: _schema_divergence_remediation_playbook (insights.py:3821-3830) builds its playbook text as:
dimension_text = ", ".join(dimension_labels) if dimension_labels else "the recorded schema details"
return [..., f"Fix the {dimension_text} mismatch(es) per tool."]
When dimension_labels is empty, the default string "the recorded schema details" already contains "the", so the composed sentence reads "Fix the the recorded schema details mismatch(es) per tool." (doubled "the"). This default is only reachable in exactly the same empty-divergence case the fix above now suppresses entirely, so after Fix 2 this text should never render in production — but fix the string anyway (defense in depth, in case another call path reaches it): change the default from "the recorded schema details" to "recorded schema details" (drop the leading "the", since the caller's f-string already supplies it).
---
Fix 3: publisher-facing remediation rows survive on a server that was never contacted (V2, first half)
File: verify/src/mcp_verify/insights.py, function build_remediations.
The bug, in three parts. server_unreachable (computed via _initialize_never_received_a_response, used at line 4298) currently only suppresses 3 checks:
# insights.py, near line 4191
NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS = {"initialize", "oauth_protected_resource", "server_card"}
# insights.py:4329-4330, inside the check-loop
if server_unreachable and check_name in NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS:
continue
tools_listhasfires_on_absence=Trueby design inREMEDIATION_RULES(a deliberate "no news is itself worth flagging" case for the normal scenario) and is not in the suppression set above — so "Ensure tools/list succeeds consistently" fires unconditionally, even when the server was never reached at all (structurally,tools/listcould never have even been attempted ifinitializenever got a response).prompts_listanddeterminism_probehavefires_on_absence=Falseand, wheninitializenever responds, get markedstatus="skipped"byservice.py's_skip_checkhelper (service.py:73-79, called at e.g.service.py:809,824,868). But the general absence-gate at line 4320 only testsstatus in {"missing", "not_assessed"}— never"skipped"— so these rows slip through the exact gate meant to catch them. This produces "Stabilize repeated tools/list responses" (fromdeterminism_probe) and "Repair prompts/list or stop advertising prompts" (fromprompts_list, confirmed live onplaywright— which never advertised prompts at all).- The 5 "Raise X score" rows come from a separate loop, not the check-name loop above, around lines 4386-4403:
for category in score_decomposition:
if category["score"] <= max(4.0, category["max_score"] * 0.35):
code = f"improve_{category['key']}"
items.setdefault(code, {...})
This loop has no server_unreachable check at all — there's no scoring signal worth acting on for a server that was never reached, but the rows fire anyway.
The fix, three parts:
3a. Near NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS (~insights.py:4191), add a second, broader set. Do not just add items to the existing set — it's also used for addressee-routing (insights.py:4343-4345, deciding whether a row reads as "publisher" or "registry" addressed when non_mcp_endpoint is true, a related-but-distinct question from "was the server ever reached at all"). Keep that routing check on the original, narrower set, unchanged:
# A superset of NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS: every check
# that structurally cannot have run if initialize never got a response,
# not just the 3 checks that also need registry-addressed reframing.
SERVER_UNREACHABLE_SUPPRESSED_CHECKS = NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS | {
"tools_list", "prompts_list", "resources_list", "determinism_probe",
}
Change the suppression check at line 4329 to use this broader set:
if server_unreachable and check_name in SERVER_UNREACHABLE_SUPPRESSED_CHECKS:
continue
(Leave line 4343-4345's addressee-routing check — addressee = "registry" if non_mcp_endpoint and check_name in NON_MCP_ENDPOINT_REGISTRY_ADDRESSED_CHECKS else "publisher" — exactly as it is, referencing the original set.)
3b. Also add "skipped" to the general absence-gate at line 4320 (defense in depth — a check can be genuinely "missing"/"skipped" for reasons other than total server unreachability, and a fires_on_absence=False check should never fire on a probe that was never actually attempted, regardless of cause):
if not fires_on_absence and status in {"missing", "not_assessed", "skipped"}:
continue
3c. Wrap the score-category loop (insights.py:4386-4403) so it's skipped entirely when server_unreachable:
if not server_unreachable:
for category in score_decomposition:
if category["score"] <= max(4.0, category["max_score"] * 0.35):
code = f"improve_{category['key']}"
items.setdefault(code, { ... }) # existing body unchanged
---
Fix 4: declared alert-count badge disagrees with the actual alert list (V2, second half)
Files: verify/src/mcp_verify/main.py (the badge) and verify/src/mcp_verify/insights.py (the severity table).
The bug: render_production_readiness (main.py:28410-28463) computes a displayed "N high/critical alerts" badge as:
# main.py:28447-28452
remediation_critical_count = sum(
1 for item in (remediations or [])
if str(item.get("severity") or "") in {"critical", "high"} and item.get("addressee") in (None, "publisher")
)
displayed_critical_count = max(int(readiness.get("critical_alerts") or 0), remediation_critical_count)
readiness["critical_alerts"] is correctly derived from the real active_alerts list. remediation_critical_count is a second, independent count built from remediations — and each remediation's severity only gets reconciled against what's actually in active_alerts for 2 check types, via:
# insights.py:4205-4208
CHECK_TO_ALERT_CODES = {
"schema_divergence_probe": ("server_card_schema_drifted", "card_omits_tool_schemas"),
"provenance_divergence_probe": ("official_registry_entry_drifted",),
}
...consumed by resolve_remediation_severities (insights.py:4211-4244). Every other check type keeps its static REMEDIATION_RULES severity (e.g. initialize → "critical", tools_list → "critical", server_card/oauth_* → "high") regardless of what actually survived build_active_alerts' own suppression/dedup logic into the final active_alerts list. Confirmed live: tracepass declares "1 high/critical alert" while its one actual active alert is severity "low"; playwright declares "4" while its one actual active alert is severity "medium".
**User's explicitly chosen fix direction (conservative — keep the existing max()-union design, just make the reconciliation complete)**: extend CHECK_TO_ALERT_CODES to cover every check type in REMEDIATION_RULES (insights.py:139+) that can produce a corresponding active_alerts entry, mapping each check name to the actual alert code(s) build_active_alerts emits for it. This requires reading both tables carefully — REMEDIATION_RULES (starting insights.py:139) for the full list of check names and their static severities, and build_active_alerts (search for alerts.append( inside it) for the actual code string(s) each check type can produce — and building an accurate name → code(s) mapping for each. Do this as a careful, complete pass, not a quick partial patch; an incomplete mapping just narrows the gap instead of closing it.
Once CHECK_TO_ALERT_CODES is complete, resolve_remediation_severities will cap every remediation's severity at what's real in active_alerts for every check type, so remediation_critical_count can never exceed what active_alerts itself supports — closing the gap without touching the max()-union structure in main.py at all.
---
Investigated and intentionally NOT included in this fix set
Two items from the original feedback were investigated and determined not to need a code change. Do not "fix" these — they were checked carefully and are either working as designed or unreproducible:
- Smithery shows "compatible" at score 80.0 with one blocker/missing-requirement present. Not a defect.
_make_profile(insights.py:5061-5069) setscompatibility = "compatible"atscore >= 80— the same threshold used for every client profile in this file. "Compatible" has never meant zero missing requirements; Smithery's profile has 5 criteria, and 4-of-5 (80%) is exactly the threshold the design tolerates. Separately: Smithery is not inbuild_client_readiness_verdicts's broader blocker-reconciliation loop (onlyopenai_connectors/claude_desktopare,insights.py:1720-1723), but Smithery has no correspondingpublishability_policy_profile/OAuth-connector policy gate in this codebase for that broader reconciliation to meaningfully add — so the omission is inert, not a live gap.
- **
evidence_confidence.scorereads null whilefreshness.confidence_scorereads 50.0 in the same tracepass snapshot.** Investigated live twice (~15 seconds apart) ontracepass's real (non-fallback)/policyresponse —current_snapshot.evidence_confidence.scoreandcurrent_snapshot.freshness.confidence_scorewere consistently equal both times (both null in a fallback-served response, both a real non-null number once promoted to the real computed value).build_trust_snapshot(main.py:18266-18309) derives both fields from the identicalpayload.get("evidence_confidence")in a single pass, so they cannot structurally drift apart within the samecurrent_snapshot. The split originally reported was very likely captured during the same/policycold-cache fallback window that was already fixed elsewhere this session (the fallback path hand-rolls its own separate placeholderevidence_confidencedict, unrelated to the real computation) — this does not currently reproduce and no fix is proposed. If it's seen again, re-confirm live (fetch/policytwice, ~10-15s apart, checking for apartial: true/partial: falsetransition) before assuming a code bug.
---
Implementation order
Recommended order (independent fixes, but this order keeps related context together and de-risks the highest-priority item first):
- Fix 1 (compatibility scoring) — standalone, single file/function, most important.
- Fix 2 (V1 schema-sync suppression + typo) — standalone, single function.
- Fix 3 (V2 publisher-rows-on-unreachable-server) — standalone, single function, three sub-changes.
- Fix 4 (V2 alert-count reconciliation) — needs the most careful cross-referencing of two tables; do last so it isn't rushed.
Verification
- After each fix:
PYTHONPATH=verify/src:verify/tests python -m pytest verify/tests -q(must stay green — currently 770 passing before this work starts). - New tests to add (one or more per fix, in
verify/tests/, following existing patterns in that directory — e.g.test_api.pyortest_score_integrity.py): - Fix 1: a fixture with
step_up_auth_probestatus"missing"on an otherwise-perfect OpenAI Connectors profile — assert the score is no longer100.0and the criterion appears inmissing_requirements. - Fix 2: a fixture with
tool_divergences: []and only an omitted-schema finding — assert the "sync tool schemas" remediation row is absent fromremediations, while the existing low-severity "completeness gap" alert is still present inactive_alerts. - Fix 3: a fixture matching playwright's shape (
initializenever responded /server_unreachabletrue) — asserttools_list,prompts_list,determinism_probe, and allimprove_*(score-category) rows are absent fromremediations. - Fix 4: a fixture exercising the newly-covered part of
CHECK_TO_ALERT_CODES— a check type outside the original 2 whose realactive_alertsseverity is lower than its staticREMEDIATION_RULESseverity — assert the displayed critical-alert count no longer exceeds whatactive_alertssupports. python -m pytest -m unit -qfrom repo root (pre-commit will also run this automatically ongit commit).- Live, after deploy (only if/when the user asks for a deploy):
curl -s "https://verify.sentinelsignal.io/v1/servers/awesome-malinoto/tracepass-mcp-server/report"— confirm the schema-sync remediation row is gone, the declared critical-alert count now matches its actual (low-severity) alert, and its OpenAI Connectors / Claude Desktop compatibility scores are no longer100.0from unavailablestep_up_auth_probeevidence.curl -s "https://verify.sentinelsignal.io/v1/servers/ryanatindago/playwright-mcp-example/report"— confirm thetools_list/prompts_list/determinism_probe/score-category publisher-addressed rows are gone fromremediations.- Spot-check 2-3 other live servers whose
step_up_auth_probeis currently"missing"to confirm Fix 1's score change looks coherent (not just "different").
Commit / push (only if the user asks for this — do not push automatically)
Standard flow: stage the changed files specifically (not git add -A), commit with a message explaining the "why" (this doc's context section has the substance), push to origin master. The pre-commit hook will auto-bump VERSION and run pytest -m unit — if that fails, fix the issue and create a new commit, don't amend.