Sentinel Signal

Verify: MCP tool fixes from the 2026-10-01 traffic review

Source: docs/mcp-verify-mcp-tool-fixes.md

Document Content

Verify: MCP tool fixes from the 2026-10-01 traffic review

What the traffic showed

The 2026-10-01 daily report showed 84 sampled tools/call invocations, up from 30. The full day (Loki, not the 25k-event sample) had 1,761 /mcp requests and 88 tools/call. Most of the calls came from indexers and dataset generators, not end-user agents:

| Client (user agent) | Calls | Pattern | |---|---|---| | Toucan-Datagen/1.0 (Cloudflare egress) | 53 | Synthetic tool-use trajectories (search → recommend → compare → fetch) | | BrickBlueBot (agentic-web registry) | 18 | Every tool exactly once | | No user agent | 9 | Every tool exactly once | | Tendle, husksecurity agent-index-prober, rokmcp | 7 | Indexers |

So the multi-step sessions are training-data generation, not a demand signal. They did expose three real defects.

Fixes

  1. **compare_servers failed on every call** (6/6, JSON-RPC -32000: int() argument must be ... not 'Query'). The MCP handler called the FastAPI route function compare_servers_api directly and omitted its Query-defaulted parameters. A direct Python call never resolves those defaults, so they arrived as Query(...) marker objects.
  • Fix: pass every parameter explicitly.
  • Same bug elsewhere: the /teams/{slug} page called get_team_digest without window. "daily" never equalled the Query object, so the page silently showed a 7-day digest instead of a daily one. Fixed the same way.
  • Guard: a test walks main.py's syntax tree and fails if any direct call to a route function omits a FastAPI-defaulted parameter.
  1. **search_servers, search and recommend_servers took 7–13 s** (every slow /mcp sample). Each call rebuilt the full trust match context for every scanned candidate:
  • ~179 servers and ~3.5k validation rows
  • ~7.4 s of CPU, against ~13 ms for scoring and payload

The context depends only on the server, never the query. agent_context_cache.py now caches contexts per server:

  • Validity: keyed by the server row's (last_validated_at, updated_at, current_score), so a new validation rebuilds that server immediately.
  • TTL: MCP_VERIFY_AGENT_CONTEXT_CACHE_SECONDS, default 300. Setting it to 0 restores the old rebuild-every-call behaviour.
  • Concurrency: single-flight builds, stale-while-refresh, a bounded cold wait, and a back-off after a failed refresh (the 09-28 stampede lesson).
  • Memory: LRU of 5,000 entries. raw_evidence is dropped once machine_summary has been derived, taking a 179-server set from ~21 MB to ~5.5 MB.
  • Correctness: scoring and ranking still run on every request, and tests show cached results are identical to uncached ones.
  1. **Intermittent test_machine_payload_routes_use_fast_fallback_when_enabled** (~1 run in 6, including on the pre-change commit). The test's in-memory SQLite StaticPool shares one connection across threads, and a background route-cache refresh overlapping the next request corrupted its cursor state. The test now defers spawned refreshes until the response has been returned. Production code is unchanged.

Not changed

  • The 404 and 400 errors on agent tools and fetch come from probers sending made-up IDs or missing arguments. Those responses are correct.
  • The per-process cache limitation is the same as for observed attention: production verify-web runs a single uvicorn process.