Verify: MCP tool fixes from the 2026-10-01 traffic review
What the traffic showed
The 2026-10-01 daily report showed 84 sampled tools/call invocations, up from 30. The full day (Loki, not the 25k-event sample) had 1,761 /mcp requests and 88 tools/call. Most of the calls came from indexers and dataset generators, not end-user agents:
| Client (user agent) | Calls | Pattern | |---|---|---| | Toucan-Datagen/1.0 (Cloudflare egress) | 53 | Synthetic tool-use trajectories (search → recommend → compare → fetch) | | BrickBlueBot (agentic-web registry) | 18 | Every tool exactly once | | No user agent | 9 | Every tool exactly once | | Tendle, husksecurity agent-index-prober, rokmcp | 7 | Indexers |
So the multi-step sessions are training-data generation, not a demand signal. They did expose three real defects.
Fixes
- **
compare_serversfailed on every call** (6/6, JSON-RPC -32000:int() argument must be ... not 'Query'). The MCP handler called the FastAPI route functioncompare_servers_apidirectly and omitted itsQuery-defaulted parameters. A direct Python call never resolves those defaults, so they arrived asQuery(...)marker objects.
- Fix: pass every parameter explicitly.
- Same bug elsewhere: the
/teams/{slug}page calledget_team_digestwithoutwindow."daily"never equalled theQueryobject, so the page silently showed a 7-day digest instead of a daily one. Fixed the same way. - Guard: a test walks
main.py's syntax tree and fails if any direct call to a route function omits a FastAPI-defaulted parameter.
- **
search_servers,searchandrecommend_serverstook 7–13 s** (every slow/mcpsample). Each call rebuilt the full trust match context for every scanned candidate:
- ~179 servers and ~3.5k validation rows
- ~7.4 s of CPU, against ~13 ms for scoring and payload
The context depends only on the server, never the query. agent_context_cache.py now caches contexts per server:
- Validity: keyed by the server row's
(last_validated_at, updated_at, current_score), so a new validation rebuilds that server immediately. - TTL:
MCP_VERIFY_AGENT_CONTEXT_CACHE_SECONDS, default 300. Setting it to 0 restores the old rebuild-every-call behaviour. - Concurrency: single-flight builds, stale-while-refresh, a bounded cold wait, and a back-off after a failed refresh (the 09-28 stampede lesson).
- Memory: LRU of 5,000 entries.
raw_evidenceis dropped oncemachine_summaryhas been derived, taking a 179-server set from ~21 MB to ~5.5 MB. - Correctness: scoring and ranking still run on every request, and tests show cached results are identical to uncached ones.
- **Intermittent
test_machine_payload_routes_use_fast_fallback_when_enabled** (~1 run in 6, including on the pre-change commit). The test's in-memory SQLiteStaticPoolshares one connection across threads, and a background route-cache refresh overlapping the next request corrupted its cursor state. The test now defers spawned refreshes until the response has been returned. Production code is unchanged.
Not changed
- The 404 and 400 errors on agent tools and
fetchcome from probers sending made-up IDs or missing arguments. Those responses are correct. - The per-process cache limitation is the same as for observed attention: production verify-web runs a single uvicorn process.