Sentinel Signal

Agent Reliability documentation

Source: docs/agent-reliability-hosted-api.md

Document Content

Agent Reliability documentation

This mirrors the public /docs/agent-reliability-hosted-api page. Operator/deployment material (env var names, secret paths, kill switches, Stripe plumbing, GitHub App registration steps) lives in docs/agent-reliability-production-runbook.md, not here -- this file is written for the developer using the product, per the Agent Reliability productization pass.

Overview

Agent Reliability remembers the MCP tool contract your team approved and compares every future change against it. Install the GitHub App, register a target, and every push or pull request gets a GitHub Check that explains what changed, why it matters, and whether your deployment policy allows it.

Quickstart

  1. Install the GitHub App: /agent-reliability
  2. Register at least one target (a public HTTPS MCP endpoint) for the repository.
  3. Open a pull request that changes your MCP server -- the Check reports the result.

The first scan for a repository establishes its baseline; every scan after that is compared against it.

Repository setup

Install the GitHub App and grant it access to the repositories you want covered. Each installed repository becomes available to register targets against.

Targets

A target is the MCP endpoint Agent Reliability scans on your behalf. Targets are public HTTPS MCP endpoints today; each repository can register multiple targets (for example, staging and production). Each target can optionally be put on a scheduled-scan interval independent of repository activity.

Baselines

Agent Reliability remembers the tool contract your team approved. Every future scan is compared against that accepted state, not against whatever the server happens to look like today. A completed, non-superseded, non-pull-request scan against your default branch that passes policy automatically promotes its snapshot as the new baseline -- pull-request scans never do, so a change under review can't silently redefine what "approved" means before it's merged. Any completed, non-superseded run that isn't blocked can also be promoted manually (POST /v1/agent-reliability/runs/{id}/baseline/promote), except a pull-request run -- promoting a baseline from an unmerged PR would have the same silent-redefinition problem the automatic case exists to prevent.

A pull-request run that lands on approval_required therefore cannot be resolved by promoting a baseline at all. POST /v1/agent-reliability/runs/{id}/approve is the separate action for that: a recorded human sign-off (actor, reason, timestamp) that flips the run's GitHub Check Run to a passing conclusion without changing any target's baseline. The run's own decision is left as the historical automated verdict -- approval is recorded alongside it, not instead of it.

Semantic diff

Every scan produces a semantic contract comparison, not a raw JSON diff. Changes are classified two ways: how severe (info/low/medium/high/critical) and whether they're breaking for existing callers -- these are independent, because a low-severity change can still break callers and a high-severity one can be fully backward compatible.

  • Tool surface: tools added, removed, or renamed; changed descriptions.
  • Capability & risk: new write capability, new destructive operation, permission expansion.
  • Contract: required field added, optional -> required, field removed, type or enum change.
  • Authentication: anonymous -> authenticated (or the reverse), OAuth/API-key changes.
  • Transport & runtime: endpoint/transport changes, latency regression.
  • Trust & evidence: Verify score or readiness regression on the underlying server.

Policies

You define what an unacceptable change looks like in a repository deployment policy. Each scan evaluates to one decision: pass, warn, approval_required, or block.

Example rules: block new destructive tools, warn on new write capabilities, block breaking schema changes, require a minimum Verify readiness, block severe authentication regression.

GitHub Checks

The Check is where you see the result. It reports a title summarizing the decision and change count, a one-line summary of the most important finding, and a full Markdown body organized by severity (Critical/High/Medium/Low/Info) naming exactly which policy rule fired and what to do next.

Scheduled monitoring

Your code does not have to change for your agent dependencies to change. Once a target is registered, you can enable scheduled scans that recheck it on an interval, independent of repository activity, and alert when an upstream MCP server introduces contract, permission, authentication, or trust changes outside your repository.

CLI / operator workflow

The verify-agent CLI (pip install sentinel-verify-agent) can drive the full operator lifecycle against the hosted API above, not just local scan/diff. A typical sequence:

verify-agent policy validate .verify-agent.yml
verify-agent monitor run --repository <repo-id> --target <target-id> --wait
verify-agent monitor status --run <run-id>
verify-agent approve <run-id> --reason "reviewed and safe to merge"   # only if the run landed on approval_required
verify-agent baseline promote --run <run-id> --reason "manual review passed"
verify-agent baseline rollback --repository <repo-id> --target <target-id> --baseline <baseline-id> --reason "bad deploy"

approve is the operation that closes a real gap: a pull-request-triggered run stuck at approval_required cannot be resolved by promoting a baseline (see "Baselines" above) -- approve is the separate, correct action, a recorded human sign-off that flips the run's GitHub Check without touching any baseline. See verify-agent --help for the full command reference, or sentinel-verify-agent on PyPI to install it.

Evidence

Every run produces a structured evidence artifact (JSON and Markdown) recording the diff, the policy evaluation, and both the baseline and candidate snapshot fingerprints. Evidence is hash-chain protected today. Evidence artifacts can be cryptographically signed for independent verification on supported deployments.

Security & privacy

The GitHub App requests the minimum permissions needed to do its job -- nothing more:

  • Metadata (read): required by GitHub for any App installation; used to identify installed repositories.
  • Checks (read/write): required to post the Agent Reliability Check result -- decision, diff summary, and next steps -- on your commits and pull requests.
  • Contents (read): used only to fetch a repository-committed deployment policy file, when that feature is enabled.
  • Pull requests (read): required to associate a run with its pull request context.

Installation access tokens are short-lived and kept in process memory only -- never persisted to application tables, logs, or evidence. Credentials for scanning authenticated targets are encrypted at rest using the same vault pattern as Verify's other stored credentials. Uninstalling the GitHub App from a repository stops all future scans for it immediately.

Deletion

You can permanently delete your Agent Reliability workspace data yourself, at any time, from your workspace dashboard -- no need to contact support. This is separate from uninstalling the GitHub App: uninstalling stops future scans, deletion removes stored data. Deletion is irreversible and does not affect your Team or any other Verify data. See the privacy policy for the full list of what's deleted.

Quotas

Free Beta and Team are the two available plans today. Confirm current numbers on /docs/agent-reliability-hosted-api or /agent-reliability -- both render the real, live-configured limits rather than a value hardcoded independently in this file.

Troubleshooting

A scan result is always one of three different kinds of outcome, and they are never confused with each other: a policy decision (pass/warn/approval_required/block) is a finding about your MCP server's contract; a configuration error means the scan couldn't run because of how it's set up (for example, an invalid policy file); an infrastructure error means the scan itself failed to complete. A broken scan or an infrastructure problem is reported as needing attention -- it is never presented as a policy block, so it never reads as "your MCP server is unsafe" when the real problem was on the scanning side.

API

Supported external routes. All require authentication except the public overview and install pages.

  • GET /v1/agent-reliability/overview
  • GET /v1/agent-reliability/dashboard
  • GET /v1/agent-reliability/billing
  • POST /v1/agent-reliability/billing/checkout
  • GET/POST /v1/agent-reliability/repositories
  • GET/PUT/DELETE /v1/agent-reliability/repositories/{id}
  • GET/POST /v1/agent-reliability/repositories/{id}/targets
  • PUT/DELETE /v1/agent-reliability/repositories/{id}/targets/{target_id}
  • GET/PUT /v1/agent-reliability/repositories/{id}/targets/{target_id}/schedule
  • GET/PUT /v1/agent-reliability/repositories/{id}/deployment-policy
  • POST /v1/agent-reliability/scans
  • GET /v1/agent-reliability/runs
  • GET /v1/agent-reliability/runs/{id}
  • GET /v1/agent-reliability/runs/{id}/evidence
  • GET /v1/agent-reliability/runs/{id}/evidence.html
  • POST /v1/agent-reliability/runs/{id}/baseline/promote
  • POST /v1/agent-reliability/repositories/{id}/targets/{target_id}/baselines/promote
  • POST /v1/agent-reliability/repositories/{id}/targets/{target_id}/baselines/rollback
  • POST /v1/agent-reliability/runs/{id}/approve
  • POST /v1/agent-reliability/diffs
  • POST /v1/agent-reliability/deployment-policy/evaluate

Scope model

The router requires these scopes: agent-reliability:read, agent-reliability:write, agent-reliability:run, agent-reliability:admin. The adapter prefers the existing Sentinel API auth stack when available and falls back to a Verify-local JWT scope checker when the Verify package is run independently.