MCP Verify methodology
How MCP Verify scores servers, handles freshness, reports confidence, and keeps paid status separate from objective trust evidence.
What the score means
The public Verify score is evidence-based and derived from validation, metadata, transport, auth, tool-surface, compatibility, and maintenance signals.
- Scores are objective evidence summaries, not paid placement.
- Billing and claim state are never scoring inputs: computing a server's score never reads whether it is claimed, paid, or on any plan tier. A CI test recomputes rankings with billing state stripped and asserts identical ordering on every run.
- Authenticated validation is a paid feature that inspects real authenticated behavior unauthenticated probing cannot reach, so it can produce stronger evidence -- and stronger evidence can change a score, the same way any new evidence would for an unclaimed server that gets revalidated. This is a real, disclosed evidence-path effect, not billing state itself affecting the score.
- Server pages show the score with percentile, evidence age, confidence, and risk drivers so the number is not interpreted in isolation.
Status, score, and verdict
Status, score, and verdict are three views of the same evidence, not three independent judgments. Status is Healthy, Degraded, or Failing, from the live validation checks. Verdict (Safe for production / Safe for evaluation / Needs remediation / Metadata only) is a single gate derived from status, score, freshness, and tool count -- it is computed in exactly one place in the codebase so a row can never show a status and a verdict that contradict each other.
- Healthy requires a clean initialize and tool listing, with consistent OAuth discovery metadata.
- Degraded means the server works but something is inconsistent, most commonly mismatched OAuth protected-resource/authorization-server metadata (the dominant real-world cause; a server merely requiring auth does not by itself degrade status).
- Failing means the core initialize or tool-listing check did not succeed.
- Safe for production requires Healthy status, score at least 80, and fresh evidence.
- Safe for evaluation requires Healthy or Degraded status and score at least 60 -- Degraded can reach this tier but never Safe for production.
- Metadata only applies whenever there is no completed validation or zero inspected tools, regardless of score.
- Score carries an explicit status penalty (Degraded and Failing are marked down after the component average, not just diluted as one signal among many) and a reduced ceiling when zero tools were inspected, so an uninspected or degraded server cannot outscore an equivalent healthy, tool-bearing one.
- Scoring dimensions that require tools to inspect (destructive-operation, network-egress, execution-sandbox, data-exfiltration, least-privilege, and secret-handling safety) award low, not high, credit when a server has zero tools -- absence of evidence is not evidence of safety.
Production readiness versus recommended runtime policy
These are two different questions Verify answers about the same server, and neither one overrides the other. Production readiness (the verdict above -- Safe for production / Safe for evaluation / Needs remediation / Metadata only) answers: how suitable is this server for evaluation or production adoption, based on current evidence? Recommended runtime policy (shown on the server page's Risks tier, and available in full via the policy export endpoint) answers a narrower, different question: under what runtime authorization policy should an agent be allowed to call this server's tools right now?
- A server can be "Safe for evaluation" and still have its recommended policy read "Allow with approval" -- readiness is about the server as a whole; policy can additionally gate on a specific tool's write/risk profile, evidence-age policy breach, or OAuth posture.
- The policy shown on a server page uses a default stance (medium risk ceiling, write actions require human approval, no OAuth requirement) -- the policy export endpoint lets a caller customize every one of those inputs for their own risk tolerance.
- Do not read the runtime policy panel as a second, competing verdict. If the two ever seem to disagree, production readiness describes the server; recommended policy describes what an agent may do under specific constraints.
Freshness
The words "fresh" and "stale" refer to exactly one window everywhere on this site: 24 hours, matching the Trust Index Fresh coverage stat and the results-table Freshness column. Every other window below serves a distinct, differently-named purpose and never uses the word fresh or stale.
- Two stats can both be built on this same 24h window and still answer different questions: the homepage's "Healthy & Fresh" figure is a compound count -- validated within the window AND currently healthy -- while the Trust Index's "Fresh" figure counts any validation attempt inside the window, outcome not required. Same window, different population; look at each stat's own hint text for which one it is.
- Fresh evidence is preferred in default ranking; the fresh/stale label is never shown for two different thresholds on the same page.
- Validation history is shown as a trend only after at least two validation runs.
Percentile and confidence
Percentile compares a server to the currently scored public catalog. Confidence communicates evidence completeness, recency, and validation density.
- Percentile is a catalog-relative signal and may move as the catalog grows.
- Confidence is not a replacement for the score or the production decision.
- Buyers should review top risk drivers before approving a server for production use.
Limits
Verify cannot prove every runtime behavior from public metadata alone.
- Authenticated validation provides stronger evidence than unauthenticated metadata.
- Write and execution-capable tools require policy review even when the server initializes cleanly.
- Publishers can improve profiles by claiming ownership, adding metadata, and keeping validation evidence fresh.
Opted-out servers
Verify honors robots.txt. When a server's own robots.txt disallows Verify's validator, that is an explicit operator opt-out, not a validation failure, and it is never treated as one. The server keeps its listing entry -- name, registry provenance, and transport are still shown -- but no score, verdict, decision label, or risk claim is published for it, and the page is excluded from search indexing until it carries an explicit opt-out notice with a route for the owner to enable validation. The share of the catalog that has opted out is reported as its own line in the coverage stats on the homepage, not folded into any other status.
More on scoring
This page is what the score means today. For exact scoring dimensions, equations, weights, floors, caps, and zero-point behavior, see Scoring specification. For the history of scoring corrections and methodology changes, see Methodology changelog.