MCP Verify scoring specification
Exact scoring dimensions, weights, zero-point behavior, and edge-case semantics behind the Verify score, for engineers, security reviewers, auditors, and technically sophisticated buyers.
What this page covers
This page documents the exact mechanics behind the Verify score: dimensions, zero-point behavior, and experimental candidates. For what a Verify score means in plain terms -- without the implementation detail -- see Methodology instead.
Windows in use
Every time-based window used anywhere on the site, and what it means. Only the first row uses the words "fresh"/"stale".
| Window | Name | Used for |
|---|---|---|
| 24h | Fresh / Stale (the only public freshness claim) | Trust Index Fresh coverage stat, Status page's Validation fresh bucket, results-table Freshness column, active-alert threshold, Freshly Validated badge, homepage Healthy & Fresh stat (which additionally requires status=healthy -- the other five uses do not) |
| 48h | Strong live window | "Strong live candidates" stat/filter; publisher launch checklist ("Recent validation (48h)") |
| 7d (168h) | Evidence retained: within 7d | Evidence-age tier display and confidence-decay weighting only -- not a freshness claim |
| 30d (720h) | Evidence retained: within 30d / display-score suppression | Evidence-age tier display; scores older than this are suppressed from display entirely |
| 24h / 168h / 720h | Validation cadence target -- guaranteed only for Enterprise (Enterprise / Pro / Community) | Not a contractual promise for Community/Pro; see /pricing and /trustops for current guarantee status |
The Trust Index "Fresh" stat and the Status page's "Validation fresh" bucket both require a recorded endpoint and both use the same 24h window above -- they can still show different counts because Trust Index Fresh counts every indexed registry-source row that passes the freshness check, including duplicate rows where the same server appears under more than one registry source, while Status page freshness counts one canonical row per server identity, after deduplicating those registry-source duplicates. Trust Index Fresh is therefore a catalog-inventory count; Status freshness is a canonical-population count -- with the same eligibility requirements, differing only by that dedup. Neither number is wrong; they answer the same question at two different levels of deduplication.
Subscore zero points
Every one of the 49 scoring dimensions defines a zero point: the value it is assigned when the check its evidence depends on errored or never ran. Per the mid-2026 scoring remediation review, that value must be 0 unless there is a written reason the dimension's evidence is genuinely independent of live check success (for example, registry-provided metadata that does not require a live connection). A server that fails to complete an MCP handshake at all will score at or near the zero point on nearly every dimension, not receive default credit for the dimensions nobody had a chance to observe. A third case, distinct from both credit and zero: when a server's own robots.txt disallows Verify's validator, the checks that would require probing beyond the core handshake are marked not_assessed rather than run, and every dimension whose evidence depends on one of those checks is excluded from the composite denominator entirely -- not scored as a failure, and not credited as a pass. See "Opted-out servers" on the Methodology page.
| Dimension | Zero point | Why |
|---|---|---|
| Auth Operability | 0.0/10 | No unconditional default; credit is only added for confirmed initialize/OAuth states. |
| Error Contract Quality | 0.0/10 | R1: no error responses observed is only informative if the server was reachable enough to error cleanly. |
| Rate-Limit Semantics | 6.0/10 | Justified non-zero: absence of a rate-limit response is absence of a rare event, not absence of core evidence -- most healthy servers never trigger one. |
| Schema Completeness | 0.0/10 | Zero when no tools were discovered at all. |
| Backward Compatibility | 6.0/10 | Justified non-zero: a first-ever validation has no prior snapshot to diverge from, so 'no drift detected' is definitionally true, not defaulted. |
| SLO Health | 0.0/10 | No default branch; the availability/latency/stability formula naturally floors near zero without any successful runs. |
| Security Hygiene | 0.0/10 | No unconditional default; credit is only added for observed HTTPS/header evidence. |
| Task Success | 0.0/10 | Criteria-based; every criterion requires a confirmed check status, so full failure naturally floors at 0. |
| Trust Confidence | 0.0/10 | Wilson lower-bound formula; naturally near zero with no confirmed successes. |
| Abuse/Noise Resilience | 2.0/10 | Already gated on current_core_success inline (7.0 vs 2.0); the failure-path value is a genuine floor, not a generous default. |
| Prompt Contract | 6.0/10 | Justified non-zero for a genuinely-attempted-and-failed check (error/warning/auth_required without supports_prompts) -- a server that actively responded is not the same as one that was never asked. R23 (round 10) narrowed this: when the check specifically never ran at all (missing) on a confirmed-healthy server, it excluded (None) a server that never claimed prompt support, zero-anchoring (0) one that did. R39.1 (round 13) reverted that exclusion -- missing now zero-anchors at 0 unconditionally, regardless of whether the server claims prompt support, and stays in the composite denominator; None-exclusion is reserved for genuine not_assessed (Verify itself couldn't check), not "server doesn't need this." |
| Resource Contract | 6.0/10 | Justified non-zero for a genuinely-attempted-and-failed check, same reasoning as prompt_contract_score, including the R23-then-R39.1 history: the missing-check branch zero-anchors at 0 unconditionally now, no longer excluding (None) servers that don't claim resource support. |
| Discovery Metadata | 0.0/10 | Justified non-zero in practice (registry-provided title/description/homepage fields are independent of live reachability), but the zero point when a server genuinely has no registry listing at all is 0. |
| Registry Consistency | 0.0/10 | Comparison-based across sources; naturally reflects disagreement, no unconditional generous default. R25.3-pattern correction (round 11 requirements): this entry's claim was previously false in practice -- the underlying ratio helpers (text_similarity, host_match_ratio, transport_match_ratio, numeric_consistency_ratio, bool_consistency_ratio) each returned 0.6 when both sides were absent, silently crediting comparisons that never ran; a sparse server with little registry/metadata data averaged in several of these. Fixed by returning None (excluded from the average) instead, so this component is now genuinely comparison-based, matching what this entry always claimed. |
| Installability | 0.0/10 | Weighted from protocol-conformance/execution-success sub-scores, both of which require confirmed initialize success to score above floor. |
| Session Semantics | 0.0/10 | No unconditional default; credit is only added for confirmed initialize/tools_list success. |
| Tool Surface Design | 0.0/10 | Zero when no tools were discovered; non-zero for genuinely-discovered tool schemas regardless of overall connectivity, since tool shape itself is real evidence once tools are known. |
| Result Shape Stability | 0.0/10 | Zero when no current tools are known. |
| OAuth Interop | 0.0/10 | R23 (round 11 requirements): tightened the 'protected-resource check failed but initialize succeeded' branch from a 7.0 default to 0 -- this metric measures depth *beyond* the protected-resource check, so it has nothing to credit when that check itself is error/missing, regardless of whether the overall handshake succeeded. Built from score_oauth_maturity, which starts at 0 and only credits confirmed OAuth/OIDC endpoint checks. |
| Recovery Semantics | 0.0/10 | R1: tightened from a 5.0 default (no confirmed handshake) to 0. |
| Maintenance Signal | 0.0/10 | Justified non-zero in practice (registry version/update-recency metadata is independent of live reachability); genuinely 0 only when the registry has none of that metadata either. |
| Adoption Signal | 0.0/10 | Justified non-zero in practice (registry-source and directory-presence signals are independent of live reachability); genuinely near 0 only with no registry signal at all. |
| Freshness Confidence | 0.0/10 | Decay component naturally floors at 0 with no confirmed successes; the density component is deliberately about revalidation recency, not success. |
| Transport Fidelity | 2.0/10 | Correction: the +2.0/+1.5 declared/inferred-transport credit is static/inferred metadata (server.transport_type, URL shape), independent of live reachability -- same reasoning as discovery_metadata_score. This entry previously claimed 0.0 ("confirmed evidence only"), which the code never actually did for the declared/inferred component; corrected to match actual behavior rather than silently changing that behavior. Content-type/handshake credit remains live-check-gated as documented. |
| Spec Recency | 0.0/10 | R1: tightened the missing-probe branch from a 2.0 default (no confirmed handshake) to 0. |
| Session Resume | 0.0/10 | R1: the missing-probe branch now requires a confirmed handshake before crediting a transport-based default. |
| Step-Up Auth | 0.0/10 | R1: tightened the no-probe and missing-probe branches from 6.0/6.5/4.5 defaults (no confirmed handshake) to 0. R23 (round 10) additionally excluded (None) a confirmed-healthy server that never claims OAuth. R39.1 (round 13) reverted that exclusion -- missing zero-anchors at 0 unconditionally now, regardless of server.has_oauth, and stays in the composite denominator; a server with nothing to step up from is not a not_assessed case. |
| Transport Compliance | 0.0/10 | Already correctly gated on confirmed core success for the no-probe branch. R17.1: when the source check is not_assessed (owner_opt_out), this component is None -- excluded from the composite denominator entirely, not scored 0. |
| Utility Coverage | 0.0/10 | R1: tightened from a 1.5/1.0 partial-credit default to a full 0 -- no probe run is not evidence advanced utilities are present. |
| Advanced Capability Coverage | 0.0/10 | Already effectively 0 at the public 0-4 point scale (1.0/10 rounds down to 0 points); left as a small distinguishing floor internally. |
| Connector Publishability | 0.0/10 | Already correctly gated on confirmed core success for the no-probe branch. |
| Tool Snapshot Churn | 7.0/10 | Justified non-zero: a first-ever validation has no prior tool snapshot to have churned from. |
| Connector Replay | 0.0/10 | R1: tightened the no-probe and missing-probe branches from 6.5/7.0 defaults (no confirmed handshake) to 0. |
| Request Association | 0.0/10 | R1: tightened the no-probe and missing-probe branches from 6.5/7.0/5.0 defaults (no confirmed handshake) to 0. R23 (round 10) additionally excluded (None) a confirmed-healthy server whose probe reports no roots/sampling/elicitation advertised at all. R39.1 (round 13) reverted that exclusion -- missing zero-anchors at 0 unconditionally now, regardless of advertised capabilities, and stays in the composite denominator. |
| Interactive Flow Safety | 0.0/10 | R1: the no-probe metadata-text fallback now requires a confirmed handshake before crediting 6.5/8.0 -- that credit is unaffected by R39.1, it's a separate mechanism for servers that DO claim elicitation/sampling. R23 (round 10) additionally excluded (None) a confirmed-healthy server that advertises neither at all; R39.1 (round 13) reverted that specific exclusion -- it zero-anchors at 0 unconditionally now and stays in the composite denominator. |
| Action Safety | 0.0/10 | R1: the no-probe branch now requires a confirmed handshake before crediting up to 8.0 for 'no high-risk tools flagged'. R18: the probe itself now returns not_assessed (not ok) when its finding derives from an empty, unconfirmed observation set -- excluded from the composite denominator entirely in that case, not scored 0. |
| Official Registry Presence | 4.0/10 | Justified non-zero: registry_source is a static classification independent of live reachability. |
| Provenance Divergence | 0.0/10 | R1: tightened the no-probe and missing-probe branches from 6.0/5.5 defaults to 0. R25.3: when fewer than two sources (registry, server_card) are readable, this component is None -- excluded from the composite denominator entirely, not scored 0 and not scored a pass; a comparison that never ran is not confirmed agreement. |
| Safety Transparency | 0.0/10 | Criteria-based; most criteria are static metadata (justified) but the live-auth criterion requires a confirmed check, and full failure floors low. |
| Tool Capability Clarity | 0.0/10 | Zero when no tools were discovered; non-zero for genuinely-discovered tool schemas, same reasoning as tool_surface_design_score. |
| Destructive Operation Safety | 1.5/10 | Zero-tool floor already fixed in a prior round. R1: the 'no destructive tools flagged' branch (8.5) now also requires a confirmed handshake, else 0 -- a fallback-recovered tool list proves nothing about runtime behavior. |
| Egress / SSRF Resilience | 1.5/10 | R1: see destructive_operation_safety_score -- same fix applied to the 'no egress flagged' branch. |
| Execution / Sandbox Safety | 1.5/10 | R1: see destructive_operation_safety_score. This is the specific dimension the mid-2026 scoring remediation review named as the clearest floor contributor (scored 4/4 on a server that never completed initialize). |
| Data Exfiltration Resilience | 1.0/10 | R1: see destructive_operation_safety_score -- same fix applied to the 'no bulk/export access flagged' branch. |
| Least Privilege Scope | 1.5/10 | Correction: zero-tool floor already fixed in a prior round (same reasoning as destructive_operation_safety_score) -- an empty tool_inventory returns 1.5 unconditionally, before the non-OAuth branches this entry described are ever reached. This entry previously claimed 0.0, which the code never actually reached for the (common) zero-tool case; corrected to match actual behavior rather than silently changing it. |
| Secret Handling Hygiene | 1.5/10 | R1: see destructive_operation_safety_score -- same fix applied to the 'no secret access flagged' branch. See also R5 for the classifier-accuracy fix on the flip side of this same branch. |
| Supply Chain Signal | 0.0/10 | Justified non-zero in practice (repository/changelog/license/version metadata is independent of live reachability); genuinely 0 only when none of that registry metadata exists. |
| Input Sanitization Safety | 0.0/10 | Zero when no tools were discovered; non-zero for genuinely-discovered tool schemas, same reasoning as tool_surface_design_score. |
| Tool Namespace Clarity | 0.0/10 | Zero when no tool names were discovered; non-zero for genuinely-discovered tool names, same reasoning as tool_surface_design_score. |
Experimental candidate components
Candidate components are reported for empirical study only. They have weight zero and do not affect public scores, verdicts, gates, percentiles, rankings, or release decisions until a matching evaluation report promotes them.
| Candidate | Weight | Evidence |
|---|---|---|
| personal_data_exposure_score | 0 | Export and bulk-access classifications plus schema-declared batch bounds only; no prose inference of personal data. |
See Methodology for what these mechanics mean in plain terms, or Methodology changelog for their history.