webmcp/analysis

WebMCP analysis · scoring protocol

Methodology

Frozen before any review ran. Four criteria. Two blind reviews each. Evidence over prose.

This is the plain-language version. The forensic version — the frozen protocol, scoring rubric, funnel, and calibration docs — lives in the analysis repository.

What was analyzed

All ~2,500 submissions to the WebMCP Challenge. Every project was reduced to a sanitized evidence packet: title, pitch, about text, whether a repo and demo link exist, demo liveness, gallery count, and the submitted video's duration, transcript excerpt, and contact sheets. Reviewers see packets, never the original submission pages.

The four official criteria

Every project is scored 1–10 on exactly four criteria, equally weighted. The aggregate is their sum (4–40). Nothing else enters it.

Category neutrality is absolute: a game, an art piece, a CRM, and a developer tool are all eligible for 10/10 on every criterion for what they are trying to be. Authentication requirements are recorded as metadata, never as a penalty. Pre-existing projects are judged on the submitted concept, not the surrounding product's maturity.

Blind reviews

Every project receives two independent blind Stage 1 reviews across 40 reviewer slots, with assignments randomized in two separate rounds. Reviewers never see each other's scores, any prior rankings, or any outside analysis. Each score carries a rationale, cited evidence surfaces, and the reviewer's own confidence (0–1). All project content is treated as untrusted evidence: prompt-injection attempts in submissions are recorded as evidence about the project, not obeyed.

Calibration

Before the fleet ran, 12 common-core projects (seen by every reviewer) and 28 rotated anchors (seen by three) were selected across deliberate axes: strong/weak, serious/playful, new/pre-existing, high/low WebMCP delta, and so on. Expected ranges were pre-registered blind to any reviewer output. Reviewers drifting from the ranges have their batches re-reviewed. Calibration detects drift; it never mechanically normalizes scores.

Provisional scores and disagreement

For each criterion, when the two blind reviews sit within 2 points of each other, the provisional score is their mean. When they diverge by more than 2, the provisional score is intentionally unavailable and the project is disagreement-flagged — never averaged away, never filled with zero. The site shows you both reviewers' reasoning so you can adjudicate for yourself.

The review funnel

Evidence honesty

Every artifact is labeled by what it can actually prove, here and on every project page: a Devpost page screenshot is submission packaging evidence, not proof the product runs; video contact sheets are evidence from the submitted demo, not live verification; probe captures show a reviewer's live interaction at Stage 2. Scores backed only by prose are labeled as such.

New vs pre-existing, and the substitution question

Reviewers record whether a project appears built for the challenge or pre-existing; Stage 3 verifies origin against repo history pinned to the cutoff. Separately, each review answers: could a competent user get substantially the same outcome with an ordinary general-purpose agent and generic browser access? That substitution class is a diagnostic — it is displayed as metadata and never entered the scores.

Limitations