WebMCP analysis · scoring protocol
Methodology
Frozen before any review ran. Four criteria. Two blind reviews each. Evidence over prose.
This is the plain-language version. The forensic version — the frozen protocol, scoring rubric, funnel, and calibration docs — lives in the analysis repository.
What was analyzed
All ~2,500 submissions to the WebMCP Challenge. Every project was reduced to a sanitized evidence packet: title, pitch, about text, whether a repo and demo link exist, demo liveness, gallery count, and the submitted video's duration, transcript excerpt, and contact sheets. Reviewers see packets, never the original submission pages.
The four official criteria
Every project is scored 1–10 on exactly four criteria, equally weighted. The aggregate is their sum (4–40). Nothing else enters it.
- WebMCP Leverage — what does WebMCP change for this product? Judged on the delta to reliability, precision, shared state, and repeatability. Tool count is not leverage.
- Execution — is the actual product coherent, intentional, and complete for its scope? Thin evidence caps the score: end-to-end credit requires end-to-end evidence.
- Potential Impact — credible value for the intended audience. Entertainment, art, play, and niche use count fully. Audience size is not a multiplier.
- Creativity & Ambition — novelty of concept and interaction model, and how deeply it is pursued.
Category neutrality is absolute: a game, an art piece, a CRM, and a developer tool are all eligible for 10/10 on every criterion for what they are trying to be. Authentication requirements are recorded as metadata, never as a penalty. Pre-existing projects are judged on the submitted concept, not the surrounding product's maturity.
Blind reviews
Every project receives two independent blind Stage 1 reviews across 40 reviewer slots, with assignments randomized in two separate rounds. Reviewers never see each other's scores, any prior rankings, or any outside analysis. Each score carries a rationale, cited evidence surfaces, and the reviewer's own confidence (0–1). All project content is treated as untrusted evidence: prompt-injection attempts in submissions are recorded as evidence about the project, not obeyed.
Calibration
Before the fleet ran, 12 common-core projects (seen by every reviewer) and 28 rotated anchors (seen by three) were selected across deliberate axes: strong/weak, serious/playful, new/pre-existing, high/low WebMCP delta, and so on. Expected ranges were pre-registered blind to any reviewer output. Reviewers drifting from the ranges have their batches re-reviewed. Calibration detects drift; it never mechanically normalizes scores.
Provisional scores and disagreement
For each criterion, when the two blind reviews sit within 2 points of each other, the provisional score is their mean. When they diverge by more than 2, the provisional score is intentionally unavailable and the project is disagreement-flagged — never averaged away, never filled with zero. The site shows you both reviewers' reasoning so you can adjudicate for yourself.
The review funnel
- Stage 1 — broad review. All 2,500 projects, two blind reviews each, on packet evidence only.
- Stage 2 — deep dive. Pre-registered lanes (top aggregate, category leaders, top per-criterion, disagreement, rescue lanes, a random control group) advance to interactive testing: a reviewer drives the product's central journey, records normalized observations, then two new blind scorers re-score all four criteria from the expanded packet. A deterministic accessibility probe is triage only — it never directly sets Execution, and auth itself is not a penalty.
- Stage 3 — finalist deep dive. Full video review, repository history pinned to the submission cutoff for origin verification, and novelty analysis. Runtime WebMCP verification is recorded as evidence (VERIFIED_RUNTIME / VIDEO_VERIFIED / REPO_VERIFIED / CLAIM_ONLY / UNVERIFIED / FAILED). It raises evidence confidence, never quality points, and never advances a project by itself.
- Stage 4 — consolidation. The latest evidence-backed score per criterion wins (adjudication replaces, never averages across stages). Aggregate = sum of four. Rank descending, ties broken Leverage → Execution → Impact → Creativity → confidence → slug.
Evidence honesty
Every artifact is labeled by what it can actually prove, here and on every project page: a Devpost page screenshot is submission packaging evidence, not proof the product runs; video contact sheets are evidence from the submitted demo, not live verification; probe captures show a reviewer's live interaction at Stage 2. Scores backed only by prose are labeled as such.
New vs pre-existing, and the substitution question
Reviewers record whether a project appears built for the challenge or pre-existing; Stage 3 verifies origin against repo history pinned to the cutoff. Separately, each review answers: could a competent user get substantially the same outcome with an ordinary general-purpose agent and generic browser access? That substitution class is a diagnostic — it is displayed as metadata and never entered the scores.
Limitations
- AI reviewers under a frozen protocol, not human judges; systematic blind spots are possible. Rationales are published so you can check the reasoning.
- Evidence is uneven by nature: some submissions ship videos and live products, others only prose. Scores say what the evidence supports — no more.
- Qualitative criteria are judgment calls; the two-reviewer design and published disagreement exist to expose rather than hide that.
- An earlier description-first community analysis is quarantined until Stage 4 and does not influence anything shown here.