webmcp/analysis

WebMCP analysis · compute ledger

Volume

What it cost to look at everything honestly — receipts included.

01 The scale

Every number on this site came out of a fresh-context AI agent that re-read its instructions, rubric, and evidence from scratch. No reviewer saw another's output. Isolation is expensive; these are the receipts.

Projects analyzed
2,500
Independent judgments
9,150
Subagent reviewers
1,032
Cumulative agent time
80.2h
API calls
10,193
Tokens (billable-shape)
59.3M
Wall-clock compression
~20×

02 Where the ranking comes from

Each project's final score rests on the strongest evidence available for it — not the same pipeline for everyone, but the same bar.

Full corpus 2,500 every public submission
S2 · live-product tested 716 observer drove the real product; 217 runtime-verified
S1R · re-scored 1,353 video-evidence defect corrected, single fresh reviewer each
S1 · as first reviewed 431 two blind reviewers, no defect found

03 9,150 judgments

Every judgment was an isolated agent reading its own full packet. This is what "we looked at everything" quantifies to.

S1 round 1 blind reviews
2,824
S1 round 2 blind reviews
2,824
S2 live observations
717
S2 blind re-scores
1,432
S1R remediation re-scores
1,353

04 What the re-score changed

1,353 projects were re-scored after a data bug hid their demo videos from reviewers. Aggregate-score movement: 1,022 rose, 192 fell, 139 held. Corrected evidence helped most projects — and honestly hurt the ones whose videos didn't survive a second look.

-4..
-3
-2
-1
0
+1
+2
+3
+4
+5..

Aggregate delta (new − prior), 1,353 re-scored projects. Coral = rose, gray = fell or unchanged.

05 Runtime verification, among the 716 live-tested

Runtime verified
217
Video verified
69
Repo verified
15
Claim only
236
Unverified
72
Failed at test time
107

06 The incident ledger

Four defects were found and fixed during the run, each with a scoped re-run rather than a silent patch. They're listed in the repo's DEVIATIONS.md with full before/after.

S2 adjudication drop-out 8 projects restored via mean-capped adjudication
Video-metadata gap (YouTube IP wall) 1,353 projects re-scored with corrected evidence (S1R)
Music-audio modality penalty audio-neutral re-scores + published sensitivity table
Ranking sort defect tracked builder with fail-closed monotonic assertions

07 Why so many tokens

Blind review is bought with re-read tokens: ~46k true input tokens per subagent, every time, because no reviewer may inherit another's context. Re-runs were scoped to affected subsets — 1,353 of 2,500 re-scored rather than everything. The full ledger, including cache-accounting methodology, lives in the repo as VOLUME.md.