WebMCP analysis · compute ledger
Volume
What it cost to look at everything honestly — receipts included.
01 The scale
Every number on this site came out of a fresh-context AI agent that re-read its instructions, rubric, and evidence from scratch. No reviewer saw another's output. Isolation is expensive; these are the receipts.
02 Where the ranking comes from
Each project's final score rests on the strongest evidence available for it — not the same pipeline for everyone, but the same bar.
03 9,150 judgments
Every judgment was an isolated agent reading its own full packet. This is what "we looked at everything" quantifies to.
04 What the re-score changed
1,353 projects were re-scored after a data bug hid their demo videos from reviewers. Aggregate-score movement: 1,022 rose, 192 fell, 139 held. Corrected evidence helped most projects — and honestly hurt the ones whose videos didn't survive a second look.
Aggregate delta (new − prior), 1,353 re-scored projects. Coral = rose, gray = fell or unchanged.
05 Runtime verification, among the 716 live-tested
06 The incident ledger
Four defects were found and fixed during the run, each with a scoped re-run rather than a silent patch. They're listed in the repo's DEVIATIONS.md with full before/after.
07 Why so many tokens
Blind review is bought with re-read tokens: ~46k true input tokens per subagent, every time, because no reviewer may inherit another's context. Re-runs were scoped to affected subsets — 1,353 of 2,500 re-scored rather than everything. The full ledger, including cache-accounting methodology, lives in the repo as VOLUME.md.