Living Evidence uses WebMCP to turn scientific papers into agent-operable tools, enabling AI to audit provenance, rerun analyses, challenge claims, and propose human-approved updates.
Aggregate36
Leverage9.5
Execution8.5
Impact9
Creativity9
Each criterion 1–10, equally weighted; aggregate is their sum. Ranking is the pipeline's consolidated output.
Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.
02 The two blind reviews
Two independent reviewers scored this project blind, from a sanitized evidence packet.
Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.
Reviewer A (round 1)
confidence 96%
Leverage
10 /10
The paper itself becomes a structured scientific tool surface; this materially changes agent access from reconstructing prose/PDFs to bounded reruns, provenance inspection, and shared rendered results.
Evidence cited
Transcript says 15 typed WebMCP tools come directly from the page, not table scraping or a copied notebook.
Frames show manifests, forest plots, audit/provenance panels, and evidence maps.
Execution
9 /10
All major surfaces agree and the long demo transcript describes an end-to-end workflow with explicit limitations, human approval, provenance, analysis, and evidence mapping.
Researchers, reviewers, and evidence synthesists have a serious problem with unverifiable claims and expensive reproduction; this directly lowers inspection and challenge cost while preserving human governance.
Evidence cited
About text identifies AI reconstruction errors and costly verification as the problem.
Demo handles a historical meta-analysis with 19 records across 18 experiments and explicit provenance/risk-of-bias tracking.
Creativity
9 /10
The concept of an executable, cross-examinable paper with bounded agent operations and human-governed evidence admission is unusually original and pursued deeply.
Evidence cited
Scientific document is treated as an agent-operable tool surface rather than static prose.
Evidence board generalizes beyond meta-analysis while preserving original-language quotes.
Reviewer B (round 2)
confidence 94%
Leverage
10 /10
The document itself becomes a typed scientific tool surface; reruns, sensitivity questions, provenance, audit receipts, and shared rendered results are impractical to reproduce comparably through generic page driving.
Evidence cited
Video transcript says the page registers 15 typed WebMCP tools.
SHA-256 ledger records exact inputs and result digest; agents propose while humans approve.
Execution
9 /10
This is strongly evidenced by a 170-second product demo transcript, alive demo, extensive gallery, and detailed coherent scope; minor uncertainty remains because visual frames were not inspectable.
Evidence cited
Video transcript describes manifest inspection, failed claim rule, sensitivity rerun, forest plot, ledger, and approval workflow.
Demo alive, public repository, 13 gallery images, and 170-second video reported.
Scope includes meta-analysis, evidence board, provenance gaps, risk of bias, and multilingual source display.
Impact
10 /10
Researchers and evidence users gain a credible way to inspect, reproduce, challenge, and update claims while explicitly avoiding false certainty.
Evidence cited
Historical Pygmalion meta-analysis is reconstructed from 19 records/18 experiments.
Transcript shows a headline claim failing its registered rule, bounded rerun, explicit provenance/risk gaps, and human-approved extraction.
Creativity
10 /10
Turning a scientific paper into an agent-operable, human-governed executable evidence surface is genuinely novel and pursued deeply.
Evidence cited
Document-native typed operations go beyond chat summarization or copied notebooks.
Claim bookkeeping, provenance, sensitivity composition, receipts, and multilingual evidence are integrated into the concept.
03 Review highlights
Standouts across reviewers
Strong evidence honesty: reproducibility is explicitly separated from truth.
Human approval cannot be bypassed by agent tools.
Best-in-packet WebMCP indispensability.
Explicitly refuses to equate reproducibility with truth.
Red flags
Benchmark section explicitly has zero runs, so no speed or accuracy claim is evidenced.
The benchmark explicitly starts at zero runs and makes no speed/accuracy claim; this is honest but limits performance conclusions.
04 Evidence
What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs;
video frames are evidence from the submitted demo, not live verification.
Probe captures come from Stage 2 interactive testing of the live product by a reviewer.
EX-01 Devpost page capture · packaging evidence, not runtime proof
· sourceEX-02 Live probe at Stage 2 · observed product behavior, reviewer-drivenEX-03 Live probe after interaction · observed product behavior
EX-V Submitted demo video — “Living Evidence Demo — Scientific Documents Agents Can Cross-Examine”
Contact sheets from the 170s video the team submitted. This is what reviewers were shown;
it demonstrates the product in motion but is not independent verification.
· watch the original
EX-V1 Sheet 1 of 3 · reported video evidenceEX-V2 Sheet 2 of 3 · reported video evidenceEX-V3 Sheet 3 of 3 · reported video evidence