webmcp/analysis

Living Evidence

Living Evidence uses WebMCP to turn scientific papers into agent-operable tools, enabling AI to audit provenance, rerun analyses, challenge claims, and propose human-approved updates.

Aggregate 36
Leverage 9.5
Execution 8.5
Impact 9
Creativity 9

Each criterion 1–10, equally weighted; aggregate is their sum. Ranking is the pipeline's consolidated output.

01 Links & metadata

Category
research / research
Origin
built for the challenge / built for the challenge
Access
no auth
Eligibility
LIKELY_ELIGIBLE / LIKELY_ELIGIBLE
Substitution
TRANSFORMATIVE / TRANSFORMATIVE
Demo liveness
alive

Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.

02 The two blind reviews

Two independent reviewers scored this project blind, from a sanitized evidence packet. Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.

Reviewer A (round 1)

confidence 96%
Leverage
10 /10

The paper itself becomes a structured scientific tool surface; this materially changes agent access from reconstructing prose/PDFs to bounded reruns, provenance inspection, and shared rendered results.

Evidence cited
  • Transcript says 15 typed WebMCP tools come directly from the page, not table scraping or a copied notebook.
  • Transcript describes bounded reruns, claim evaluation, rendered forest plots, and SHA-256 input/result receipts.
  • Frames show manifests, forest plots, audit/provenance panels, and evidence maps.
Execution
9 /10

All major surfaces agree and the long demo transcript describes an end-to-end workflow with explicit limitations, human approval, provenance, analysis, and evidence mapping.

Evidence cited
  • 170-second transcript covers inspection, claim challenge, rerun, sensitivity plot, ledger, authoring, approval, and multilingual evidence map.
  • Frames show structured extraction fields, risk-of-bias data, plots, workbench, and audit controls.
  • Devpost screenshot visibly shows source locators, quotes, derivations, estimates, and bias fields.
Impact
9 /10

Researchers, reviewers, and evidence synthesists have a serious problem with unverifiable claims and expensive reproduction; this directly lowers inspection and challenge cost while preserving human governance.

Evidence cited
  • About text identifies AI reconstruction errors and costly verification as the problem.
  • Demo handles a historical meta-analysis with 19 records across 18 experiments and explicit provenance/risk-of-bias tracking.
Creativity
9 /10

The concept of an executable, cross-examinable paper with bounded agent operations and human-governed evidence admission is unusually original and pursued deeply.

Evidence cited
  • Scientific document is treated as an agent-operable tool surface rather than static prose.
  • Evidence board generalizes beyond meta-analysis while preserving original-language quotes.

Reviewer B (round 2)

confidence 94%
Leverage
10 /10

The document itself becomes a typed scientific tool surface; reruns, sensitivity questions, provenance, audit receipts, and shared rendered results are impractical to reproduce comparably through generic page driving.

Evidence cited
  • Video transcript says the page registers 15 typed WebMCP tools.
  • Tools perform deterministic analysis, reruns, sensitivity checks, claim rules, provenance, audit, and receipts.
  • SHA-256 ledger records exact inputs and result digest; agents propose while humans approve.
Execution
9 /10

This is strongly evidenced by a 170-second product demo transcript, alive demo, extensive gallery, and detailed coherent scope; minor uncertainty remains because visual frames were not inspectable.

Evidence cited
  • Video transcript describes manifest inspection, failed claim rule, sensitivity rerun, forest plot, ledger, and approval workflow.
  • Demo alive, public repository, 13 gallery images, and 170-second video reported.
  • Scope includes meta-analysis, evidence board, provenance gaps, risk of bias, and multilingual source display.
Impact
10 /10

Researchers and evidence users gain a credible way to inspect, reproduce, challenge, and update claims while explicitly avoiding false certainty.

Evidence cited
  • Historical Pygmalion meta-analysis is reconstructed from 19 records/18 experiments.
  • Transcript shows a headline claim failing its registered rule, bounded rerun, explicit provenance/risk gaps, and human-approved extraction.
Creativity
10 /10

Turning a scientific paper into an agent-operable, human-governed executable evidence surface is genuinely novel and pursued deeply.

Evidence cited
  • Document-native typed operations go beyond chat summarization or copied notebooks.
  • Claim bookkeeping, provenance, sensitivity composition, receipts, and multilingual evidence are integrated into the concept.

03 Review highlights

Standouts across reviewers

  • Strong evidence honesty: reproducibility is explicitly separated from truth.
  • Human approval cannot be bypassed by agent tools.
  • Best-in-packet WebMCP indispensability.
  • Explicitly refuses to equate reproducibility with truth.

Red flags

  • Benchmark section explicitly has zero runs, so no speed or accuracy claim is evidenced.
  • The benchmark explicitly starts at zero runs and makes no speed/accuracy claim; this is honest but limits performance conclusions.

04 Evidence

What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs; video frames are evidence from the submitted demo, not live verification. Probe captures come from Stage 2 interactive testing of the live product by a reviewer.

Devpost page capture for Living Evidence
EX-01 Devpost page capture · packaging evidence, not runtime proof · source
Live probe of Living Evidence before interaction
EX-02 Live probe at Stage 2 · observed product behavior, reviewer-driven
Live probe of Living Evidence after interaction
EX-03 Live probe after interaction · observed product behavior

EX-V Submitted demo video — “Living Evidence Demo — Scientific Documents Agents Can Cross-Examine”

Contact sheets from the 170s video the team submitted. This is what reviewers were shown; it demonstrates the product in motion but is not independent verification. · watch the original

Contact sheet 1 from the Living Evidence demo video
EX-V1 Sheet 1 of 3 · reported video evidence
Contact sheet 2 from the Living Evidence demo video
EX-V2 Sheet 2 of 3 · reported video evidence
Contact sheet 3 from the Living Evidence demo video
EX-V3 Sheet 3 of 3 · reported video evidence