webmcp/analysis

Substrate

A radiology viewer built so an AI agent can do everything except read the scan.

Aggregate 37.5
Leverage 10
Execution 7.5
Impact 10
Creativity 10

Each criterion 1–10, equally weighted; aggregate is their sum. Ranking is the pipeline's consolidated output.

01 Links & metadata

Category
health / health
Origin
built for the challenge / built for the challenge
Access
no auth
Eligibility
LIKELY_ELIGIBLE / LIKELY_ELIGIBLE
Substitution
TRANSFORMATIVE / TRANSFORMATIVE
Demo liveness
alive

Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.

02 The two blind reviews

Two independent reviewers scored this project blind, from a sanitized evidence packet. Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.

Reviewer A (round 1)

confidence 94%
Leverage
9 /10

Structured access to WebGL viewer state—studies, slices, windows, measurements—makes precise agent assistance possible without granting pixel interpretation; generic UI driving is fragile and unsafe.

Evidence cited
  • Transcript explicitly shows tool calls moving two studies to slice 80, 100, and 120, with human-in-loop approval.
  • About text says no tool returns pixels and propose_measurement uses IDs not coordinates.
  • Frames show paired CT view, slice progression, and conversational panel.
Execution
9 /10

The 179-second transcript describes a complete central workflow with live UI changes, approvals, undo, measurements, and report drafting; evidence is unusually strong.

Evidence cited
  • Transcript narrates side-by-side prior/current setup, slice synchronization, human approval, undo, measurement propagation, and draft findings.
  • Frames show radiology viewer states and agent conversation.
  • Devpost screenshot shows viewer and activity/tool panel.
Impact
10 /10

Radiologists have a specific high-value workflow burden in locating and comparing studies; accelerating navigation and grounding reports while preventing agent diagnosis is credible and safety-conscious.

Evidence cited
  • Concrete prior-scan comparison, measurement propagation, and evidence-linked report use case.
  • Safety boundary explicitly prevents the agent from reading or diagnosing scans.
  • Transcript demonstrates the workflow rather than merely claiming it.
Creativity
9 /10

The deliberate inversion—agent flies the viewer but never sees the scan—creates a novel and memorable human-agent division of labor pursued deeply.

Evidence cited
  • Measurement chips and signature-bound report export connect every sentence to accepted evidence.
  • Human measurement ownership is preserved while navigation is automated.

Reviewer B (round 2)

confidence 93%
Leverage
10 /10

WebGL viewer state is difficult to scrape or manipulate reliably; typed tools expose safe, precise actions while explicitly preventing pixel access and diagnosis.

Evidence cited
  • Transcript shows tool calls moving both studies to requested slices and applying human approval.
  • About text states no tool returns pixels and measurement proposals use IDs rather than coordinates.
Execution
9 /10

The transcript provides unusually strong end-to-end evidence: natural-language request, live tool calls, synchronized viewer changes, human-in-loop approval, undo, measurements, and grounded report drafting.

Evidence cited
  • Video transcript explicitly describes real-time tool invocation and synchronized slice changes.
  • Transcript covers approval toggle, undoability, measurements, and report drafting.
Impact
9 /10

Radiologists face a real, high-friction workflow problem, and the product reduces navigation burden while keeping clinical judgment and pixel interpretation human-controlled.

Evidence cited
  • About text identifies prior-study arrangement and state management as workflow burdens.
  • Every report sentence is tied to accepted measurements and export integrity is enforced.
Creativity
9 /10

The safety boundary—agent can operate everything except read the scan—creates a novel and deeply pursued human-agent collaboration model.

Evidence cited
  • Agent propagates measurements and drafts findings without receiving pixels.
  • Signature binds report text to measurements and invalidates altered exports.

03 Review highlights

Standouts across reviewers

  • Exceptional safety/product boundary: agent can operate the viewer but cannot diagnose.
  • Transcript provides direct end-to-end evidence.
  • Best packet-level evidence and exceptionally clear safety model.

Red flags

  • Mock data is used; clinical deployment claims should not be inferred.

04 Evidence

What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs; video frames are evidence from the submitted demo, not live verification. Probe captures come from Stage 2 interactive testing of the live product by a reviewer.

Devpost page capture for Substrate
EX-01 Devpost page capture · packaging evidence, not runtime proof · source
Live probe of Substrate before interaction
EX-02 Live probe at Stage 2 · observed product behavior, reviewer-driven
Live probe of Substrate after interaction
EX-03 Live probe after interaction · observed product behavior

EX-V Submitted demo video — “Substrate Demo: OpenAI Webmcp Hackathon”

Contact sheets from the 179s video the team submitted. This is what reviewers were shown; it demonstrates the product in motion but is not independent verification. · watch the original

Contact sheet 1 from the Substrate demo video
EX-V1 Sheet 1 of 3 · reported video evidence
Contact sheet 2 from the Substrate demo video
EX-V2 Sheet 2 of 3 · reported video evidence
Contact sheet 3 from the Substrate demo video
EX-V3 Sheet 3 of 3 · reported video evidence