Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.
02 The two blind reviews
Two independent reviewers scored this project blind, from a sanitized evidence packet.
Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.
Reviewer A (round 1)
confidence 94%
Leverage
9 /10
Structured access to WebGL viewer state—studies, slices, windows, measurements—makes precise agent assistance possible without granting pixel interpretation; generic UI driving is fragile and unsafe.
Evidence cited
Transcript explicitly shows tool calls moving two studies to slice 80, 100, and 120, with human-in-loop approval.
About text says no tool returns pixels and propose_measurement uses IDs not coordinates.
Frames show paired CT view, slice progression, and conversational panel.
Execution
9 /10
The 179-second transcript describes a complete central workflow with live UI changes, approvals, undo, measurements, and report drafting; evidence is unusually strong.
Evidence cited
Transcript narrates side-by-side prior/current setup, slice synchronization, human approval, undo, measurement propagation, and draft findings.
Frames show radiology viewer states and agent conversation.
Devpost screenshot shows viewer and activity/tool panel.
Impact
10 /10
Radiologists have a specific high-value workflow burden in locating and comparing studies; accelerating navigation and grounding reports while preventing agent diagnosis is credible and safety-conscious.
Evidence cited
Concrete prior-scan comparison, measurement propagation, and evidence-linked report use case.
Safety boundary explicitly prevents the agent from reading or diagnosing scans.
Transcript demonstrates the workflow rather than merely claiming it.
Creativity
9 /10
The deliberate inversion—agent flies the viewer but never sees the scan—creates a novel and memorable human-agent division of labor pursued deeply.
Evidence cited
Measurement chips and signature-bound report export connect every sentence to accepted evidence.
Human measurement ownership is preserved while navigation is automated.
Reviewer B (round 2)
confidence 93%
Leverage
10 /10
WebGL viewer state is difficult to scrape or manipulate reliably; typed tools expose safe, precise actions while explicitly preventing pixel access and diagnosis.
Evidence cited
Transcript shows tool calls moving both studies to requested slices and applying human approval.
About text states no tool returns pixels and measurement proposals use IDs rather than coordinates.
Execution
9 /10
The transcript provides unusually strong end-to-end evidence: natural-language request, live tool calls, synchronized viewer changes, human-in-loop approval, undo, measurements, and grounded report drafting.
Evidence cited
Video transcript explicitly describes real-time tool invocation and synchronized slice changes.
Transcript covers approval toggle, undoability, measurements, and report drafting.
Impact
9 /10
Radiologists face a real, high-friction workflow problem, and the product reduces navigation burden while keeping clinical judgment and pixel interpretation human-controlled.
Evidence cited
About text identifies prior-study arrangement and state management as workflow burdens.
Every report sentence is tied to accepted measurements and export integrity is enforced.
Creativity
9 /10
The safety boundary—agent can operate everything except read the scan—creates a novel and deeply pursued human-agent collaboration model.
Evidence cited
Agent propagates measurements and drafts findings without receiving pixels.
Signature binds report text to measurements and invalidates altered exports.
03 Review highlights
Standouts across reviewers
Exceptional safety/product boundary: agent can operate the viewer but cannot diagnose.
Transcript provides direct end-to-end evidence.
Best packet-level evidence and exceptionally clear safety model.
Red flags
Mock data is used; clinical deployment claims should not be inferred.
04 Evidence
What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs;
video frames are evidence from the submitted demo, not live verification.
Probe captures come from Stage 2 interactive testing of the live product by a reviewer.
EX-01 Devpost page capture · packaging evidence, not runtime proof
· sourceEX-02 Live probe at Stage 2 · observed product behavior, reviewer-drivenEX-03 Live probe after interaction · observed product behavior
EX-V Submitted demo video — “Substrate Demo: OpenAI Webmcp Hackathon”
Contact sheets from the 179s video the team submitted. This is what reviewers were shown;
it demonstrates the product in motion but is not independent verification.
· watch the original
EX-V1 Sheet 1 of 3 · reported video evidenceEX-V2 Sheet 2 of 3 · reported video evidenceEX-V3 Sheet 3 of 3 · reported video evidence