webmcp/analysis

Relay

Crisis coordination where every volunteer's own AI agent triages — and only the human commits.

Aggregate 34
Leverage 9
Execution 8
Impact 9
Creativity 8

Each criterion 1–10, equally weighted; aggregate is their sum. Ranking is the pipeline's consolidated output.

01 Links & metadata

Category
social / productivity
Origin
built for the challenge / built for the challenge
Access
login required
Eligibility
LIKELY_ELIGIBLE / LIKELY_ELIGIBLE
Substitution
TRANSFORMATIVE / MAJOR_DELTA
Demo liveness
alive

Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.

02 The two blind reviews

Two independent reviewers scored this project blind, from a sanitized evidence packet. Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.

Reviewer A (round 1)

confidence 89%
Leverage
9 /10

The live shared backlog, participant-specific capability constraints, structured triage, and risk-gated draft/commit boundary make agent coordination materially safer and more reliable than generic UI automation.

Evidence cited
  • Six named WebMCP tools cover state, needs, clarification, offers, drafts, and review blocks.
  • Server compiles needs into Routine/Review/Human-only levels for each participant.
  • Agents coordinate but humans commit, with batch confirmation restricted by risk level.
Execution
8 /10

The scenario, participant profiles, risk policy, and frame sheets indicate a serious coherent workflow; actual external crisis operation is appropriately not claimed.

Evidence cited
  • Alive demo and public repository.
  • Three frame sheets show a crisis-board UI paired with agent workflow evidence.
  • Detailed synthetic flood-response scenario and explicit safety boundaries.
Impact
9 /10

Crisis responders and volunteers face a real coordination bottleneck, and profile-aware triage with human commitment could provide high value even in a synthetic demonstration.

Evidence cited
  • Concrete flood-response coordination audience and needs.
  • Profiles encode transport, range, skills, availability, and exclusions.
  • Risk levels distinguish logistics from hazardous, clinical, or safeguarding actions.
Creativity
9 /10

Applying personal browser agents to distributed emergency coordination while enforcing per-action human authority is ambitious, memorable, and deeply pursued.

Evidence cited
  • Agent-per-volunteer model rather than one central bot.
  • Safety policy is integrated into the interaction model.

Reviewer B (round 2)

confidence 70%
Leverage
8 /10

Structured coordination state and bounded draft/review tools make agent triage safer and more repeatable than UI driving; human commitment remains explicit.

Evidence cited
  • Six named WebMCP tools include get_coordination_state, read_need, offer_resource, draft_commitment, and get_review_block.
  • Profile exclusions and L0/L1/L2 review gates are compiled server-side.
Execution
7 /10

The packet presents a coherent end-to-end crisis workflow and a visible review-oriented demo surface, though the supplied packet has no actual video evidence despite a referenced frame sheet.

Evidence cited
  • Demo link is alive and public repository is reported.
  • Devpost screenshot visibly shows matched response drafts and human review controls.
Impact
9 /10

Crisis volunteers and local responders have a concrete, high-stakes coordination problem; the human-in-the-loop design credibly addresses triage overload and unsafe commitments.

Evidence cited
  • Scenario covers meals, medicine, roads, rumors, and volunteer capability constraints.
  • Human-only handling is specified for clinical, evacuation, safeguarding, and hazardous cases.
Creativity
8 /10

The agent-coordinates/human-commits framing is a strong application of WebMCP to safety-critical mutual aid, with meaningful policy states rather than a generic chatbot.

Evidence cited
  • Participant-owned agents triage a shared board while only humans commit.
  • Three-level attention and commitment policy is part of the interaction model.

03 Review highlights

Standouts across reviewers

  • Excellent agent/human authority boundary.
  • Participant-specific risk compilation.
  • Strong safety boundary between agent drafts and human commitments.
  • Server-side capability and risk classification.

Red flags

  • Scenario and data are explicitly synthetic and not affiliated with authorities or NGOs.
  • Frames cannot prove real-world deployment or safety outcomes.
  • No transcript and no submitted video flag, while frame sheets are referenced; end-to-end execution evidence is limited.
  • Scenario is explicitly fictionalized.

04 Evidence

What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs; video frames are evidence from the submitted demo, not live verification. Probe captures come from Stage 2 interactive testing of the live product by a reviewer.

Devpost page capture for Relay
EX-01 Devpost page capture · packaging evidence, not runtime proof · source
Live probe of Relay before interaction
EX-02 Live probe at Stage 2 · observed product behavior, reviewer-driven
Live probe of Relay after interaction
EX-03 Live probe after interaction · observed product behavior

EX-V Submitted demo video

Contact sheets from the ?s video the team submitted. This is what reviewers were shown; it demonstrates the product in motion but is not independent verification. · watch the original

Contact sheet 1 from the Relay demo video
EX-V1 Sheet 1 of 3 · reported video evidence
Contact sheet 2 from the Relay demo video
EX-V2 Sheet 2 of 3 · reported video evidence
Contact sheet 3 from the Relay demo video
EX-V3 Sheet 3 of 3 · reported video evidence