Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.
02 The two blind reviews
Two independent reviewers scored this project blind, from a sanitized evidence packet.
Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.
Reviewer A (round 1)
confidence 89%
Leverage
9 /10
The live shared backlog, participant-specific capability constraints, structured triage, and risk-gated draft/commit boundary make agent coordination materially safer and more reliable than generic UI automation.
Evidence cited
Six named WebMCP tools cover state, needs, clarification, offers, drafts, and review blocks.
Server compiles needs into Routine/Review/Human-only levels for each participant.
Agents coordinate but humans commit, with batch confirmation restricted by risk level.
Execution
8 /10
The scenario, participant profiles, risk policy, and frame sheets indicate a serious coherent workflow; actual external crisis operation is appropriately not claimed.
Evidence cited
Alive demo and public repository.
Three frame sheets show a crisis-board UI paired with agent workflow evidence.
Detailed synthetic flood-response scenario and explicit safety boundaries.
Impact
9 /10
Crisis responders and volunteers face a real coordination bottleneck, and profile-aware triage with human commitment could provide high value even in a synthetic demonstration.
Evidence cited
Concrete flood-response coordination audience and needs.
Profiles encode transport, range, skills, availability, and exclusions.
Risk levels distinguish logistics from hazardous, clinical, or safeguarding actions.
Creativity
9 /10
Applying personal browser agents to distributed emergency coordination while enforcing per-action human authority is ambitious, memorable, and deeply pursued.
Evidence cited
Agent-per-volunteer model rather than one central bot.
Safety policy is integrated into the interaction model.
Reviewer B (round 2)
confidence 70%
Leverage
8 /10
Structured coordination state and bounded draft/review tools make agent triage safer and more repeatable than UI driving; human commitment remains explicit.
Evidence cited
Six named WebMCP tools include get_coordination_state, read_need, offer_resource, draft_commitment, and get_review_block.
Profile exclusions and L0/L1/L2 review gates are compiled server-side.
Execution
7 /10
The packet presents a coherent end-to-end crisis workflow and a visible review-oriented demo surface, though the supplied packet has no actual video evidence despite a referenced frame sheet.
Evidence cited
Demo link is alive and public repository is reported.
Devpost screenshot visibly shows matched response drafts and human review controls.
Impact
9 /10
Crisis volunteers and local responders have a concrete, high-stakes coordination problem; the human-in-the-loop design credibly addresses triage overload and unsafe commitments.
Evidence cited
Scenario covers meals, medicine, roads, rumors, and volunteer capability constraints.
Human-only handling is specified for clinical, evacuation, safeguarding, and hazardous cases.
Creativity
8 /10
The agent-coordinates/human-commits framing is a strong application of WebMCP to safety-critical mutual aid, with meaningful policy states rather than a generic chatbot.
Evidence cited
Participant-owned agents triage a shared board while only humans commit.
Three-level attention and commitment policy is part of the interaction model.
03 Review highlights
Standouts across reviewers
Excellent agent/human authority boundary.
Participant-specific risk compilation.
Strong safety boundary between agent drafts and human commitments.
Server-side capability and risk classification.
Red flags
Scenario and data are explicitly synthetic and not affiliated with authorities or NGOs.
Frames cannot prove real-world deployment or safety outcomes.
No transcript and no submitted video flag, while frame sheets are referenced; end-to-end execution evidence is limited.
Scenario is explicitly fictionalized.
04 Evidence
What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs;
video frames are evidence from the submitted demo, not live verification.
Probe captures come from Stage 2 interactive testing of the live product by a reviewer.
EX-01 Devpost page capture · packaging evidence, not runtime proof
· sourceEX-02 Live probe at Stage 2 · observed product behavior, reviewer-drivenEX-03 Live probe after interaction · observed product behavior
EX-V Submitted demo video
Contact sheets from the ?s video the team submitted. This is what reviewers were shown;
it demonstrates the product in motion but is not independent verification.
· watch the original
EX-V1 Sheet 1 of 3 · reported video evidenceEX-V2 Sheet 2 of 3 · reported video evidenceEX-V3 Sheet 3 of 3 · reported video evidence