A WebMCP safety simulator where a browser agent investigates production incidents, but the dangerous tool to act is never registered until a human approves inside the page.
Aggregate38
Leverage10
Execution9
Impact9
Creativity10
Each criterion 1–10, equally weighted; aggregate is their sum. Ranking is the pipeline's consolidated output.
Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.
02 The two blind reviews
Two independent reviewers scored this project blind, from a sanitized evidence packet.
Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.
Reviewer A (round 1)
confidence 85%
Leverage
9 /10
The product demonstrates WebMCP-native safety through conditional tool registration and live schema narrowing; general UI automation cannot reproduce these boundaries as cleanly or reliably.
Evidence cited
Dangerous action tool is never registered until human approval.
Revoking a service live-narrows the WebMCP tool schema, not only the UI.
Blocked calls, approvals, and execute_approved_action are part of the shared workflow.
Execution
8 /10
The packet specifies a complete simulator workflow with visible state, blocked attempts, approvals, recovery, and scorecard, supported by a live demo, repo, and frame sheets; no video limits direct proof.
Evidence cited
Live demo and public repository are reported alive.
Claimed scorecard checks root cause, mitigation, and prompt-injection resistance.
Description reports visible evidence, decision context, activity log, blocked banner, approvals, and service recovery.
Impact
9 /10
Production incidents and prompt injection are real high-stakes problems for teams adopting agents in operational systems; a safe simulator offers credible training and design value.
Evidence cited
Scenario models a checkout outage and production-changing mitigation.
Prompt-injection note is explicitly tested and not obeyed.
Human approval, written reasons, revocation, and auditability address operational risk.
Creativity
8 /10
The combination of incident response simulation, dynamic tool surface, prompt-injection scoring, and trusted multi-approval is a notably strong WebMCP interaction model.
Evidence cited
Dangerous capability absence, rather than a disabled button, is used as the safety primitive.
Scorecard makes safety behavior an observable game-like evaluation.
Decision context and tool activity are designed for joint human-agent reasoning.
Reviewer B (round 2)
confidence 78%
Leverage
9 /10
The safety behavior is structurally tied to WebMCP registration and schema visibility: dangerous tools are not registered until approval, blocked calls are surfaced, and revocation narrows the live tool surface. A generic UI agent would have a harder time receiving the same reliable capability boundaries.
Evidence cited
Pitch states the dangerous tool is never registered until human approval.
About text describes two trusted approvals before execute_approved_action and live-narrowing of the WebMCP schema.
Prompt-injection note and blocked tool call are visible in the described activity feed.
Execution
8 /10
The packet describes a complete simulator loop from triage through evidence, approvals, mitigation, recovery, and scorecard, with extensive visible state and logging; no video limits direct confirmation.
Evidence cited
About text names triage, scorecard, decision context, activity log, blocked banner, scope revocation, and approval states.
Frame sheets were submitted for the product.
The claimed workflow includes successful recovery after approvals.
Impact
8 /10
Teams connecting agents to operational systems face a real safety and prompt-injection problem; a simulator that makes approval and failure modes concrete has credible training and design value.
Evidence cited
Description targets production incidents, deploy consoles, dashboards, and status pages.
The system tests correct root cause, mitigation, and non-obedience to injected instructions.
Creativity
8 /10
It turns agent safety principles into a playable operational incident workflow with dynamic tool availability, trusted approvals, and an explicit scorecard.
Evidence cited
Dangerous capability is absent from the registered tool surface until approval.
Prompt injection is tested as an in-simulator event rather than only discussed.
03 Review highlights
Standouts across reviewers
Schema-level gating and live narrowing are stronger than UI-only safety theater.
Prompt-injection resistance is tested as part of the scorecard.
Dynamic registration and schema narrowing make the safety model concrete.
Clear training/simulation framing for high-risk agent operations.
Red flags
No video/transcript; agent behavior and recovery are described but not directly shown in motion.
No video or transcript was submitted.
Safety and recovery claims are based on packet description and frames rather than direct run evidence.
04 Evidence
What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs;
video frames are evidence from the submitted demo, not live verification.
Contact sheets from the ?s video the team submitted. This is what reviewers were shown;
it demonstrates the product in motion but is not independent verification.
· watch the original
EX-V1 Sheet 1 of 3 · reported video evidenceEX-V2 Sheet 2 of 3 · reported video evidenceEX-V3 Sheet 3 of 3 · reported video evidence