An AI agent can now act on a site as you, in your browser. Grenz decides what it may do: the site's own rules, approvals a script can't fake, and every call on the record.
Aggregate38
Leverage10
Execution8.5
Impact9.5
Creativity10
Each criterion 1–10, equally weighted; aggregate is their sum. Ranking is the pipeline's consolidated output.
Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.
02 The two blind reviews
Two independent reviewers scored this project blind, from a sanitized evidence packet.
Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.
Reviewer A (round 1)
confidence 94%
Leverage
10 /10
Policy enforcement, authenticity-bound approval, and refusal of dangerous or injected tool calls change what WebMCP can safely do; this is not a UI convenience layer.
Evidence cited
Transcript demonstrates harmless read allowed, door unlock requiring human approval, self-approval rejected, permanent access denied, and injection ignored.
About text describes coverage of both registration paths, closed-shadow-root approval card, and per-call record.
Video is 167 seconds and directly depicts the policy workflow.
Execution
9 /10
The video transcript and frame sheets show a coherent end-to-end policy scenario with multiple decision outcomes, including the policy-off contrast.
Evidence cited
Submitted video with detailed transcript and alive demo.
Frame sheets show smart-home dashboard, approval modal, and BARGE-like policy states.
Agent actions in a user’s browser can carry real authority; preventing prompt injection and requiring trustworthy human approval addresses a high-stakes, broadly relevant safety problem.
Evidence cited
Concrete front-door, alarm, contractor, and permanent-access scenario.
Audit trail and explicit deny/approval semantics are visible in the described demo.
Creativity
9 /10
The policy-layer framing, site-authored action descriptions, anti-self-approval mechanism, and injection test are an unusually deep response to WebMCP’s trust boundary.
Evidence cited
Three-way Runs freely/Asks me first/Never model.
Approval authenticity and hostile door-log content are pursued as core interaction design.
Reviewer B (round 2)
confidence 95%
Leverage
10 /10
WebMCP is the governed action surface: Grenz makes agent calls policy-checked, state-aware, consent-bound, and auditable in the browser session; ordinary UI driving cannot provide comparable typed interception and authority boundaries.
Evidence cited
Pitch says every call goes through the policy layer and site rules decide Runs freely/Asks me first/Never.
Transcript shows agent approval self-click refused because browser cannot prove a person.
Transcript shows prompt injection in a door log ignored and permanent access denied by an explicit house rule.
Policy-off comparison demonstrates the unguarded door opens.
Execution
9 /10
The 167-second video transcript presents a clear end-to-end series of allowed, approval-required, self-approval-refused, prompt-injection-resistant, explicit-deny, and policy-off states.
Evidence cited
Alive demo, eight gallery images, and 167-second video reported.
Transcript enumerates multiple successful/refused tool outcomes and audit record.
ABOUT text describes closed shadow-root approval card and both registration paths.
Impact
9 /10
Agent authority, prompt injection, and irreversible smart-home actions are a concrete safety problem; enforceable human approval and permanent-deny rules directly reduce risk.
Evidence cited
Front-door unlock and alarm/access scenarios are consequential.
Transcript demonstrates hostile door-log instructions being treated as untrusted words.
Audit trail records reasons for every attempted action.
Creativity
9 /10
The project turns WebMCP's security caveat into a vivid, adversarial, policy-controlled interaction model and demonstrates it with a compelling smart-home narrative.
Evidence cited
Three-level site-authored policy model is legible and memorable.
Human presence/authenticator semantics prevent an agent from approving itself.
Policy-on/off contrast makes the intervention concrete.
03 Review highlights
Standouts across reviewers
Best direct end-to-end WebMCP evidence in the packet.
Safety model is concrete, legible, and central to the product.
Strongest security and least-authority demonstration in the packet.
Direct prompt-injection test with visible refusal.
Red flags
Claims of script-proof approval are demonstrated by transcript but implementation is not independently inspected.
Smart-home environment is a demo rather than a deployed device integration.
04 Evidence
What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs;
video frames are evidence from the submitted demo, not live verification.
Probe captures come from Stage 2 interactive testing of the live product by a reviewer.
EX-01 Devpost page capture · packaging evidence, not runtime proof
· sourceEX-02 Live probe at Stage 2 · observed product behavior, reviewer-drivenEX-03 Live probe after interaction · observed product behavior
EX-V Submitted demo video — “Grenz: a policy layer for WebMCP”
Contact sheets from the 167s video the team submitted. This is what reviewers were shown;
it demonstrates the product in motion but is not independent verification.
· watch the original
EX-V1 Sheet 1 of 3 · reported video evidenceEX-V2 Sheet 2 of 3 · reported video evidenceEX-V3 Sheet 3 of 3 · reported video evidence