webmcp/analysis

SpendGate

Agents move the work. Policy makes the call. An agent triages 40 expenses in one WebMCP call; the server (not the agent) decides and authorizes every money action.

Aggregate 35
Leverage 10
Execution 8
Impact 9
Creativity 8

Each criterion 1–10, equally weighted; aggregate is their sum. Ranking is the pipeline's consolidated output.

01 Links & metadata

Open the source material

Category
finance / finance
Origin
built for the challenge / built for the challenge
Access
login required
Eligibility
LIKELY_ELIGIBLE / LIKELY_ELIGIBLE
Substitution
MAJOR_DELTA / TRANSFORMATIVE
Demo liveness
unknown

Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.

02 The two blind reviews

Two independent reviewers scored this project blind, from a sanitized evidence packet. Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.

Reviewer A (round 1)

confidence 71%
Leverage
9 /10

WebMCP exposes the real authenticated, role-scoped session while preserving server authority and structured refusal; this is materially safer and more repeatable than an agent clicking through finance UI.

Evidence cited
  • Claims one WebMCP call triages 40 expenses.
  • Tool surface changes by analyst versus manager role.
  • Server, not model, evaluates policy and authorizes actions; poisoned memo remains outside policy evaluation.
Execution
6 /10

The product story is coherent and the frame sheet shows a concrete policy console, but demo_alive is unknown and there is no transcript, so end-to-end behavior is only claimed.

Evidence cited
  • Three frame sheets and a Devpost screenshot show the stated finance workflow and WebMCP contract.
  • Facts report demo_alive=unknown.
  • No repository is available.
Impact
9 /10

Finance teams need safe, auditable automation around high-volume expense queues where model authority and prompt injection are material risks.

Evidence cited
  • Claims explicit caps, receipts, duplicates, role limits, approvals, review, and rejection.
  • The scenario directly tests whether an agent can be allowed near money.
Creativity
8 /10

The poisoned-record demonstration and server-authorized batch triage make a familiar expense console a focused governance experiment.

Evidence cited
  • Adversarial memo is deliberately isolated from policy evaluation.
  • Least privilege and safe escalation are visible product beats.

Reviewer B (round 2)

confidence 78%
Leverage
9 /10

The structured tool surface, role scoping, server policy evaluation, and safe escalation materially change whether an agent can operate near money; generic UI automation cannot provide the same reliable authority boundary.

Evidence cited
  • One WebMCP call triages 40 expenses under server-enforced rules
  • Tool surface changes from analyst to manager role
  • Poisoned memo is treated as untrusted data and over-limit cases are escalated
Execution
8 /10

The packet gives a specific seeded queue, explicit policy outcomes, role behavior, and visible beats, but no video and demo status is unknown.

Evidence cited
  • About text specifies 40 seeded expenses and Approved/Needs review/Rejected outcomes
  • Two gallery images and three listed frame sheets provide visual packaging evidence
  • Description details authenticated session and structured refusal
Impact
8 /10

Finance teams have a real risk-sensitive approval workload, and the product addresses throughput while preserving policy and human escalation.

Evidence cited
  • Specific expense approval queue and category/receipt/duplicate/role rules
  • Focus on safe use of agents near money
Creativity
8 /10

The combination of policy-as-authority, least privilege, adversarial memo handling, and one-call triage is a strong, timely interaction concept.

Evidence cited
  • Server rather than model decides authorization
  • Poisoned record is deliberately included as a governance test

03 Review highlights

Standouts across reviewers

  • Strong server-side authority boundary.
  • Concrete prompt-injection handling in a consequential domain.
  • Excellent safety model for agentic financial operations
  • Concrete adversarial and least-privilege scenarios

Red flags

  • Demo liveness is unknown and there is no transcript.
  • Authenticated flow may limit accessibility for reviewers.
  • Demo status is unknown and no submitted video transcript

04 Evidence

What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs; video frames are evidence from the submitted demo, not live verification.

Devpost page capture for SpendGate
EX-01 Devpost page capture · packaging evidence, not runtime proof · source

EX-V Submitted demo video

Contact sheets from the ?s video the team submitted. This is what reviewers were shown; it demonstrates the product in motion but is not independent verification. · watch the original

Contact sheet 1 from the SpendGate demo video
EX-V1 Sheet 1 of 3 · reported video evidence
Contact sheet 2 from the SpendGate demo video
EX-V2 Sheet 2 of 3 · reported video evidence
Contact sheet 3 from the SpendGate demo video
EX-V3 Sheet 3 of 3 · reported video evidence