Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.
02 The two blind reviews
Two independent reviewers scored this project blind, from a sanitized evidence packet.
Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.
Reviewer A (round 1)
confidence 93%
Leverage
10 /10
The agent and human share a deterministic live case, while typed tools add dimensions, investigate, change views, and expose state; structurally absent approval keeps the human boundary enforceable.
Evidence cited
Transcript records tool-driven selection, adding an unmodeled seat question, recalculation, disagreement, and refusal of an unregistered approval tool.
About text describes typed capabilities that reshape the workspace and feed deterministic analysis.
Screenshot visibly shows an approval request refused because sift_review_proposal is not registered.
Execution
9 /10
The long transcript and frame sheets evidence a complete, coherent decision workflow with parallel specialists, live updates, uncertainty, and approval refusal.
Evidence cited
174-second demo transcript describes concrete runtime events and state changes.
Frames show decision workspace, tool cards, changing options, and investigation request.
Packaging screenshot shows the human-only approval constraint.
Impact
9 /10
Families and other decision makers have real value in an auditable workspace that exposes tradeoffs and preserves final responsibility.
Evidence cited
Demo uses a family car decision with safety, child-seat fit, and changing priorities.
Product preserves evidence, uncertainty, disagreement, and human approval.
Creativity
9 /10
The adaptive chalkboard-like decision model, where conversation and deterministic interface continuously reshape each other, is unusually original and deeply pursued.
Evidence cited
Adds new decision dimensions not present in the initial pack and recomputes rankings.
Agent can disagree with its own model and expose what would flip the outcome.
Reviewer B (round 2)
confidence 94%
Leverage
10 /10
WebMCP is the two-way steering channel: the agent reads and reshapes live decision state, introduces dimensions, updates evidence, and cannot approve; a general UI agent would not replicate the typed shared-state interaction comparably well.
Evidence cited
Transcript shows selecting a car, adding a new seat-fit dimension, recalculating rankings, and changing weights through tools.
26 typed tools expose workspace state and actions.
Approval is absent from the tool catalog and structurally human-only.
Execution
9 /10
The 174-second transcript demonstrates a coherent central workflow with live recalculation, disagreement, uncertainty, parallel specialists, and human closure.
Evidence cited
Video transcript records end-to-end case evolution and result changes.
Frame sheets show dark-green decision workspace and evidence/review surfaces.
Shared state visibly includes options, criteria, and approval boundary.
Impact
9 /10
Complex personal and professional decisions are a real audience/problem, and the product directly improves traceability, comparison, and uncertainty handling.
Evidence cited
Family car decision demonstrates concrete needs, tradeoffs, unknowns, and evidence.
Decision state remains visible and reorganizes as new concerns arrive.
Creativity
10 /10
The adaptive chalkboard-like decision model, typed two-way workspace steering, specialist parallelism, and structurally unavailable approval are unusually original and deeply pursued.
Evidence cited
New dimensions can be created from conversation and immediately enter deterministic analysis.
Agent can disagree with its own recommendation and quantify what would flip it.
Recorded specialist runs and human-only approval are part of the interaction model.
03 Review highlights
Standouts across reviewers
Exceptional evidence of live shared state and structural human control.
Demonstrates uncertainty and model disagreement rather than hiding them.
WebMCP changes the interaction model, not just automation reliability.
Excellent evidence of live state, uncertainty, and human authority.
Red flags
High conceptual density may require onboarding.
04 Evidence
What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs;
video frames are evidence from the submitted demo, not live verification.
Probe captures come from Stage 2 interactive testing of the live product by a reviewer.
EX-01 Devpost page capture · packaging evidence, not runtime proof
· sourceEX-02 Live probe at Stage 2 · observed product behavior, reviewer-drivenEX-03 Live probe after interaction · observed product behavior
EX-V Submitted demo video — “Sift Hackathon Demo”
Contact sheets from the 174s video the team submitted. This is what reviewers were shown;
it demonstrates the product in motion but is not independent verification.
· watch the original
EX-V1 Sheet 1 of 3 · reported video evidenceEX-V2 Sheet 2 of 3 · reported video evidenceEX-V3 Sheet 3 of 3 · reported video evidence