webmcp/analysis

Ops Co-pilot

An AI-friendly infra dashboard where agents check real service health and act but every risky action needs a human's explicit approval before it runs.

Aggregate 37
Leverage 10
Execution 9
Impact 9
Creativity 9

Each criterion 1–10, equally weighted; aggregate is their sum. Ranking is the pipeline's consolidated output.

01 Links & metadata

Category
dev-tool / dev-tool
Origin
built for the challenge / built for the challenge
Access
no auth
Eligibility
LIKELY_ELIGIBLE / LIKELY_ELIGIBLE
Substitution
MAJOR_DELTA / MAJOR_DELTA
Demo liveness
alive

Origin, access model, eligibility, and substitution are reviewer diagnostics, not judging criteria. Authentication requirements are not penalized.

02 The two blind reviews

Two independent reviewers scored this project blind, from a sanitized evidence packet. Scores are shown separately so the reasoning stays inspectable. A withheld score means the reviewers differed by more than two points.

Reviewer A (round 1)

confidence 73%
Leverage
8 /10

Structured operational tools plus approval tokens and audit state make consequential actions safer and more reliable than an agent clicking a dashboard.

Evidence cited
  • Seven registered tools cover health, alerts, notes, and operational actions.
  • High-risk actions halt for a human dialog and single-use cryptographically bound token.
  • Claims real Render API calls against a separately deployed service.
Execution
7 /10

The architecture and safety boundary are unusually concrete, but the packet has no gallery images and frame evidence is not sufficient to verify live infrastructure actions end-to-end.

Evidence cited
  • Go backend claims service registry, alert engine, guardrails, token system, and audit log.
  • Description distinguishes immediate read-only/low-risk actions from approval-gated risky actions.
Impact
9 /10

Operators managing production services have a high-value need for fast diagnosis and controlled remediation; the product directly addresses the trust/safety trade-off.

Evidence cited
  • Targets real service uptime, DB/Redis status, errors, alerts, restart, and scaling.
  • Explicit human approval reduces blast radius for consequential actions.
Creativity
8 /10

The agent-plus-human operational control model is a strong, ambitious application of WebMCP, especially with token-bound approvals and auditability.

Evidence cited
  • Risk-tiered tool execution policy.
  • Real deployed service observability rather than purely simulated dashboard data is claimed.

Reviewer B (round 2)

confidence 74%
Leverage
8 /10

Structured live observability plus cryptographically bound approval is a meaningful reliability delta over UI driving.

Evidence cited
  • Seven tools expose health and operations.
  • Risky actions require explicit approval and single-use token.
  • Real service and audit log claimed.
Execution
7 /10

Concrete architecture and live-service scenario, but no video proves an end-to-end action.

Evidence cited
  • Alive demo and public repo.
  • Go backend, React frontend, real observability endpoints, Render calls described.
  • Three frame sheets listed.
Impact
9 /10

Directly addresses safe remediation of production infrastructure.

Evidence cited
  • Uses real uptime, DB/Redis, metrics, and errors.
  • Restart and scaling are human-gated.
Creativity
7 /10

Known copilot pattern, strengthened by live integration and exact action binding.

Evidence cited
  • Risk-based separation of immediate and gated tools.
  • Token bound to intended action.

03 Review highlights

Standouts across reviewers

  • Cryptographically bound, single-use approval token.
  • Real-service observability and audit-log framing.
  • Concrete production safety mechanism.

Red flags

  • No gallery images; no transcript; live API and approval execution are claims not directly demonstrated in supplied visuals.
  • No video.
  • Deployment dependencies may affect reproducibility.

04 Evidence

What each artifact proves is labeled on the artifact itself. A Devpost page capture is packaging evidence, not proof the product runs; video frames are evidence from the submitted demo, not live verification.

Devpost page capture for Ops Co-pilot
EX-01 Devpost page capture · packaging evidence, not runtime proof · source

EX-V Submitted demo video

Contact sheets from the ?s video the team submitted. This is what reviewers were shown; it demonstrates the product in motion but is not independent verification. · watch the original

Contact sheet 1 from the Ops Co-pilot demo video
EX-V1 Sheet 1 of 3 · reported video evidence
Contact sheet 2 from the Ops Co-pilot demo video
EX-V2 Sheet 2 of 3 · reported video evidence
Contact sheet 3 from the Ops Co-pilot demo video
EX-V3 Sheet 3 of 3 · reported video evidence