A person stays in charge

An evaluation protocol for agent output — without chasing benchmarks

A chat demo has one promotion rule: someone liked the screenshot. An agentic-first company needs a rule you can run twice and get the same answer.

This is that rule for MeltingFace Presence. It does not publish accuracy percentages. We do not have a Board-verified baseline for “the model got better.” We do have gates that fail closed.

What we measure (observable)

GatePass meansFail means
Forbidden-claimsZero critical/high hits on shipped markdownThe draft does not merge
Goldens`site-out/` matches `tests/golden/` unless an intentional refreshAccidental HTML drift
PerfHTML/CSS/JS under documented capsThe page does not ship
Dry-runPreview badge on local/CI site-check buildsMissing HITL signal
DispositionIssue `done` with a git SHAComments that reopen finished work

`bash scripts/site-check.sh` is the bundled gate. Live GitHub Pages uses a separate `BOARD_UNLOCK=LIVE` path that must not keep a corner preview chip.

Promotion path

  1. Draft — agent or human, still `preview` / not approved.
  2. Approved — front-matter `status: approved`; SSG will include it.
  3. Preview — local `site-out/` + goldens + claims scan.
  4. Board unlock — human, out of band, for a named channel.
  5. One channel live — never “all adapters” in one step.

There is no step called “the model was confident.” Confidence is not a gate.

What we refuse to count as evidence

Those are how demos leak into production.

How this meets LLM Ops

Orchestration decides *who* may draft vs approve. Routing decides *which* local model may run. RAG decides *what* may be cited. Ops is this protocol: measure the artifact, not the vibe, then wait for a human.

MeltingFace