A person stays in charge
An evaluation protocol for agent output — without chasing benchmarks
A chat demo has one promotion rule: someone liked the screenshot. An agentic-first company needs a rule you can run twice and get the same answer.
This is that rule for MeltingFace Presence. It does not publish accuracy percentages. We do not have a Board-verified baseline for “the model got better.” We do have gates that fail closed.
What we measure (observable)
| Gate | Pass means | Fail means |
|---|---|---|
| Forbidden-claims | Zero critical/high hits on shipped markdown | The draft does not merge |
| Goldens | `site-out/` matches `tests/golden/` unless an intentional refresh | Accidental HTML drift |
| Perf | HTML/CSS/JS under documented caps | The page does not ship |
| Dry-run | Preview badge on local/CI site-check builds | Missing HITL signal |
| Disposition | Issue `done` with a git SHA | Comments that reopen finished work |
`bash scripts/site-check.sh` is the bundled gate. Live GitHub Pages uses a separate `BOARD_UNLOCK=LIVE` path that must not keep a corner preview chip.
Promotion path
- Draft — agent or human, still `preview` / not approved.
- Approved — front-matter `status: approved`; SSG will include it.
- Preview — local `site-out/` + goldens + claims scan.
- Board unlock — human, out of band, for a named channel.
- One channel live — never “all adapters” in one step.
There is no step called “the model was confident.” Confidence is not a gate.
What we refuse to count as evidence
- Leaderboard scores from someone else’s benchmark.
- “Engagement will increase.”
- Compliance badges or certifications we have not earned.
- A DONE comment without a PATCH and a SHA.
Those are how demos leak into production.
How this meets LLM Ops
Orchestration decides *who* may draft vs approve. Routing decides *which* local model may run. RAG decides *what* may be cited. Ops is this protocol: measure the artifact, not the vibe, then wait for a human.
Read next
MeltingFace