A person stays in charge
RAG as corporate memory — how agents stop inventing identity
Large models are fluent. Fluency is not memory. If an agent writes a hex code, a legal notice, or a founder bio from “what sounds right,” you do not have a brand. You have a rumor with good grammar.
Retrieval-augmented generation (RAG) here means: look up house truth, then write. The corpus is small and ranked on purpose. It is not “the whole internet.”
Ranked corpora (highest first)
- Board-gated facts — `corporate-facts.json` (legal name, emails, what we will not claim).
- Policy filters — `forbidden-claims.json` and legal boilerplate. These reject sentences; they are not a writing style guide.
- Approved pages and posts — `status: approved` markdown that already survived review.
- Internal runbooks — how the host actually runs. Useful for operators; not scraped onto the public site.
If a fact is not in (1) or (3), the agent must not invent it. The correct output is a question or a `BLOCKED` note, not a plausible paragraph.
Chunking without theater
We chunk files the company already owns: kit YAML/JSON, approved markdown, runbooks. We do not scrape customer inboxes or live social threads into the public corpus. Chunks stay small enough to cite (a table, a section, a policy line) so a reviewer can see *which* file the sentence came from.
Ranking is boring on purpose: facts beat essays; approved pages beat drafts; nothing beats a Board unlock.
Generate-time filter
Retrieval is not enough. After a draft exists:
- Forbidden-claims scan on markdown (`scripts/check-forbidden-claims.py`).
- SSG HTML-escape; `javascript:` and `data:` URLs refused.
- Janus / Board review before anything leaves dry-run.
That stack does not promise platform outcomes. It makes invented identity expensive.
What this is not
- Not a promise that models “understand” the company.
- Not live channel memory (X/LinkedIn stay dry-run until Board unlock).
- Not a license to publish percentages or certifications that are not in the facts pack.
Read next
MeltingFace