A person stays in charge
LLM Routing, Data-Center Placement, and RAG as Institutional Memory
An agent army that always calls the same model for every task is like a company that always ships the same person to every meeting. You can do it. You should not.
MeltingFace researches three linked ideas: LLM routing, data-center / placement policy, and RAG as institutional memory. Together they decide *which intelligence* answers, *where it runs*, and *what truth it is allowed to use*.
LLM routing
Routing is a policy problem before it is a networking problem.
| Signal | Example route choice |
|---|---|
| Task class | Long planning → high-context local MoE; quick classify → small model |
| Privacy | Customer-sensitive → on-host only |
| Cost / budget | Agents have zero real-money spend authority; paid routes need human unlock |
| Latency | Interactive review vs overnight batch |
| Capability | Code vs copy vs vision |
A good router fails closed: if the preferred local route is down, do not silently escalate to a paid endpoint with agent credentials. Escalate to a human or queue.
Data-center and placement thinking
Even small teams inherit placement decisions:
- Laptop iGPU vs workstation
- Colocated inference box vs remote region
- Future dedicated racks for multi-agent load
Placement is not marketing architecture. It is thermal headroom, blast radius, and who can walk to the machine when a device is lost. We research placement as part of operating design for agent-first companies — how many concurrent leaves a host can honestly run, and how recovery behaves when hardware blinks.
We do not claim a public multi-region SaaS fabric in v0.1. We claim the discipline of naming the route and the recovery path.
RAG as institutional memory
Retrieval-augmented generation is how agents stop inventing the company.
Useful corpora for an agent-first org include:
- Brand kit and corporate facts — identity that agents may not rewrite
- Approved site and policy pages — what we already said in public
- Runbooks and ADRs — how systems are meant to behave
- Issue and disposition history — what was decided, not what a model guesses
RAG quality is a product surface: chunking, ranking, citation, freshness, and forbidden-claim filters after generation. Presence Bot already treats brand kit + facts as non-optional identity memory; broader retrieval research extends that pattern to engineering and ops corpora.
How the three meet
```
Task arrives
→ classify (task class, privacy, budget)
→ choose route (local / remote / human)
→ retrieve house memory (RAG)
→ generate under policy
→ human / Board gate if outbound
```
Skip retrieval and you get confident fiction. Skip routing and you get expensive fiction. Skip the gate and you get public fiction.
Read next
*Research essay. No percentage performance claims without a published protocol; no live channel claims without Board unlock.*
MeltingFace