A person stays in charge

LLM Routing, Data-Center Placement, and RAG as Institutional Memory

An agent army that always calls the same model for every task is like a company that always ships the same person to every meeting. You can do it. You should not.

MeltingFace researches three linked ideas: LLM routing, data-center / placement policy, and RAG as institutional memory. Together they decide *which intelligence* answers, *where it runs*, and *what truth it is allowed to use*.

LLM routing

Routing is a policy problem before it is a networking problem.

SignalExample route choice
Task classLong planning → high-context local MoE; quick classify → small model
PrivacyCustomer-sensitive → on-host only
Cost / budgetAgents have zero real-money spend authority; paid routes need human unlock
LatencyInteractive review vs overnight batch
CapabilityCode vs copy vs vision

A good router fails closed: if the preferred local route is down, do not silently escalate to a paid endpoint with agent credentials. Escalate to a human or queue.

Data-center and placement thinking

Even small teams inherit placement decisions:

Placement is not marketing architecture. It is thermal headroom, blast radius, and who can walk to the machine when a device is lost. We research placement as part of operating design for agent-first companies — how many concurrent leaves a host can honestly run, and how recovery behaves when hardware blinks.

We do not claim a public multi-region SaaS fabric in v0.1. We claim the discipline of naming the route and the recovery path.

RAG as institutional memory

Retrieval-augmented generation is how agents stop inventing the company.

Useful corpora for an agent-first org include:

RAG quality is a product surface: chunking, ranking, citation, freshness, and forbidden-claim filters after generation. Presence Bot already treats brand kit + facts as non-optional identity memory; broader retrieval research extends that pattern to engineering and ops corpora.

How the three meet

```

Task arrives

→ classify (task class, privacy, budget)

→ choose route (local / remote / human)

→ retrieve house memory (RAG)

→ generate under policy

→ human / Board gate if outbound

```

Skip retrieval and you get confident fiction. Skip routing and you get expensive fiction. Skip the gate and you get public fiction.

*Research essay. No percentage performance claims without a published protocol; no live channel claims without Board unlock.*

MeltingFace