A person stays in charge

Local Orchestration and Custom LLMs for Agent Workforces

If your agent army only exists inside a single vendor chat UI, you do not have an operating model — you have a subscription.

MeltingFace researches local orchestration: agent runtimes, adapters, queues, and model servers you can inspect on infrastructure you control (or at least policy you can name). Custom LLMs sit next to that stack — not as a brand stunt, but as a way to fit house tools, house style, and house context floors.

What “local orchestration” means

Local does not mean “never use the cloud.” It means:

For agent-first companies, that control plane is what lets humans stay in charge while agents run for hours.

Why custom / adapted LLMs matter

Off-the-shelf models are excellent generalists. Company work is not general:

Adaptation paths (fine-tuning, continued pretrain, preference optimization, system-prompt + tool training packs) are research tools for fitting models to that reality. Quantization and partial GPU offload are operational tools for running them next to a desktop and a display GPU without pretending the machine is a full data center.

We talk about these choices openly because agent productivity collapses when the model thrashing the GPU is misconfigured for the job.

Orchestration layers we care about

  1. Company graph — issues, parents, assignees, terminal states
  2. Agent runtime — profiles, timeouts, max iterations, sole-leaf gates
  3. Model server — context, batch, offload, recovery after device loss
  4. Policy — dry-run, no real money, brand kit read-only
  5. Human gate — Board unlock, Janus review, content status

Skip a layer and the army either idles or burns.

What we are not claiming

We claim that orchestration + model fit + gates is the productive unit of research — not the model name alone.

*Engineering research notes. Configurations and commercial offerings are Board-gated; no live multi-tenant claims in v0.1.*

MeltingFace