AI features tend to grow one route at a time: /chat, /generate-query, /explain, /dashboard. Each one has its own auth, timeout, context fetch, payload shape and trace. Add a fifth and you copy all of it again.

Instead, everything the assistant does is one of three verbs, and each verb has one engine.

Verb What it does Runs on Who reads the result
Ask reads: run a query, find an entity, search docs server the model
Act changes what the user sees: filter, sort, navigate client the user’s screen
Make produces one typed thing: a query, a dashboard, a form draft server the user, as a card

One URL, four intents

  chat ─────────┐
  /command ─────┤                          ┌─ turn ──── agent loop (model picks tools)
  in-app button ┼──▶ POST /assistant ──▶ gate ─┼─ command ─ registry entry
  MCP client ───┘     auth, trace,         ├─ make ──── producer for that kind, no model routing
                      version              └─ call ──── one tool, forced (debug)
type AgentRequest = {
  protocolVersion: 2;
  intent: "turn" | "command" | "make";
  messages?: Message[];                       // turn, command
  command?: { name: string; args?: Record<string, string> };
  make?: { kind: string; input: unknown };    // e.g. { kind: "dashboard", input: { request } }
};

A button on a page sends make with input it already has. A chat message reaches the same producer through the model’s generate(kind, input) tool. Same producer, same cache, same payload shape. A dashboard made from chat and one made from a button are the same generation.

Rules that keep it small

  • Register once. A capability is one registry entry. A new way in is an adapter on the ingress, never a copy.
  • One naming rule. Tools, commands and kinds are all kebab-case, unique across every registry. They’re valid tool names for every model and map straight onto MCP.
  • Make is one tool. The model sees a single generate(kind) instead of a tool per kind. Its kind list is what the server can make, intersected with what the client says it can render.
  • Few commands. A slash command exists only where a shortcut adds something a chat message can’t. Asking needs no command, because plain chat is asking.

Why it works

Gates, caching, tracing and versioning happen once, at the door. Exposing the assistant over MCP becomes an adapter, not a second implementation: commands are prompts, Ask and Make are tools, context is resources.