Your Agent on a Runner
A runner is a third model source alongside built-in AI and user keys - what it can run, the one call it cannot, and the single constraint your own agent tools must satisfy to reach a device.
A paired runner is a third model source, resolved by the same decision point as built-in AI and user API keys (resolveChatRun in packages/api/src/ai/byok.ts). The user picks their CLI in the model picker and the run streams from their own machine on their own subscription.
| Source | Whose credential | Draws credits |
|---|---|---|
| Built-in AI | Your operator key | Yes - the only source that bills |
| User API key (BYOK) | The user's provider key | No - they pay their provider directly |
| Runner | The user's own coding-CLI subscription | No - their subscription executes the run |
A runner's MODEL calls are never metered by construction, not by a setting: it resolves to an unwrapped model at zero rates, with no path to the ledger - see metering is structural. Its web-tool calls ARE metered, on this lane like every other: a search spends your search vendor's key (Firecrawl, Parallel or TinyFish) rather than the user's machine, at the engine's price - nothing on a free engine unless you set one. It behaves like any other provider besides, with per-run knobs on providerOptions.runner and the resumed session id on the result's provider metadata.
What a runner model can run
| Call | On a runner |
|---|---|
streamText | Works - the adapter maps the daemon's frames to stream parts as they arrive |
generateText | Works - drains that same stream into a one-shot result |
generateObject | Throws UnsupportedFunctionalityError, naming responseFormat.type = json |
generateObject is a permanent capability gap, not a TODO: a coding CLI answers in prose and offers no constrained decoding, so there is nothing to hand a JSON schema to.
How a runner differs from a cloud provider
A runner model is a real LanguageModelV4 running through the same Mastra agent, so most things are identical. The table says exactly which are not.
| Built-in / BYOK | Runner | |
|---|---|---|
| Model resolution seam | resolveChatRun | same |
| Memory threads, conversation history | yes | same |
Your agent's persona (instructions) | yes | same, on chat AND automations |
Your agent's own tools | yes | chat yes, automations not yet |
| Per-run knobs (effort, resumed session) | providerOptions | same channel |
| Metering | built-in bills credits; BYOK never does | never - the user's own subscription pays |
| Balance gating and output clamping | built-in only | never |
| Where the agent loop runs | your server | the user's machine |
Structured output (generateObject) | yes | not supported |
| Offline | not a state | refuses an interactive turn |
Two of those rows have consequences worth spelling out:
- The loop runs on the device. The CLI decides for itself when to call a tool, so your server sees one step, not a step per tool call:
stopWhen: stepCountIs(n),onStepFinish, and anything else hooking Mastra's step lifecycle will not fire per tool call. Your tools still run - the device calls back over loopback MCP - but the orchestration is the CLI's. - Offline splits by who is waiting. An interactive turn refuses with a
409rather than quietly running on a billable cloud model. An AUTOMATION makes the opposite call: nobody is watching, so it uses your configured fallback model, and fails cleanly with a recorded error when there is none.
Conversation continuity on a runner
Every run reports the CLI's own session id back and the next turn replays it, so the CLI resumes instead of starting cold. Where that handle is stored decides whether reopening an old chat resumes or starts over.
- On this lane, the handle rides your thread. The id is written into the chat thread's metadata alongside the CLI and device that minted it, so reopening a thread days later resumes the same CLI session.
- On the dashboard's direct-dispatch runner chat, it does not. That view holds the handle in memory only;
POST /runner/dispatchaccepts aconversationIdbut persists nothing, so reopening a stored chat there starts a fresh CLI session with the messages intact. - The handle belongs to one device AND one CLI. Both surfaces check both axes and start fresh on a mismatch rather than replay a foreign handle - Codex hard-fails on an unknown session id, so retargeting costs one turn of context where replaying would cost the turn itself.
- You wire none of it. If you dispatch runs yourself, the id comes back on the result's provider metadata and goes out as the next dispatch's
conversationId.
Automated runs on a runner
An automation pinned to a runner CLI dispatches straight to the device rather than running the agent loop, so the cron is not held open for the length of a CLI run.
- It carries your agent's persona, exactly as chat and cloud automations do.
- It finalizes through the same
finalizeAutomationRun, solastRunAt,lastResult,lastErrorand the history row match in shape. - It settles when the device reports back - the row is written when the daemon posts its terminal frame, not inline.
- Offline with no fallback model, it fails immediately with a recorded error rather than retrying indefinitely.
Your own agent tools reach the device
Tools you declare on the assistant agent (new Agent({ tools: { myTool } })) work on a runner CHAT run. The run advertises them over the device's loopback MCP, and your backend resolves each call server-side under the verified user, so the tool's secrets never leave your server. An app capability wins a name collision (refused outright at dispatch), and a run may only call what its dispatch advertised.
execute must be reconstructible from the agent definition. The device posts the call back over HTTP long after the dispatching request ended, possibly into a different serverless instance, so the closure the tool was declared in is gone and the agent's tools are rebuilt from scratch. Whatever execute needs must come from its own arguments plus the run's persisted scope (owner, surface, org).
This works - everything it needs arrives in its arguments:
const saveAudit = tool({
description: "Save a completed site audit.",
inputSchema: z.object({ url: z.string(), score: z.number() }),
execute: async ({ url, score }) => db.insert(audits).values({ url, score })
});This does not, though it works fine on built-in AI and BYOK, where execute runs inside the request that declared it:
new Agent({
// A dynamic `tools` function is re-resolved on the device's callback with a FRESH,
// EMPTY request context - so `tenantId` is undefined and the captured client is gone.
tools: ({ requestContext }) => {
const client = clientFor(requestContext.get("tenantId"));
return { saveAudit: tool({ /* ... */ execute: async (args) => client.save(args) }) };
}
});Three limits worth knowing:
- Runner-pinned AUTOMATIONS get no agent tools yet - they advertise the app capability set alone, so use a capability if an automated runner run needs one.
- The default assistant agent only. Mastra never tells a model which agent invoked it, so put device-bound tools on the assistant.
- Memory-managed tools are not agent tools. With working memory on,
updateWorkingMemoryis advertised to the device and answeredUnknown tool, which the CLI recovers from.
Always-on runners
Run the runner on a VPS so automated runs never wait, and what happens to dispatched runs, automations, and chat turns while a device is offline.
Dispatch from Your Code
Run device-side work on a user's paired runner straight from your product code - what pairing authorizes, the run statuses, and how results come back.