Chat
The desktop app's persisted chat - a multi-turn conversation on your hosted models or on the user's own coding CLI, as a full page and a side-panel tab.
Chat is a persisted, multi-turn conversation with the user's chosen model, reachable as the /chat page and as a side-panel tab anywhere in the app. The chosen model decides where the turn runs: a hosted model streams from your backend's /ai/chat route and is metered there, a connected coding CLI runs on the user's own machine through the agent runtime.
What the user gets
- Multi-turn, with memory: the conversation carries prior turns, so the model has the context of everything said so far.
- Model picker: offers both sources, and can re-point a conversation mid-run, across lanes. Every switch keeps the conversation; a pick is also the new device default, so the next chat opens on it, and a conversation reopened from history keeps the model it was held on.
- Streaming: the reply streams in as it is produced; Stop ends a turn early.
- Web search (hosted lane, and
config.ai.webSearch): a composer toggle letting the model search the web, on by default until the user turns it off. It is hidden on the CLI lane, and it does not bind it: a coding CLI is offeredweb_searchandweb_extractthrough the app-tools manifest whenever the deployment's switch is on and a search engine is configured, exactly as automations and the MCP server are. There is no per-conversation toggle on that lane, and those calls are metered to the signed-in user like any other. - Clear: starts a fresh conversation while keeping the session.
Where a conversation lives
Each lane keeps its history where that lane's turns are actually run - neither store can hold the other's.
| Lane | History authority | Scope |
|---|---|---|
| Hosted model | your server's agent memory, the same threads the web app reads | per user |
| Coding CLI | the on-device runtime's local store | per user and project |
A chat started on the desktop's hosted lane is therefore the one the browser shows, and continuing it recalls what the server already holds rather than re-sending the transcript. The switcher lists both, best-effort per source, so an unreachable backend never hides the local conversations.
Switching lane moves the conversation to the other store, and it is listed once throughout:
| Switch | What the new lane gets |
|---|---|
| Hosted model to coding CLI | the whole transcript on-device; the CLI's first turn reads a bounded summary of it |
| Coding CLI to hosted model | one bounded summary message on the thread, never the raw CLI turns, so later turns do not re-bill them |
Sessions
A conversation is a session: a persisted, ordered list of turns, created on the first message and surviving restarts.
- Switcher: lists every session; each can be renamed or deleted.
- Per workspace, with
multiTenantenabled: on-device sessions follow the active organization, and ones started before an organization existed stay in the default on-device bucket. - Retention:
config.desktop.agents.maxChatsPerAgent(20 by default) caps how many on-device chats are kept per agent - a new chat past the cap deletes the oldest.0means no limit. - A turn survives walking away: on the CLI lane, switching organization or closing the window mid-answer does not cancel the run - the runtime writes the question and the reply into the session itself, so the answer is there on return.
A question from the CLI
Claude Code can stop mid-turn and ask, through its own question tool. The chat draws the same question card the hosted lane uses, and the turn stays suspended until it is settled.
| Which CLIs | Claude Code only - Codex, Grok and OpenCode have no question tool, so none can ask |
| Where | On-device chat only |
| Answering | Click an option, tick several on a multi-select, or type instead; the answer goes back to the suspended tool call |
| Skip | Tells the CLI the user declined, and it carries on without the answer |
| Automated runs | Declined immediately - nobody is watching, so the run finishes instead of waiting |
- One card per question. A single call may ask up to four things; the CLI resumes once every card is settled.
- A stopped run settles its cards. Cancelling, or leaving the conversation, marks anything still open as skipped rather than leaving a button that resumes nothing.
Compaction
A long hosted conversation is condensed automatically as it approaches the model's context window: the older half is summarized into a recap the next turn carries in its place, and the server thread is reconciled so it cannot resurrect the full history. A Compact control does the same on demand.
Compaction applies to the hosted lane only. A CLI run resends no growing history from this app, so there is nothing to condense.
The composer's context meter still reads on the CLI lane, from the counts the CLI itself reports on the terminal frame of each run: the tokens its last model step carried - the uncached prompt, both cache buckets and that step's output - over the window the CLI states it ran against. A CLI that reports neither leaves the meter estimating from the visible transcript.
When a turn cannot run
- Out of credits (hosted lane): the failure names it and offers a top-up, which opens the in-app Billing screen (the web billing page, in a build without payments).
- Runtime stopped (CLI lane): the app forks it on demand and the screen reloads once it reports ready, so this is rarely visible. Settings > Models has the App runtime section with its state and a Start control.
- No CLI connected: pick a hosted model, or set a CLI up on the Models screen.
Tasks
Builder-declared AI tasks the end-user maps to a model and reasoning effort, per task and per device, stored on their own machine.
Terminal and side panel
The side panel's two built-in tabs - Chat and Terminal - how to choose which ship and which opens first, and what a Terminal session gives the user's own CLI.