GenerateSaaS

Chat

The desktop app's persisted chat - a multi-turn conversation on your hosted models or on the user's own coding CLI, as a full page and a side-panel tab.

Chat is a persisted, multi-turn conversation with the user's chosen model, reachable as the /chat page and as a side-panel tab anywhere in the app. The chosen model decides where the turn runs: a hosted model streams from your backend's /ai/chat route and is metered there, a connected coding CLI runs on the user's own machine through the agent runtime.

What the user gets

  • Multi-turn, with memory: the conversation carries prior turns, so the model has the context of everything said so far.
  • Model picker: offers both sources, and can re-point a conversation mid-run, across lanes. Every switch keeps the conversation; a pick is also the new device default, so the next chat opens on it, and a conversation reopened from history keeps the model it was held on.
  • Streaming: the reply streams in as it is produced; Stop ends a turn early.
  • Web search (hosted lane, and config.ai.webSearch): a composer toggle letting the model search the web, on by default until the user turns it off. It is hidden on the CLI lane, and it does not bind it: a coding CLI is offered web_search and web_extract through the app-tools manifest whenever the deployment's switch is on and a search engine is configured, exactly as automations and the MCP server are. There is no per-conversation toggle on that lane, and those calls are metered to the signed-in user like any other.
  • Clear: starts a fresh conversation while keeping the session.

Where a conversation lives

Each lane keeps its history where that lane's turns are actually run - neither store can hold the other's.

LaneHistory authorityScope
Hosted modelyour server's agent memory, the same threads the web app readsper user
Coding CLIthe on-device runtime's local storeper user and project

A chat started on the desktop's hosted lane is therefore the one the browser shows, and continuing it recalls what the server already holds rather than re-sending the transcript. The switcher lists both, best-effort per source, so an unreachable backend never hides the local conversations.

Switching lane moves the conversation to the other store, and it is listed once throughout:

SwitchWhat the new lane gets
Hosted model to coding CLIthe whole transcript on-device; the CLI's first turn reads a bounded summary of it
Coding CLI to hosted modelone bounded summary message on the thread, never the raw CLI turns, so later turns do not re-bill them
A move that fails keeps the conversation on its lane and says why; the pick does not apply. A conversation moved to the hosted lane keeps its full turns on screen, and reopening it after a restart shows the summary the move seeded plus everything said since.

Sessions

A conversation is a session: a persisted, ordered list of turns, created on the first message and surviving restarts.

  • Switcher: lists every session; each can be renamed or deleted.
  • Per workspace, with multiTenant enabled: on-device sessions follow the active organization, and ones started before an organization existed stay in the default on-device bucket.
  • Retention: config.desktop.agents.maxChatsPerAgent (20 by default) caps how many on-device chats are kept per agent - a new chat past the cap deletes the oldest. 0 means no limit.
  • A turn survives walking away: on the CLI lane, switching organization or closing the window mid-answer does not cancel the run - the runtime writes the question and the reply into the session itself, so the answer is there on return.

A question from the CLI

Claude Code can stop mid-turn and ask, through its own question tool. The chat draws the same question card the hosted lane uses, and the turn stays suspended until it is settled.

Which CLIsClaude Code only - Codex, Grok and OpenCode have no question tool, so none can ask
WhereOn-device chat only
AnsweringClick an option, tick several on a multi-select, or type instead; the answer goes back to the suspended tool call
SkipTells the CLI the user declined, and it carries on without the answer
Automated runsDeclined immediately - nobody is watching, so the run finishes instead of waiting
  • One card per question. A single call may ask up to four things; the CLI resumes once every card is settled.
  • A stopped run settles its cards. Cancelling, or leaving the conversation, marks anything still open as skipped rather than leaving a button that resumes nothing.

Compaction

A long hosted conversation is condensed automatically as it approaches the model's context window: the older half is summarized into a recap the next turn carries in its place, and the server thread is reconciled so it cannot resurrect the full history. A Compact control does the same on demand.

Compaction applies to the hosted lane only. A CLI run resends no growing history from this app, so there is nothing to condense.

The composer's context meter still reads on the CLI lane, from the counts the CLI itself reports on the terminal frame of each run: the tokens its last model step carried - the uncached prompt, both cache buckets and that step's output - over the window the CLI states it ran against. A CLI that reports neither leaves the meter estimating from the visible transcript.

When a turn cannot run

  • Out of credits (hosted lane): the failure names it and offers a top-up, which opens the in-app Billing screen (the web billing page, in a build without payments).
  • Runtime stopped (CLI lane): the app forks it on demand and the screen reloads once it reports ready, so this is rarely visible. Settings > Models has the App runtime section with its state and a Start control.
  • No CLI connected: pick a hosted model, or set a CLI up on the Models screen.

On this page