+
Lightspeed is open-source infrastructure for running managed agent fleets as durable workflows.
"Managed agents" is an emerging pattern that separates the core agent loops from the VM or sandbox they use. Agents survive restarts, can run for months, and stay cheap when idle. When they
@@ -126,7 +132,7 @@ The current implementation includes:
- [x] **Workflow-backed plugins**: external Temporal workflows can extend session with various tools and custom logic
- [x] **One backend binary**: run every runtime role in one process or scale them independently across Temporal workers
-**Borrowed compute**
+**Attached compute**
- [x] **Dedicated VMs**: attach an existing machine or provision one through the
included Incus provider
diff --git a/docs/documentation/deployment/configuration.md b/docs/documentation/deployment/configuration.md
index 9856a21fb..405e1d5de 100644
--- a/docs/documentation/deployment/configuration.md
+++ b/docs/documentation/deployment/configuration.md
@@ -120,11 +120,16 @@ it. Deployment-level `OPENAI_API_KEY` and `ANTHROPIC_API_KEY` are fallback
credentials when the corresponding universe record is absent. A disabled or
broken stored record blocks fallback.
-`LIGHTSPEED_CHAT_PROVIDER` and `LIGHTSPEED_CHAT_MODEL` change the runtime's
-default provider ID and model name. Its default API kind remains
-`openai:responses`; changing a provider ID alone does not switch wire formats.
-For Anthropic or another route, select provider, API kind, and model explicitly
-in a profile. See [Models and credentials](../using-lightspeed/models-and-credentials.md).
+Choose each universe's agent and speech-to-text defaults under **Models →
+Defaults**, or through `lightspeed model defaults`. Each selection names a
+provider, API kind, and model. New sessions use an explicit session model,
+then the profile model, then the universe default; existing sessions keep
+their resolved model. There is no deployment-wide model fallback.
+
+Remove `LIGHTSPEED_CHAT_PROVIDER` and `LIGHTSPEED_CHAT_MODEL` from older
+configuration: the runtime rejects them at startup. See
+[Models and credentials](../using-lightspeed/models-and-credentials.md#choose-universe-defaults)
+for setup and precedence.
## Choose the blob backend
diff --git a/docs/documentation/deployment/troubleshooting.md b/docs/documentation/deployment/troubleshooting.md
index caa77d0c6..8b3e43323 100644
--- a/docs/documentation/deployment/troubleshooting.md
+++ b/docs/documentation/deployment/troubleshooting.md
@@ -49,6 +49,7 @@ If the runtime exits before serving, check the error against these dependencies:
| Object-store configuration | A nonempty bucket when any `LIGHTSPEED_OBJECT_STORE_*` variable is set, plus the intended endpoint and credentials. |
| Role or environment routing | Valid roles and the internal environment gateway URL/token on every process without that role. |
| Secret configuration | A base64 master key decoding to 32 bytes, matching the stored encrypted state. |
+| Retired model variables | Remove `LIGHTSPEED_CHAT_PROVIDER` and `LIGHTSPEED_CHAT_MODEL`; configure each universe's model defaults through Models or `lightspeed model defaults`. |
The Platform needs its own database and authentication settings. Its migrations
run before HTTP serving. Do not point it at the runtime database or reset
@@ -130,6 +131,22 @@ not confuse removal with disabling access. A coding-agent subscription login
also does not authenticate Lightspeed's own session inference. Follow
[Models and credentials](../using-lightspeed/models-and-credentials.md).
+If session creation reports `model_default_unset`, select an explicit session
+or profile model, or configure **Models → Defaults → Agent runs**. A provider
+key alone does not choose a model. Voice preparation and dictation need the
+separate **Speech-to-text** default.
+
+For context-length errors, inspect the effective compaction policy and recovery
+attempts in session settings. Enabled sessions can compact and retry within
+the same run; Disabled and historical omitted policies do not recover
+automatically. A protected input larger than the usable window can still fail.
+See [Manage long conversations](../using-lightspeed/sessions-and-runs.md#manage-long-conversations).
+
+If a specific entry keeps making the provider reject requests, stop active
+work and use [CLI context repair](../using-lightspeed/cli.md#repair-rejected-context)
+to inspect and replace a user message or tool result. This preserves the
+event history; do not rewrite the stored log to repair active context.
+
If discovery succeeds but an agent has no expected tool, inspect its profile's
capability grants and the MCP server/tool selection. Registering a server
does not grant every session permission to use it. Also distinguish a tool
diff --git a/docs/documentation/deployment/upgrades-and-recovery.md b/docs/documentation/deployment/upgrades-and-recovery.md
index 728bf90ea..69d42b77f 100644
--- a/docs/documentation/deployment/upgrades-and-recovery.md
+++ b/docs/documentation/deployment/upgrades-and-recovery.md
@@ -56,13 +56,38 @@ not guarantee compatibility with saved session configuration or workflow
histories. Use matching clients and review daemon changes even when protocol
mismatch would not trigger an automatic daemon update.
+## Review model and context configuration
+
+Remove `LIGHTSPEED_CHAT_PROVIDER` and `LIGHTSPEED_CHAT_MODEL` from the runtime
+environment before starting current workers; both variables are retired and
+cause startup to fail. After the runtime is available, save the intended
+provider/API/model route in each universe's **Agent runs** default using
+**Models → Defaults** or `lightspeed model defaults set agent-run`. Configure
+**Speech-to-text** separately for dictation and channel voice messages.
+Deployment provider credentials remain supported, but do not choose a model.
+See [Configure universe model defaults](../using-lightspeed/cli.md#configure-universe-model-defaults).
+
+Existing sessions retain their stored model. Historical sessions with an
+omitted compaction policy also retain their disabled replay behavior. New
+sessions and configuration replacements resolve omission to automatic
+standalone compaction; preserve an explicit Disabled setting when automatic
+compaction and context-limit recovery should stay off. Review these settings
+when applying a profile or replacing configuration during an upgrade.
+See [Manage long conversations](../using-lightspeed/sessions-and-runs.md#manage-long-conversations).
+
+Audio integrations must transcribe before session admission. Run start,
+context append, and steering reject raw audio; use the
+[standalone transcription flow](../integrating-and-extending/api-and-typescript.md#transcribe-audio-before-submitting-session-input)
+and submit prepared text.
+
## Enable company sign-in on an existing Platform
The company-identity migration adds identity provenance, session revocation
versions and durable access records. Existing users remain local and keep
their memberships; the migration does not grant company or emergency access.
-The current Platform schema revision is 3. Startup applies the generated
-Platform migrations before serving requests.
+The required Platform revision is `LIGHTSPEED_PLATFORM_SCHEMA_REVISION` in
+the target release's [metadata](../../../release/metadata.env). Startup applies
+the generated Platform migrations before serving requests.
Before enabling OIDC, designate an existing local password admin through
`LIGHTSPEED_PLATFORM_ADMIN_EMAIL` and verify its emergency login. Use a
diff --git a/docs/documentation/development/testing-and-evaluation.md b/docs/documentation/development/testing-and-evaluation.md
index 1dc9f7b9c..b830a7af1 100644
--- a/docs/documentation/development/testing-and-evaluation.md
+++ b/docs/documentation/development/testing-and-evaluation.md
@@ -227,8 +227,8 @@ Without `--model`, the harness uses the provider's model environment variables
and then its compiled default. OpenAI Responses checks
`OPENAI_RESPONSES_MODEL`, Completions checks `OPENAI_COMPLETIONS_MODEL`, and
both fall back to `OPENAI_LIVE_MODEL`. Anthropic checks
-`ANTHROPIC_MESSAGES_MODEL`, then `ANTHROPIC_LIVE_MODEL`. The product's
-`LIGHTSPEED_CHAT_MODEL` setting does not configure this harness. Provider base
+`ANTHROPIC_MESSAGES_MODEL`, then `ANTHROPIC_LIVE_MODEL`. Universe model
+defaults do not configure this evaluation harness. Provider base
URL overrides are honored too; record them with a comparison.
### Add a case with observable assertions
diff --git a/docs/documentation/environments/using-environments.md b/docs/documentation/environments/using-environments.md
index 39d6b44bb..17fb18a62 100644
--- a/docs/documentation/environments/using-environments.md
+++ b/docs/documentation/environments/using-environments.md
@@ -95,11 +95,20 @@ support the operation, and its operating-system permissions still apply.
read instructions from an attached environment independently of whether it
permits commands.
-The session's context includes an **Environment catalog** listing every
-attachment with its display name, status, access, working directory, and
-which one is active, so the agent knows what it may use before calling a
-tool. After a switch, the old catalog is removed until the next idle refresh.
-Use `environment_list` or `environment_read` for current status during a run.
+The session's **Environment catalog** lists attached machines with their
+references, display names, default markers, and access grants. Selection and
+lifecycle changes do not replace this directory. Use `environment_read` for
+current selection, status, and working directory; it remains available when
+environment selection tools are disabled.
+
+Long environment IDs appear to the model as short `env:` references in
+catalogs, control results, and job handles. These resolve only against the
+session's authorized attachments. Full IDs remain accepted and are still the
+identities stored in profiles and public API records. An ambiguous short
+reference is rejected; use the full ID in that case. Short references do not
+grant additional access. Skill and prompt discovery remains tied to the
+selected machine, as described in
+[Workspaces and skills](../using-lightspeed/workspaces-and-skills.md#discover-skills-installed-on-a-machine).
## Keep files in the right domain
diff --git a/docs/documentation/getting-started/quickstart.md b/docs/documentation/getting-started/quickstart.md
index 23152b66e..80a2005e9 100644
--- a/docs/documentation/getting-started/quickstart.md
+++ b/docs/documentation/getting-started/quickstart.md
@@ -120,7 +120,7 @@ session. A saved key needs access to the model you will select.
1. Open **Sessions** and choose the plus button labeled **New session**.
2. Enter `First conversation` as the **Name** and leave **Profile** at
- **No profile (engine defaults)**.
+ **No profile (universe default)**.
3. Choose **Customize setup…**. Under **Model configuration → Model**, select
a conversational model from the provider you just connected.
4. Choose **Create session**.
@@ -130,9 +130,11 @@ session. A saved key needs access to the model you will select.
*Demo mode: choose **Customize setup…** to select a model before creating
the conversation.*
-Select the model explicitly. Adding a credential does not change the
-deployment's default model, so leaving **Deployment default** selected can
-send the request to a different provider.
+Select the model explicitly for this walkthrough. Alternatively, an Operator
+or Admin can choose **Agent runs** under **Models → Defaults** and leave
+**Universe default** selected in session setup. Adding a credential alone
+does not select that default. If neither the session nor its profile supplies
+a model and the universe has no default, session creation fails.
Send a short message:
diff --git a/docs/documentation/how-it-works/context-and-storage.md b/docs/documentation/how-it-works/context-and-storage.md
index f4bf8393a..fb6b52945 100644
--- a/docs/documentation/how-it-works/context-and-storage.md
+++ b/docs/documentation/how-it-works/context-and-storage.md
@@ -31,6 +31,20 @@ rewriting entries happens through events. An entry removed from the active
set still has the event that introduced it in session history. Replaying that
history reconstructs both the earlier state and the later removal.
+`session/context/replace` lets an authorized caller repair an active user message or
+tool result by entry ID. Replacement retains its ID, kind, and position, and
+a tool result remains paired with its call. The operation is refused during
+an active run and reports an outcome for each entry. Redaction through the
+CLI uses this same operation with a placeholder; original events and blobs
+remain retained. See [Repair rejected context](../using-lightspeed/cli.md#repair-rejected-context).
+
+Provider adapters can also transform media for a particular request without
+changing active entries. Image normalization sends a bounded copy; request
+media budgeting replaces older items with handle-bearing omission notes.
+Stored references, previews, and downloads still identify the originals.
+[Tool media](../using-lightspeed/tools-and-mcp.md#see-images-and-documents-from-tools)
+describes those limits.
+
## Keep provider-native material at the provider boundary
Provider responses contain more than visible text. They can include tool calls,
@@ -115,10 +129,26 @@ to the renderer cannot change an earlier request. Runtime catalogs are managed
by the session workflow; clients can publish their own catalogs under separate
keys through the [context API](../../../crates/api/contract/api-reference.md).
-A changed keyed catalog is appended as the current version at the context
-tail. Earlier versions remain in their original positions, with their bytes
-unchanged. The new entry identifies which catalog it supersedes. This lets a
-skill or sub-agent catalog change while preserving an earlier cached prefix.
+When a keyed catalog's title or rendered content changes, the new version is
+appended at the context tail. Earlier versions remain in their original
+positions, with their bytes unchanged. The new entry identifies which catalog
+it supersedes. A metadata-only refresh instead records an in-place replacement
+that preserves the entry ID, position, text, and supersession link. Replay
+reconstructs either path, so fresh discovery diagnostics need not append an
+identical message or disrupt a cached prefix.
+
+Catalog text describes available resources rather than constantly changing
+status. The environment directory includes references, display names, default
+markers, and access; selection, lifecycle state, and working directories come
+from `environment_read`. Skills list names, descriptions, and `SKILL.md` paths,
+sorted by visible path. VFS mounts describe workspace or snapshot storage and
+access without printing internal IDs or snapshot hashes. Revisions and
+discovery details remain structured provenance.
+
+Failed environment skill discovery can retain a matching prior catalog as
+stale, with API warnings, while leaving its rendered paths unchanged. Matching
+requires the same environment and configured discovery scope; changed roots,
+working directory, access, or a revoked attachment invalidate that fallback.
Active context retains at most five superseded versions per catalog;
compaction clears them. Instructions and other ordinary keyed entries replace
@@ -144,37 +174,71 @@ that can support later turns. It is a lossy context transformation. The original
session events and output descriptors remain in history, so this is separate
from deleting stored data.
-The core treats standalone compaction as explicit work. It records the request
-and selected context revision, waits for an adapter result, then commits the
-replacement. If the operation fails, it clears the pending state and retains
-the original entries. A stale result cannot rewrite a newer context revision.
-The core rejects new runs and context edits while standalone compaction is
-pending; the hosted workflow holds their admissions until the operation finishes.
-
-Standalone compaction can be requested manually or through an optional
-threshold. It starts only with no active or queued run. The threshold sums
-token estimates for compactable entries. It needs a valid estimate to fire;
-provider usage totals are not a substitute for the current context size.
-
-The adapter mechanism depends on the API kind:
+Automatic policy and an individual compaction operation are separate:
-| Route and mode | What performs the compaction |
+| Setting | Automatic behavior |
| --- | --- |
-| OpenAI Responses, `provider_triggered` | The ordinary generation request includes `context_management`. Returned native compaction material enters context, and older eligible conversation is pruned. |
-| OpenAI Responses, `provider_standalone` | A separate call to the Responses compact endpoint produces native compaction output. |
-| Anthropic Messages, `provider_standalone` | A summarization request with Lightspeed-authored instructions produces a plain-text replacement summary. |
-| Chat Completions, `provider_standalone` | A summarization request produces a plain-text replacement summary. |
-
-Only OpenAI Responses supports `provider_triggered` compaction. The summary
-adapters use `targetTokens` as guidance and an output budget; the Responses
-compact endpoint does not receive that setting.
-
-Instructions and current catalogs survive compaction. Skill reads and inserted
-skill text follow ordinary conversation retention. Eligible conversation and
-superseded catalogs can be removed;
-nonterminal tool work and unconsumed active input are protected. These rules
-retain the material needed to continue valid execution while reducing the
-conversation carried forward.
+| **Engine default** (omitted policy) | New sessions and configuration replacements resolve to harness-managed standalone compaction. |
+| `providerStandalone` | The harness schedules standalone work at safe boundaries between generations, including during an active run. |
+| `providerTriggered` | Supported OpenAI Responses and Anthropic Messages routes compact inside generation. Chat Completions rejects this mode. |
+| `disabled` | No automatic compaction or context-limit recovery. Explicit API compaction remains available. |
+
+Historical events with an omitted policy retain their original disabled
+behavior on replay. Replacing that session's configuration with the policy
+omitted adopts the new default; an explicit Disabled setting stays disabled.
+
+For standalone policy, an explicit `compactThresholdTokens` controls the
+proactive trigger. Otherwise it uses 80% of known usable input capacity.
+`context.inputLimitTokens` can supply an explicit capacity; without one, the
+runtime uses reported capacity where available. Unknown capacity leaves
+error-driven recovery available without inventing a numeric limit. Request
+occupancy and the effective model's capacity are recorded facts for the
+deterministic harness; cumulative billed tokens are not window occupancy.
+
+When an enabled session reaches a typed provider context-length failure, the
+harness can compact and resume the same run using already-recorded tool
+results. Recovery is bounded to two attempts per consecutive overflow
+sequence. Authentication failures and unrelated request rejections do not
+trigger this path. If protected input cannot fit, or recovery cannot produce
+a usable replacement, the run fails with the source context retained.
+
+`session/context/compact` requests one standalone operation in any mode,
+including Disabled, without changing automatic policy. Active work queues it
+until the current generation and tool work reach a safe boundary. The harness
+records the covered prefix and context revision, then commits a validated
+replacement atomically. Failed or stale results cannot erase newer context.
+
+The standalone adapter depends on the route and its supported capabilities:
+
+| Route | Standalone operation |
+| --- | --- |
+| OpenAI Responses | The Responses compact endpoint; retain the complete returned native window, including its encrypted state. |
+| Anthropic Messages | Native on-demand compaction on supported models; retain the signed block unchanged. Older models use a Lightspeed-authored summary generation. |
+| Chat Completions | A Lightspeed-authored summary generation. |
+
+An explicitly unavailable native operation can fall back to summarization on
+the same provider, endpoint, and model. Ordinary invalid requests do not
+authorize that fallback. Summary adapters use `targetTokens` as guidance and
+an output budget; native operations do not guarantee that output size.
+
+A retained Anthropic on-demand signed block requires subsequent compactions
+to use standalone execution, even if the requested policy was provider
+triggered. The effective strategy reflects that transition. Retained signed,
+encrypted, or reasoning state can also restrict otherwise permitted model
+changes within a session's pinned provider route.
+
+Instructions and current catalogs survive compaction. The harness protects
+unconsumed input and unanswered tool exchanges, and normally keeps the two
+newest settled exchanges with their preceding user input while compacting an
+older prefix. Manual compaction without an older prefix and repeated overflow
+recovery can cover more settled history. Skill reads and superseded catalogs
+follow conversation retention. The adapter processes bounded chunks, but only
+the complete validated replacement changes active context.
+
+Session settings and the API's active-context projection report effective
+mode, threshold source, queued or pending work, and recovery attempts. See
+[Sessions and runs](../using-lightspeed/sessions-and-runs.md#manage-long-conversations)
+for the user controls.
## Keep large bytes outside Temporal history
@@ -284,5 +348,18 @@ input, a tool result, or a sub-agent result. The model sees a stable `media:`
handle it can use to refer to the attachment; clients resolve that handle
against the descriptors exposed by the API.
+Explicit VFS references use a separate `file:` handle and typed file
+descriptor. `vfs_reference` resolves an immutable file version; reads and
+writes do not implicitly publish file attachments. The descriptor travels
+with a sub-agent result only when its successful final answer cites that
+known attachment. The receiving session records the descriptor and retains
+the bytes, without needing the child's workspace or session.
+
+Clients resolve these handles against recorded attachments, including those
+from historical completion pages. Unknown or ambiguous handles are
+unavailable. A source workspace and path provide optional navigation; they
+do not determine which bytes the attachment opens. Short handles grant no
+access, and a hash in prose alone is not a retention root.
+
Together, these views let the model work with a manageable context while
people inspect retained history and the runtime reconstructs execution state.
diff --git a/docs/documentation/integrating-and-extending/api-and-typescript.md b/docs/documentation/integrating-and-extending/api-and-typescript.md
index 313b178ec..b9a724b2d 100644
--- a/docs/documentation/integrating-and-extending/api-and-typescript.md
+++ b/docs/documentation/integrating-and-extending/api-and-typescript.md
@@ -124,6 +124,42 @@ The run's final message and a file it wrote are different outputs. Reuse the
session for a follow-up conversation, or create another session ID for
independent work.
+## Transcribe audio before submitting session input
+
+Run start, context append, and steering reject raw audio media. Transcribe
+audio separately, then submit ordinary `text` or `textRef` input:
+
+1. Upload the audio to the universe's content store with `blobs/put`.
+2. Call `transcriptions/start` with a stable `idempotencyKey` and an `audio`
+ object containing `blobRef`, `mime`, and `name`. Supply a complete `model`
+ route, or omit it to resolve the universe's **Speech-to-text** default.
+ The API kind must be `openai:audio-transcriptions`.
+3. Save the returned `transcriptionId` and poll `transcriptions/read` until
+ `status` is terminal. `succeeded` supplies plain UTF-8 `text` and
+ `transcriptRef`. Handle `failed`, `cancelled`, and `expired` explicitly;
+ `failure` provides details when present.
+4. Submit the prepared text through the ordinary session API. For unchanged
+ transcripts, an optional `provenanceRef` can point to the source audio blob
+ in the same universe. Admission checks that the source exists and retains
+ it with the session content.
+
+Transcription creates no session or run. A direct key needs `blobs/put` for
+upload, `transcriptions` for the job, and `session` for subsequent admission.
+Use `transcriptions/cancel` to request cancellation; abandoning an HTTP wait
+does not cancel the job.
+
+Matching retries with the same requester, idempotency key, audio, and options
+rejoin the original job even if the universe default has changed. Reusing the
+key with changed input conflicts. This identity lasts for the Temporal
+namespace's workflow-history retention period. The resolved model is fixed
+at admission. Asserted actors can read and cancel only their own drafts;
+direct universe keys retain their method-group authority.
+
+Unsubmitted audio and transcripts use ordinary CAS collection grace and can
+expire. Persist or admit the result when it is needed beyond that period.
+When users edit a transcript, send their reviewed text as an ordinary message;
+do not present the edited words as the unchanged transcription.
+
## Keep retry identity with the business operation
Persist the session ID, submission ID, request input/configuration, and returned
diff --git a/docs/documentation/reference/environment-variables.md b/docs/documentation/reference/environment-variables.md
index f03a7baad..1c6abae8d 100644
--- a/docs/documentation/reference/environment-variables.md
+++ b/docs/documentation/reference/environment-variables.md
@@ -91,10 +91,16 @@ Provider keys in the environment are deployment-wide fallback credentials.
They may be omitted when every request resolves a stored, universe-scoped
provider credential.
+Model selections belong to each universe. Set **Agent runs** and
+**Speech-to-text** under **Models → Defaults**, or use
+[`lightspeed model defaults`](../using-lightspeed/cli.md#configure-universe-model-defaults).
+`LIGHTSPEED_CHAT_PROVIDER` and `LIGHTSPEED_CHAT_MODEL` are retired: setting
+either prevents runtime startup. Remove them and save the intended complete
+provider/API/model route in each universe that needs it. Credential fallback
+does not supply an unset model default.
+
| Variable | Requirement/default | Purpose |
| --- | --- | --- |
-| `LIGHTSPEED_CHAT_PROVIDER` | `openai` | Deployment default provider ID for sessions that do not choose a model. |
-| `LIGHTSPEED_CHAT_MODEL` | `gpt-5.5` | Deployment default model for sessions that do not choose a model. |
| `OPENAI_API_KEY` | Conditional | Default OpenAI Responses, Chat Completions, and audio-transcription credential. |
| `OPENAI_BASE_URL` | `https://api.openai.com/v1` | Deployment fallback URL for the built-in `openai` provider, shared by Responses, Chat Completions, and audio transcription. Custom universe providers use their stored endpoint instead. |
| `OPENAI_ORG_ID` | Unset | Optional `OpenAI-Organization` header. |
@@ -124,6 +130,12 @@ writes above that inline limit fail.
### Audio preprocessing
+These settings control the optional transcoder used by standalone
+transcription. Select the transcription provider and model through the
+universe's **Speech-to-text** default or an explicit `transcriptions/start`
+model. Session input must already be transcribed; see
+[the integration flow](../integrating-and-extending/api-and-typescript.md#transcribe-audio-before-submitting-session-input).
+
| Variable | Requirement/default | Purpose |
| --- | --- | --- |
| `LIGHTSPEED_AUDIO_TRANSCODER` | `none` | Set to `ffmpeg` to enable audio transcoding; `none` or unset disables it. |
diff --git a/docs/documentation/using-lightspeed/chat-channels.md b/docs/documentation/using-lightspeed/chat-channels.md
index 7139d55f9..bcedb9f4f 100644
--- a/docs/documentation/using-lightspeed/chat-channels.md
+++ b/docs/documentation/using-lightspeed/chat-channels.md
@@ -146,13 +146,25 @@ Inbound media has explicit limits:
| Supported audio | 25 MiB |
At most eight attachments are admitted per message. Video processing is not
-supported. The selected model still needs to support the input type; channel
-admission does not add media capabilities to a model that lacks them.
-
-The connector prepares supported media into content-addressed storage for
-agent input. A bot does not need an execution environment just to converse
-or receive that input. Add a machine only when its task needs processes or
-the machine's filesystem.
+supported. The selected conversational model still needs to support admitted
+images and PDFs; channel admission does not add those capabilities to a model
+that lacks them.
+
+The connector prepares supported media into content-addressed storage. For
+voice messages, the conversation workflow transcribes authorized audio using
+the universe's **Speech-to-text** default before delivering the prepared
+text to the bot. Configure that route under **Models → Defaults**; the bot's
+conversational model does not need audio support. An unset or unusable speech
+route prevents voice preparation and is reported as a delivery failure.
+
+Transcription retains the original audio as provenance. Matching retries
+reuse the admitted transcription, and spoken text does not become a
+channel-management command such as `/activation`. Connector processes handle
+transport; they do not select models or obtain model credentials.
+
+A bot does not need an execution environment just to converse or receive
+prepared input. Add a machine only when its task needs processes or the
+machine's filesystem.
## Manage connection and conversation state
diff --git a/docs/documentation/using-lightspeed/cli.md b/docs/documentation/using-lightspeed/cli.md
index 3c206b8d6..bc4cae2ea 100644
--- a/docs/documentation/using-lightspeed/cli.md
+++ b/docs/documentation/using-lightspeed/cli.md
@@ -185,9 +185,11 @@ Choose a discovered route explicitly when starting a new session:
lightspeed chat --provider my-provider --api-kind openai:completions --model MODEL_ID
```
-Without those flags, a new session uses deployment defaults. Provider identity
-and API kind are fixed for each session; `/model` can choose another model
-within that route. The TUI footer shows the connection name and universe slug
+Without those flags, a new session uses its profile's model, if supplied, then
+the universe's agent-run default. Configure that default below or select a
+route explicitly. Provider identity and API kind are fixed for each session;
+`/model` can choose another compatible model within that route. Retained native
+state can restrict that choice. The TUI footer shows the connection name and universe slug
for saved or explicit runtime connections; generated local-development
connections omit this label. The footer also shows the current session ID,
shortening UUIDs while preserving readable IDs. If the universe slug is
@@ -241,6 +243,33 @@ and preserve other attachments and configuration. Repeating `--upload` creates
another workspace. Resuming with `chat -s SESSION_ID` alone retains the session's
existing attachments. Every `--session` option also accepts `-s`.
+## Configure universe model defaults
+
+Model defaults belong to the selected universe, separately from provider
+credentials and the CLI's saved connection:
+
+```bash
+lightspeed model defaults read
+lightspeed model defaults set agent-run --provider my-provider --api-kind openai:completions --model MODEL_ID
+lightspeed model defaults set speech-to-text --provider openai --api-kind openai:audio-transcriptions --model TRANSCRIPTION_MODEL_ID
+lightspeed model defaults clear speech-to-text
+```
+
+Use a model supported by the configured provider. These commands require the
+`models` method group. `set` and `clear` read the current revision before
+writing; use `--expected-revision REVISION` to require a revision you already
+read. All three commands support `--json`.
+
+An explicit session model takes precedence over a profile model, followed by
+the `agent-run` default. Without any of these, creation fails with
+`model_default_unset`. Changing or clearing defaults leaves existing sessions
+and admitted transcription jobs unchanged. A resumed chat uses its stored
+model. See [Models and credentials](models-and-credentials.md#choose-universe-defaults).
+
+`LIGHTSPEED_CHAT_PROVIDER` and `LIGHTSPEED_CHAT_MODEL` are retired and prevent
+runtime startup. Remove them from deployment configuration and save the
+intended routes with these commands.
+
## Administer keys and add Platform later
A deployment key with `deployment/api-keys` can issue narrower keys:
@@ -345,9 +374,46 @@ Resource operations accept `--json` for scripts. OAuth authorization instruction
are written to stderr so JSON stdout remains parseable. Profile export always
emits a reusable JSON document; profile list/read/import/check use human output
unless `--json` is requested. Configuration `put` replaces a complete document;
-omitted fields return to their defaults. It checks the revision read immediately
+omitted fields return to their defaults, except an omitted model preserves the
+existing session model. It checks the revision read immediately
before the write, or the explicit `--expected-revision` supplied by the caller.
+## Repair rejected context
+
+If a provider repeatedly rejects a particular user message or tool result,
+inspect the active context after the run has stopped:
+
+```bash
+lightspeed session context list SESSION_ID
+lightspeed session context list SESSION_ID --json
+```
+
+The list follows model order and includes entry IDs such as `item_12`, kinds,
+media handles when present, and bounded previews. Use the IDs actually
+returned for that session. Replace an entry with useful text, or redact it
+with a standard removed-by-operator placeholder:
+
+```bash
+lightspeed session context replace SESSION_ID item_12 "The command failed; its oversized diagnostic output was removed."
+lightspeed session context redact SESSION_ID item_15 item_16
+```
+
+Only active user messages and tool results can be replaced. A tool result
+keeps its call identity and position, so the provider still sees the call
+answered. Assistant messages, tool calls, and provider-native state cannot
+be edited this way. Replacement is refused while a run is active.
+
+Inspect each reported result: `replaced`, `unchanged`, `absent`, or `failed`.
+An entry that has already left active context is absent; a batch can report
+different outcomes for different entries. List the context again before
+continuing the conversation. These commands also support `--json`; API
+clients use `session/context/replace`.
+
+Redaction changes what future requests receive. Original events and blobs
+remain in retained history, so this is not a data-erasure operation. For a
+full context window rather than a specific rejected entry, use the
+[compaction controls](sessions-and-runs.md#manage-long-conversations).
+
## Provision environments from the CLI
A deployment administrator registers a provider controller and binds it to a
diff --git a/docs/documentation/using-lightspeed/models-and-credentials.md b/docs/documentation/using-lightspeed/models-and-credentials.md
index 5ae46bfa3..a6308e5bb 100644
--- a/docs/documentation/using-lightspeed/models-and-credentials.md
+++ b/docs/documentation/using-lightspeed/models-and-credentials.md
@@ -42,8 +42,9 @@ to use a subscription inside a machine.
## Connect an OpenAI-compatible provider
Compatible providers let Lightspeed use a service implementing OpenAI-style
-Responses or Chat Completions endpoints. Compatibility describes the request
-format; individual services and models can support different features.
+Responses, Chat Completions, or Audio Transcriptions endpoints. Compatibility
+describes the request format; individual services and models can support
+different features.
Open **Models → Add provider → OpenAI-compatible
provider**. Choose a **Provider** preset for DeepSeek, OpenRouter, Ollama, or
@@ -54,7 +55,7 @@ vLLM, or choose **Custom provider**. Configure:
| Custom provider ID | For a custom provider, a stable identifier used by model selections. Presets supply their IDs automatically; `deepseek` and `openrouter` select the corresponding compatibility rules. |
| Base URL | The API base URL advertised by the service, including its API path when required. |
| API key (optional) | The service's key, or leave it empty for an endpoint that does not require authentication. |
-| API kinds | The endpoints the service actually implements: Chat Completions, Responses, or both. The form defaults to Chat Completions. |
+| API kinds | The endpoints the service actually implements: Chat Completions, Responses, and/or Audio Transcriptions. The form defaults to Chat Completions. |
| Extra headers | Non-secret service-specific headers, if required. Authentication has its own field. |
Choose **Save provider**, check model discovery, and select a model in your
@@ -93,15 +94,71 @@ Chat Completions instead of the picker's preferred Responses route. Select
an API kind implemented by that provider.
For an existing session, change compatible model settings only while it is
-idle. The API kind is fixed for that session because its conversation is
-stored in the provider's native format. Create a new session to use another
-API kind. A profile change can select a different kind for future sessions.
+idle. Provider identity and API kind are fixed for that session because its
+conversation can contain provider-native state. Retained reasoning or native
+compaction state can also restrict model changes within that route. Create a
+new session for an incompatible model or route. A profile change can select
+a different route for future sessions.
Leave optional reasoning and generation settings unset until you need them
and have verified the chosen model supports them. A provider can accept a
model name while refusing an incompatible parameter.
-## Understand deployment defaults
+
+
+## Choose universe defaults
+
+Under **Models → Defaults**, an Operator or Admin can choose, change, or clear
+two independent defaults:
+
+| Default | Used for |
+| --- | --- |
+| **Agent runs** | New sessions whose setup and profile leave the model unset. |
+| **Speech-to-text** | Transcription jobs without an explicit model, including web dictation and channel voice messages. |
+
+Choose **Choose model** or **Change**, select a discovered model or enter its
+provider, API kind, and model name, then choose **Save default**. Manual models
+do not need to appear in discovery. Adding a provider credential alone does
+not choose a default.
+
+For a new session, an explicit session model takes precedence over the
+profile's model, followed by the universe's **Agent runs** default. If all
+three are unset, creation fails with `model_default_unset`; Lightspeed does
+not assume an OpenAI route. An explicit model works without a universe default.
+
+The resolved model is saved in the session. Changing or clearing a default
+does not retarget existing sessions, forks, clones, or admitted work. Applying
+a profile or replacing an existing session's configuration with the model
+omitted preserves that session's current model. Other configuration fields
+retain their replacement semantics.
+
+The [CLI](cli.md#configure-universe-model-defaults) and
+`models/defaults/read` and `models/defaults/put` APIs manage the same defaults.
+Updates check a revision; if someone changes the defaults first, reload and
+review their selection before saving again.
+
+## Configure speech-to-text
+
+To use web dictation or channel voice messages, configure a provider that
+supports `openai:audio-transcriptions`, then choose its transcription model
+under **Models → Defaults → Speech-to-text**. The built-in OpenAI provider and
+compatible endpoints use this protocol. A compatible endpoint uses its saved
+URL, headers, and authentication, including an explicitly credentialless
+configuration; it does not fall back to OpenAI if that setup fails.
+
+Choose a discovered speech model or enter the exact model name manually.
+Some providers omit transcription models from discovery. The speech default
+is independent of **Agent runs**, so the conversational model does not need
+to accept audio. Setting only a conversational default leaves dictation
+unavailable.
+
+The microphone in the [session composer](sessions-and-runs.md#dictate-a-message)
+reports why dictation is unavailable. Check the selected speech route and its
+credential, browser recording support, microphone permission, and a secure
+browser origin. [Channel voice messages](chat-channels.md#understand-replies-and-media)
+use the same universe default before delivering text to the bot.
+
+## Understand deployment credentials
An operator can supply deployment-level OpenAI or Anthropic credentials.
Those are fallback credentials when the universe has no corresponding model
@@ -113,12 +170,10 @@ using the deployment's credential. Removing the built-in provider record
allows fallback again. Consider that difference when rotating or disabling
keys.
-The runtime defaults to provider `openai` and API kind `openai:responses`.
-`LIGHTSPEED_CHAT_PROVIDER` and `LIGHTSPEED_CHAT_MODEL` change the default
-provider ID and model name. The runtime's default API kind remains Responses,
-so setting an Anthropic provider ID alone does not create an Anthropic route.
-Select the full route in a profile when using Anthropic or a compatible
-service. Exact deployment settings are in the
+`LIGHTSPEED_CHAT_PROVIDER` and `LIGHTSPEED_CHAT_MODEL` are retired. The runtime
+refuses startup if either is set. Remove them and configure each universe's
+model defaults instead. Provider credentials and transport settings remain
+deployment options, as listed in the
[environment-variable reference](../reference/environment-variables.md).
The API also supports model OAuth records that refer to a suitable stored
@@ -143,6 +198,8 @@ Verify the input types used by your actual tasks.
| Symptom | What to check |
| --- | --- |
+| Creating a session reports `model_default_unset` | Choose the universe's Agent runs default, or select a model explicitly in the session or profile. |
+| Runtime startup rejects `LIGHTSPEED_CHAT_PROVIDER` or `LIGHTSPEED_CHAT_MODEL` | Remove the retired variables and configure universe defaults through Models or the CLI. |
| The provider connects but the model is missing | Refresh discovery or enter the exact route manually. Confirm that the credential can access that model. |
| Authentication fails despite a deployment key | Check for an existing universe provider record. Disabled or unusable records block fallback. |
| A compatible service is unreachable | Test reachability from the gateway and session-worker network, and check the base URL and API path. |
diff --git a/docs/documentation/using-lightspeed/profiles-and-instructions.md b/docs/documentation/using-lightspeed/profiles-and-instructions.md
index 90db20172..818369e25 100644
--- a/docs/documentation/using-lightspeed/profiles-and-instructions.md
+++ b/docs/documentation/using-lightspeed/profiles-and-instructions.md
@@ -136,15 +136,17 @@ lightspeed session profile apply "" --profile release-reviewer
```
The API equivalent is `session/profiles/apply`. The session must be open with
-no active or queued runs. Its API kind is fixed for the lifetime of the
-session, so the applied profile must use that same kind. Create a new session
-when changing API kinds.
+no active or queued runs. Its provider identity and API kind are fixed for the
+lifetime of the session. An explicit profile model must use that same route
+and be compatible with retained native state. An omitted model preserves the
+session's current model; it does not resolve the universe default again.
+Create a new session for an incompatible route or model.
Applying a profile is not a deep merge of every field:
| Profile content | Effect on an existing session |
| --- | --- |
-| `config` present | Replaces the session configuration as a whole. Include the capabilities and attachments you intend to retain. |
+| `config` present | Replaces the session configuration, preserving the current model if the new model is omitted. Include the capabilities and attachments you intend to retain. |
| `config` absent | Leaves the current configuration in place. |
| `instructions` present or absent | Replaces or clears the profile instruction layer. Sourced prompt files follow the resulting VFS setup. |
| The active environment is no longer attached | Clears the active environment. |
diff --git a/docs/documentation/using-lightspeed/sessions-and-runs.md b/docs/documentation/using-lightspeed/sessions-and-runs.md
index 721c13747..63e73818e 100644
--- a/docs/documentation/using-lightspeed/sessions-and-runs.md
+++ b/docs/documentation/using-lightspeed/sessions-and-runs.md
@@ -37,6 +37,7 @@ use the earlier conversation and its linked files. Starting a new session
from the same profile gives you a fresh conversation; workspace attachments may
still point to the same shared files.
+
## Share a session with the universe
New standalone sessions are private: through the Platform, only their creator
@@ -129,6 +130,57 @@ calls while leaving the retained transcript available to inspect. It does
not mean that the agent will reproduce every earlier detail from memory; keep
important source material in files it can read again.
+## Manage long conversations
+
+New sessions use automatic standalone compaction by default. It reduces older
+conversation at safe boundaries between model turns, including during a run.
+When input capacity is known, the default trigger is 80% of that capacity;
+otherwise Lightspeed can recover when the provider reports a full context
+window. Recovery resumes the same run with completed tool results retained.
+It can still fail when the protected input is too large or its bounded
+attempts cannot make enough room.
+
+In the profile or idle session's model setup, open **Customize run controls →
+Context compaction**. **Engine default** and **Engine managed standalone**
+enable standalone compaction; **Provider triggered** uses supported OpenAI
+Responses or Anthropic Messages generation compaction. **Disabled** turns off
+automatic compaction and recovery. **Input limit tokens** overrides usable
+input capacity; leave it blank to use reported capacity or error-driven recovery.
+Session settings show the effective mode, threshold, queued or pending state,
+and recovery attempts.
+
+Older sessions whose saved setup omitted compaction keep their historical
+disabled behavior until configuration is replaced. Applying setup with
+**Engine default** adopts the current default. Explicit Disabled remains
+disabled. An API caller can request one `session/context/compact` operation
+in any mode; active work queues it until a safe boundary.
+
+Compaction can summarize away details, so keep important source material in
+workspaces. It preserves retained history and original blobs. The
+[context guide](../how-it-works/context-and-storage.md#compact-the-active-conversation)
+explains provider behavior and protected input. For a specific rejected entry,
+use [context repair](cli.md#repair-rejected-context) instead.
+
+## Dictate a message
+
+With a [speech-to-text default](models-and-credentials.md#configure-speech-to-text)
+configured, choose the composer's microphone (**Dictate message**) and allow
+microphone access. Choose **Stop recording** to transcribe and insert the text
+into your draft. Review or edit it, then send normally. Recording stops
+automatically after ten minutes; dictation audio is limited to 25 MiB.
+
+You can keep editing while transcription runs. Recording or stopping alone
+does not start a run. If you choose **Send** or press **Enter** while recording
+or transcribing, the composer waits for the transcript and then submits the
+message. During an active run, that action queues the next run; choosing
+steering instead sends it to the active run when ready.
+
+**Cancel dictation** or **Esc** discards the recording or pending transcription
+and clears any send waiting for it. Failures preserve your existing draft;
+use **Retry transcription** when available. The microphone control explains
+unavailable recording or model configuration. Browsers require microphone
+permission and a secure origin, such as HTTPS or localhost.
+
## Find and change a session
Open **Filter sessions** in the session list. Under **Include**, select
@@ -146,9 +198,10 @@ A session keeps its configured provider identity and API kind for its entire
lifetime, including before its first run. You can switch model names within
that route, such as between OpenAI models, and switch back later. Changing
provider or API kind requires a new session because conversation context may
-contain provider-native opaque data. An aggregator such as OpenRouter counts
-as one configured provider; model changes within it are allowed without
-checking the underlying model vendor.
+contain provider-native opaque data. Retained reasoning or native compaction
+state can also prevent an incompatible model change within that route. An
+aggregator such as OpenRouter counts as one configured provider; the underlying
+vendor name alone does not determine compatibility.
Existing ordinary sessions keep the setup they received at creation. Editing
their source profile does not update them automatically. See
diff --git a/docs/documentation/using-lightspeed/subagents-and-federation.md b/docs/documentation/using-lightspeed/subagents-and-federation.md
index 1366bc074..3c88a9c40 100644
--- a/docs/documentation/using-lightspeed/subagents-and-federation.md
+++ b/docs/documentation/using-lightspeed/subagents-and-federation.md
@@ -123,6 +123,28 @@ documents referenced in the child's answer can also pass back to the parent.
A child returns one run's result and closes automatically. Delegate a new
task when more work is needed.
+## Return file attachments
+
+A child can deliver a file even when its parent cannot access the child's
+workspace. Ask it to save the deliverable, call `vfs_reference`, and cite the
+returned `file:` link in its successful final answer. Both joined `agent_run`
+results and spawned/awaited results carry descriptors for the known attachments
+referenced in that answer. Intermediate files that are not cited stay with
+the child. A parent handing the result to its own parent follows the same rule.
+
+The descriptor identifies immutable bytes. A later workspace edit or deletion
+does not retarget the attachment, and the receiving session retains the
+referenced content. A digest written only in prose does not establish an
+attachment or keep its bytes alive. Access to downloads still follows the
+universe's blob permissions.
+
+Image and PDF media use the existing `media:` links and native-media admission
+rules. File references preserve a downloadable attachment without loading it
+as model input. Either `[label](file:HANDLE)` or image Markdown for a returned
+image handle can select an attachment for handoff. See
+[Workspaces and skills](workspaces-and-skills.md#share-an-immutable-file-in-an-answer)
+for creating references and capturing outputs from a machine.
+
## Bound the delegation tree
The default limits are depth `2`, `16` total descendants, `4` concurrent open
diff --git a/docs/documentation/using-lightspeed/tools-and-mcp.md b/docs/documentation/using-lightspeed/tools-and-mcp.md
index 5a6543df2..82573cb39 100644
--- a/docs/documentation/using-lightspeed/tools-and-mcp.md
+++ b/docs/documentation/using-lightspeed/tools-and-mcp.md
@@ -173,6 +173,11 @@ authorize tool-call egress. See the
## See images and documents from tools
+For a downloadable file version, ask the agent to call `vfs_reference` and
+cite the returned `file:` link. This explicitly publishes an attachment;
+writing or reading a file alone does not. See
+[Share an immutable file in an answer](workspaces-and-skills.md#share-an-immutable-file-in-an-answer).
+
Reading a PNG, JPEG, GIF, WebP, or PDF can return the file to the model as
media. MCP servers can also supply image blocks and embedded PDFs. The
transcript shows those items as thumbnails or document links, and the agent
@@ -181,12 +186,23 @@ their results.
Lightspeed accepts up to eight media items per tool result, each at most
10 MiB. Unsupported types and oversized items produce an explanatory note.
-The bytes are passed through without resizing or conversion, so the selected
-model must support the format. A text-only model receives a note instead;
-provider-specific limits or refusals can still reject a model request. Test a
-representative document before choosing a model for a document-heavy task.
-Claude Opus 5 has refused some tool-produced PDF follow-ups in Lightspeed's
-live tests; verify that combination with the documents your task will use.
+The original bytes remain stored for previews and downloads. Before model
+requests, Lightspeed normalizes oversized images to at most 2,000 pixels per
+side and a 3.75 MiB image-byte budget. It can re-encode them as JPEG; images
+already within the limits pass through unchanged. A resized image's model
+announcement includes its original and displayed dimensions. PDFs are not
+resized or converted.
+
+Across a request, media is limited to 100 items and 24 MiB of encoded payload.
+When needed, older media is replaced with omission notes that retain its
+`media:` handle. Those notes change that request, not stored content or the
+retained transcript. The model may therefore remember an attachment's handle
+without receiving its pixels or document bytes on every turn.
+
+The selected model must still support the format. A text-only model receives
+a note instead; provider-specific limits or refusals can still reject a
+request. Test representative documents with the intended model. To withdraw
+a problematic active entry, use [context repair](cli.md#repair-rejected-context).
## Require approval for tool calls
diff --git a/docs/documentation/using-lightspeed/workspaces-and-skills.md b/docs/documentation/using-lightspeed/workspaces-and-skills.md
index a84d8fc4c..6d64856f1 100644
--- a/docs/documentation/using-lightspeed/workspaces-and-skills.md
+++ b/docs/documentation/using-lightspeed/workspaces-and-skills.md
@@ -20,6 +20,23 @@ Create or select a workspace under **Workspaces**. **New file** accepts a path
relative to that workspace, and creates directories in the path as needed.
Open a file, edit its contents, and choose **Save**.
+To bring existing files into the workspace, open **Workspace actions** or a
+folder's actions menu and choose **Upload files** or **Upload folder**. You
+can also drop files and folders onto the workspace or a destination folder.
+An upload supports up to 32 MiB of file content and 10,000 file/directory
+entries. This is a one-time copy; later local edits are not synchronized.
+
+Uploads with no conflicts publish directly. If files already exist, review
+the collisions and choose whether to replace them or skip them and upload the
+remaining items. A file's **Upload replacement…** action targets that file.
+Publication checks the workspace revision, so concurrent edits require
+retrying against the latest tree rather than silently overwriting them.
+
+The same menus provide **New folder** and **Rename…** for files and folders.
+Choose **Download** for one file or **Download as ZIP** for a folder or the
+whole workspace. Uploaded binary files can be previewed where supported and
+downloaded without converting them to text.
+
In a profile's **Virtual File System: Files, Instructions, Skills** section,
enabling VFS turns on **Prompt loading** and **Skill discovery**. Choose **Add
workspace** under **Workspace attachments** and configure its path and access.
@@ -46,6 +63,38 @@ snapshot with read-only access. For tasks
that must produce independent artifacts, create separate workspaces or use
different output paths deliberately.
+## Share an immutable file in an answer
+
+After writing a deliverable, ask the agent to obtain a reference with
+`vfs_reference` and include the returned link in its answer. For example:
+
+```text
+Save the finished notes to /workspace/release-notes.md. Use vfs_reference
+for that file, then include the returned file link in your final answer.
+```
+
+The tool records an attachment to that exact file version without reading
+its contents or changing it. Ordinary reads, writes, edits, and transfers do
+not create these file attachments automatically. A returned `file:` handle
+can be used as `[Release notes](file:HANDLE)`; use the actual handle from the
+tool. For an image, `` displays it inline. Image syntax
+for another file type falls back to a file link.
+
+The link keeps opening the referenced bytes after the workspace file is
+renamed, overwritten, deleted, or detached. The blob viewer can also link to
+the source workspace path; that navigation opens the current path, which may
+have changed since the attachment was created. A file reference to an image
+does not itself load the image into model context.
+
+For an output on an execution environment, first use `vfs_capture`, then
+reference the captured file. The reference tool can also use a returned
+snapshot reference and a path inside it if workspace publication encountered
+a revision conflict. See [VFS transfer](../environments/vfs-transfer.md).
+
+To pass an attachment from a sub-agent to its parent, the child must cite it
+in its successful final answer. Simply creating it is not enough. See
+[Sub-agents and federation](subagents-and-federation.md#return-file-attachments).
+
## Add project instructions
In **Workspaces → Release notes → New file**, enter
@@ -257,9 +306,17 @@ catalog during that run. Switching environments removes the old machine's
catalog; the new machine is scanned at the next eligible refresh.
The machine must be online and support filesystem scanning. Discovery does
-not wake it. If a scan fails or exceeds its limits, the catalog reports that
-source as unavailable with diagnostics. Check the machine's state, root
-paths, and read permissions before trying again.
+not wake it. If a scan fails or is incomplete, Lightspeed can retain the last
+successful catalog for the same environment and discovery scope, marking it
+stale with a warning in the API. The model's menu continues to list discovered
+paths; reading a skill still accesses the machine's current file.
+
+Fallback is discarded when the environment or configured scope changes,
+including roots, working directory, or access. Removing the attachment also
+removes its advertised paths. Without a matching prior observation, the source
+is unavailable. Check the machine's state, roots, and permissions, then retry
+discovery at the next eligible refresh. A stale listing does not prove that
+a file is currently reachable.
## Update files and handle concurrent edits
diff --git a/docs/roadmap/p165-documentation.md b/docs/roadmap/p165-documentation.md
index 1c60ee653..59b07116c 100644
--- a/docs/roadmap/p165-documentation.md
+++ b/docs/roadmap/p165-documentation.md
@@ -457,6 +457,30 @@ for 93 shell and 23 JSON examples without executing their operations.
Documentation whitespace checks passed. No live service or credentialed tests
were run.
+### Recent runtime and workspace documentation refresh
+
+Updated the manual against the implemented model-default, transcription,
+context-reliability, and workspace changes on 2026-10-04:
+
+- Replaced retired deployment model settings with universe defaults and CLI
+ setup, including profile precedence and existing-session preservation.
+- Explained standalone compaction defaults, active-run scheduling, bounded
+ context-limit recovery, native provider behavior, and historical policy.
+- Added speech configuration, dictation controls, channel voice preparation,
+ and the standalone transcription integration flow.
+- Documented context repair, request-time image normalization and media
+ budgeting, while distinguishing active context from retained history.
+- Added workspace upload/download procedures, immutable file references,
+ selected sub-agent attachment handoff, and stale skill-catalog behavior.
+- Corrected environment catalog contents and short-reference semantics, and
+ pointed upgrade guidance to authoritative release schema metadata.
+
+The root README and generated contracts are unchanged. `npm run check:docs`
+passed all nine adapter tests, Astro diagnostics, and the production build,
+including verification of 55 HTML/Markdown pages, 15 diagrams, links, anchors,
+assets, search, sitemap, and Markdown exports. `git diff --check` passed.
+No live services or credentialed tests were needed.
+
### Screenshot refresh after the manual review
Refreshed all seven existing screenshots with Playwright against the current
diff --git a/platform/backend/src/routes/gateway-compaction.test.ts b/platform/backend/src/routes/gateway-compaction.test.ts
new file mode 100644
index 000000000..2c62ae9f7
--- /dev/null
+++ b/platform/backend/src/routes/gateway-compaction.test.ts
@@ -0,0 +1,51 @@
+import { Hono } from "hono";
+import { afterEach, beforeEach, expect, it, vi } from "vitest";
+import type { ApiVariables, AppContext } from "../context.js";
+import { gatewayRoutes } from "./gateway.js";
+
+const access = vi.hoisted(() => ({ role: "contributor", owner: "user" }));
+vi.mock("./universes.js", () => ({
+ universeForSession: vi.fn(async () => ({
+ universe: { lightspeedUniverseId: "universe", gatewayUrl: "https://engine.example/rpc" },
+ role: access.role,
+ member: { userId: "user", role: access.role },
+ })),
+}));
+let requests: Array<{ method: string; params: unknown }>;
+beforeEach(() => {
+ access.role = "contributor";
+ access.owner = "user";
+ requests = [];
+ vi.stubGlobal("fetch", vi.fn(async (_url: unknown, init: RequestInit) => {
+ const rpc = JSON.parse(String(init.body));
+ requests.push(rpc);
+ const session = { id: "session", status: "idle", access: { visibility: "restricted", createdBy: { kind: "actor", id: access.owner } } };
+ return Response.json({ id: rpc.id, result: { result: { session }, notifications: [] } });
+ }));
+});
+afterEach(() => vi.unstubAllGlobals());
+async function compact() {
+ const app = new Hono<{ Variables: ApiVariables }>();
+ app.use("*", async (c, next) => {
+ c.set("session", { user: { id: "user" } } as ApiVariables["session"]);
+ await next();
+ });
+ app.route("/", gatewayRoutes({ env: { lightspeedApiUrl: "https://engine.example/rpc", lightspeedApiKey: "lsk_fixture" } } as AppContext));
+ return app.request("/universe/sessions/session/context/compact", { method: "POST" });
+}
+it("forwards a contributor's manual compaction to the core API", async () => {
+ const response = await compact();
+ expect(response.status).toBe(200);
+ expect(await response.json()).toMatchObject({ session: { id: "session" } });
+ expect(requests).toContainEqual(expect.objectContaining({ method: "session/context/compact", params: { sessionId: "session" } }));
+});
+it("rejects viewers before requesting compaction", async () => {
+ access.role = "viewer";
+ expect((await compact()).status).toBe(403);
+ expect(requests.some((request) => request.method === "session/context/compact")).toBe(false);
+});
+it("does not compact another user's private session", async () => {
+ access.owner = "someone-else";
+ expect((await compact()).status).toBe(404);
+ expect(requests.some((request) => request.method === "session/context/compact")).toBe(false);
+});
diff --git a/platform/backend/src/routes/gateway.ts b/platform/backend/src/routes/gateway.ts
index 2f854ce16..86a435a48 100644
--- a/platform/backend/src/routes/gateway.ts
+++ b/platform/backend/src/routes/gateway.ts
@@ -498,6 +498,17 @@ export function gatewayRoutes(ctx: AppContext) {
});
});
+ app.post("/:id/sessions/:sessionId/context/compact", async (c) => {
+ const access = await universeForSession(ctx, c, c.req.param("id"));
+ if (!access) return c.json({ error: "not found" }, 404);
+ return withGateway(c, async () => {
+ const response = await engineClientFor(ctx, access).call("session/context/compact", {
+ sessionId: c.req.param("sessionId"),
+ });
+ return c.json(response.result);
+ });
+ });
+
/// Closing is a lifecycle transition that retains session history.
/// `force=true` also cancels active/queued work.
app.post("/:id/sessions/:sessionId/close", async (c) => {
diff --git a/platform/web/src/components/bot/detail.tsx b/platform/web/src/components/bot/detail.tsx
index 58e783856..70c6ae1ae 100644
--- a/platform/web/src/components/bot/detail.tsx
+++ b/platform/web/src/components/bot/detail.tsx
@@ -1,3 +1,4 @@
+import { useSessionCompaction } from "@/lib/sessions/compaction";
import { useState } from "react";
import { useMutation, useQuery, useQueryClient } from "@tanstack/react-query";
import {
@@ -409,6 +410,8 @@ function ConversationMenu({
`/api/v1/universes/${universeId}/sessions/${encodeURIComponent(sessionId)}`,
),
});
+ const compaction = useSessionCompaction(universeId, sessionId, session.data);
+ const canCompact = useActionPermissions(universeId).can("control_session");
const reset = useMutation({
mutationFn: () =>
api(
@@ -427,7 +430,8 @@ function ConversationMenu({
variant="tab"
sessionId={sessionId}
metadata={session.data?.metadata}
- pending={reset.isPending}
+ onCompact={canCompact && session.data && session.data.status !== "closed" ? compaction.compact : undefined}
+ compactionLabel={compaction.label}
open={{
label: "Open on the Sessions page",
href: `/u/${slug}/sessions/${encodeURIComponent(sessionId)}`,
diff --git a/platform/web/src/components/session/session-actions-menu.test.tsx b/platform/web/src/components/session/session-actions-menu.test.tsx
index a314d448b..bd06a5b93 100644
--- a/platform/web/src/components/session/session-actions-menu.test.tsx
+++ b/platform/web/src/components/session/session-actions-menu.test.tsx
@@ -11,11 +11,11 @@ vi.mock("@/components/ui/dropdown-menu", () => {
const pass = ({ children }: { children?: ReactNode }) =>