Files
b875c537be fix(codex): restore installed skills with catalog-backed models (#140422)
* fix(codex): restore installed skills with catalog-backed models

* fix(codex): keep skill catalog refreshable on live incognito threads

Moving the installed-skill catalog into thread developer instructions made
every catalog refresh look like a generic policy change. On a live incognito
thread the warm-reuse guard compared the whole creation string, so editing a
skill description and sending another message threw
CodexIncognitoPolicyChangeError before turn/start.

The catalog is now a separate lifecycle input (`skillsInstructions`). The
request builder joins it onto the wire developer carrier after the generic
policy, so catalog-backed models still receive it and persistent threads keep
cold-resuming through the existing policy handoff. The retained ephemeral
record splits the immutable generic policy from the catalog last delivered;
incognito warm reuse compares only the generic policy (unchanged guard, still
fail-closed) and delivers a changed catalog with a `thread/inject_items`
developer message, tracked so an unchanged catalog is not re-sent.

Docs: the runtime guide, system-prompt concept, and token-use reference still
said the catalog rode turn-scoped collaboration instructions; they now match
the thread carrier and the incognito refresh.

Codex protocol behavior checked in sibling ../codex (main 459a79e):
core/src/context/world_state/collaboration_mode.rs:30-42 (catalog messages
replace caller collaboration instructions), app-server-protocol v2 thread.rs
(thread/settings/update has no developerInstructions), turn_processor.rs
thread_inject_items, context-fragments/src/additional_context.rs (1,000-token
cap rules out additionalContext as the catalog carrier).

* test(codex): cover catalog refresh on reused persistent threads

Adds the missing half of the reused-thread coverage. The incognito case
already proved an edited catalog reaches a live ephemeral thread through
thread/inject_items. This covers the persistent case: the catalog rides the
thread developer carrier, so it participates in the thread-config
fingerprint (thread-fingerprints.ts:121). An edited skill therefore
invalidates warm reuse and cold-resumes the SAME thread with the new
catalog, instead of either losing the conversation or answering from the
stale catalog.

Mutation-checked: dropping options.skillsInstructions from the join in
buildCodexThreadConfiguration fails both this test and the incognito
two-turn test, so neither assertion is vacuous.

Also adapts the native skills test to upstream's changed shutdown contract.
closeAndWait() returns CodexAppServerCloseResult rather than boolean since
bcef3b3526 (#139657); the assertion now uses the same
toMatchObject({ exited: true }) form as the other native tests
(thread-lifecycle.native.test.ts:142, settled-turn-finalizer.native.test.ts:57).

* fix(codex): break app-server import cycle and register skills test lanes

Extract joinPresentSections into a zero-import leaf module so
thread-requests.ts no longer imports run-attempt-state.ts, which closed a
10-file strongly connected component in extensions/codex/src/app-server.
The merge base is clean, so the cycle was introduced by this branch.

Register the two new skills test files in the vitest project that already
owns the comparable run-attempt.native-hook-relay test, so full-suite file
ownership is exactly one project per file.

* fix(codex): keep refreshed incognito skill catalogs across compaction

An incognito thread carries its skill catalog in creation-time native
developer instructions. A mid-session catalog change is delivered as a
client-authored developer message, but Codex remote compaction drops those
messages unless retain_client_developer_messages is enabled (off by default)
and rebuilds initial context from the unchanged native instructions. The host
had already recorded the refreshed catalog as delivered, so later turns
skipped reinjection and the conversation silently reverted to the catalog it
started with: edited skills kept stale descriptions and removed skills
stayed visible.

Track the creation-time catalog separately from the current one and
re-deliver the current catalog after every compaction, including compaction
inside a turn, via the existing onContextCompacted hook. An unchanged catalog
still sends nothing, because compaction rebuilds it natively.

* test(codex): cover catalog restoration across native compaction

Drive the packaged Codex app-server binary through a real compaction on a
live incognito thread: turn two refreshes the catalog, turn three exceeds the
resolved auto-compact limit, and turn four must still show the model the
current catalog.

The continuation request after summarization is asserted to carry the
creation-time catalog and not the refresh, which reproduces the reversion
natively. The creation-time catalog stays in the ephemeral thread's immutable
developer instructions and is re-emitted on every request, so restoration is
pinned by catalog ordering rather than absence.

Pin nativeSkillsInstructions in the ephemeral policy binding assertions so a
refresh is shown never to rewrite what compaction will restore.

* fix(codex): re-deliver the catalog after a failed post-compaction restore

A failed or aborted restore was logged and dropped while the retained record
still named the refreshed catalog as delivered. Warm reuse compares that
record, so every later turn skipped the refresh and the conversation stayed
stranded on the creation-time catalog for the rest of its life.

Record what compaction actually restored when the injection fails, so the
next turn delivers the current catalog again.

* docs(codex): state catalog restore timing precisely

The restore is issued as soon as a compaction completes, but that turn's own
continuation request may already have been built, so the guaranteed floor is
the following request rather than the remainder of the compacting turn.

* fix(codex): keep incognito creation policy across standalone compaction

Standalone compaction re-retained the live thread without its ephemeral
policy, so the record lost the creation-time generic policy and the catalog
it had delivered. The next turn then read the live incognito thread as policy
drift, and the discarded catalog refresh was never re-sent.

Carry the ephemeral policy back through the re-retain and record the catalog
compaction actually restored, so an edited or withdrawn skill is delivered
again on the next turn instead of reverting for the rest of the conversation.

* fix(codex): revert the catalog on the incognito compaction path

Standalone compaction only claims and re-retains a thread for non-incognito
session keys, so the previous reversion never ran for the conversations it was
written for. An incognito thread keeps its own live subscription, leaving the
retained record still naming the discarded refresh as delivered.

Correct the retained record in place after a completed compaction, and cover
both session kinds so an incognito key can no longer be masked by an ordinary
one in the regression.

* fix(codex): preserve trajectory helper import after rebase

* test(codex): budget the compaction history read in skill catalog tests

The three post-compaction catalog re-delivery tests never observed
`thread/inject_items`, which read as the repair being broken. It was not.

Compaction runs the real `before_compaction` hook, which performs an actual
mirrored-session history read; its first worker start costs several seconds.
`createParams` defaults `timeoutMs` to 5s, so the attempt execution budget
expired inside that read, aborted projection, and the `onContextCompacted`
restore never ran -- reproducing precisely the symptom these tests assert
against.

Give these turns the same budget the sibling compaction tests already use in
`run-attempt.skills.native.test.ts` and `run-attempt.hooks.test.ts`. Every
assertion is unchanged.

* fix(codex): preserve parent-only skills on managed connections

* test(codex): validate managed instructions before collection

---------

Co-authored-by: Rook <rook@openclaw.invalid>
Co-authored-by: Colton Harris <287090614+Colton-Harris@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
2026-09-24 14:59:15 -07:00

16 KiB

summary, read_when, title
summary read_when title
How OpenClaw builds prompt context and reports token usage + costs
Explaining token usage, costs, or context windows
Debugging context growth or compaction behavior
Token use and costs

OpenClaw tracks tokens, not characters. Tokens are model-specific, but most OpenAI-style models average ~4 characters per token for English text.

How the system prompt is built

OpenClaw assembles its own system prompt on every run. It includes:

  • Tool list + short descriptions
  • Skills list (metadata only; instructions load on demand with read). Native Codex turns on the managed bundled app-server get the compact skills block in parent-local model request instructions. Connections without that relay use thread developer instructions; other harnesses get it in the normal prompt surface. Bounded by skills.limits.maxSkillsPromptChars, with optional per-agent override at agents.entries.*.skillsLimits.maxSkillsPromptChars.
  • Self-update instructions
  • Workspace + bootstrap files (AGENTS.md, SOUL.md, IDENTITY.md, USER.md, BOOTSTRAP.md when new, plus MEMORY.md when present). Large injected files are truncated by agents.defaults.bootstrapMaxChars (default: 20000); total bootstrap injection is capped by agents.defaults.bootstrapTotalMaxChars (default: 60000).
    • Native Codex turns do not paste raw MEMORY.md when memory tools are available for that workspace; they get a small memory pointer in parent-local request instructions instead and use memory tools on demand. If tools are disabled, memory search is unavailable, or the active workspace differs from the agent memory workspace, MEMORY.md falls back to the normal bounded turn-context path.
    • Lowercase root memory.md is never injected. It is legacy repair input for openclaw doctor --fix, which migrates it into MEMORY.md.
    • memory/*.md daily files are not part of the normal bootstrap prompt; they stay on-demand via memory tools on ordinary turns. Reset/startup model runs can prepend a one-shot startup-context block with recent daily memory for that first turn, controlled by agents.defaults.startupContext. Bare chat /new and /reset are acknowledged without invoking the model.
    • Post-compaction AGENTS.md excerpts require explicit agents.defaults.compaction.postCompactionSections opt-in; plugins can add other context through before_prompt_build.
  • Time (UTC + user timezone)
  • Reply tags + heartbeat behavior
  • Runtime metadata (host/OS/model/thinking)

See the full breakdown in System Prompt.

When documenting credentials or auth snippets, use the Secret Placeholder Conventions to avoid secret-scanner false positives in docs-only changes.

What counts in the context window

Everything the model receives counts toward the context limit:

  • System prompt (all sections above)
  • Conversation history (user + assistant messages)
  • Tool calls and tool results
  • Attachments/transcripts (images, audio, files)
  • Compaction summaries and pruning artifacts
  • Provider wrappers or safety headers (not visible, but still counted)

Runtime-heavy surfaces have their own explicit caps under agents.defaults.contextLimits (per-agent overrides under agents.entries.*.contextLimits):

Key Purpose
memoryGetMaxChars Max characters memory_get returns before truncation.
postCompactionMaxChars Max characters retained from AGENTS.md during post-compaction refresh.

These are bounded runtime excerpts and injected runtime-owned blocks, separate from bootstrap limits, startup-context limits, and skills prompt limits.

OpenClaw derives the live tool-result cap from the effective model context window: 16000 chars below 100K tokens, 32000 chars at 100K+ tokens, 64000 chars at 200K+ tokens. The runtime context-share guard also caps a single tool result at 30% of the context window.

Large provider windows are not enabled automatically when they materially change cost or latency. For example, direct OpenAI GPT-5.5 and GPT-5.6 models publish a 1050000 token total window, but OpenClaw defaults their active runtime budget to 272000 tokens. The opt-in 922000 input budget reserves the full 128000 output allowance, and OpenAI applies higher long-context pricing to the entire request once input exceeds 272000 tokens. See OpenAI context window defaults.

For images, OpenClaw downscales transcript/tool image payloads before provider calls. Tune with agents.defaults.imageMaxDimensionPx (default: 1200):

  • Lower values reduce vision-token usage and payload size.
  • Higher values preserve more visual detail for OCR/UI-heavy screenshots.

For a practical breakdown (per injected file, tools, skills, and system prompt size), use /context list or /context detail. See Context.

How to see current token usage

In chat:

  • /status -> emoji-rich status card with the session model, context usage, last response input/output tokens, and cost from recorded billing or local pricing for the active model.
  • /usage off|tokens|full -> appends a per-response usage footer to every reply. Persists per session (stored as responseUsage).
    • /usage reset (aliases: inherit, clear, default) clears the session override so it re-inherits the configured default.
    • /usage tokens shows turn token/cache details.
    • /usage full shows compact model/context/cost details. Cost comes from a recorded amount or usage metadata with local pricing for the active model. Custom messages.usageTemplate layouts can include token/cache fields.
  • /usage cost -> local cost summary from OpenClaw session logs.

Other surfaces:

  • Control UI: the working indicator and completed-run recap show cumulative output tokens for that run, including its model calls across tool use and retries. Counts update when the runtime reports completed-response usage, not on every streamed text fragment. Reloading an active run restores its latest count. This counter excludes input tokens and is separate from the composer context-window meter and persisted billing summaries.
  • TUI/Web TUI: /status and /usage are supported.
  • CLI: openclaw status --usage and openclaw channels list show normalized provider quota windows (X% left, not per-response costs). Usage-window providers, checked against 2026.9.3: Claude (Anthropic), ClawRouter, Copilot (GitHub), DeepSeek, MiniMax, OpenAI, OpenRouter, Venice, xAI, Xiaomi, Xiaomi Token Plan, and z.ai. Provider plugins supply these snapshots, so an installed plugin can add one.

Usage surfaces normalize common provider-native field aliases before display. For OpenAI-family Responses traffic, that includes both input_tokens/output_tokens and prompt_tokens/completion_tokens, so transport-specific field names do not change /status, /usage, or session summaries. Gemini CLI usage is normalized too: the default stream-json parser reads assistant message events, and stats.cached maps to cacheRead, with stats.input_tokens - stats.cached used when the CLI omits an explicit stats.input field. Legacy JSON overrides still read reply text from response.

For native OpenAI-family Responses traffic, WebSocket/SSE usage aliases normalize the same way, and totals fall back to normalized input + output when total_tokens is missing or 0.

When the current session snapshot is sparse, /status and session_status can recover token/cache counters and the active runtime model label from the most recent transcript usage log. Existing nonzero live values still take precedence over transcript fallback values, and larger prompt-oriented transcript totals can win when stored totals are missing or smaller.

Usage auth for provider quota windows comes from provider-specific hooks first; if a provider has no hook (or the hook does not resolve a token), OpenClaw falls back to matching OAuth/API-key credentials from auth profiles, env, or config.

Assistant transcript entries persist the same normalized usage shape, including usage.cost when the runtime calculates an estimate or the provider reports a billed amount. This gives /usage cost and transcript-backed session status a stable source even after the live runtime state is gone.

OpenClaw keeps provider usage accounting separate from the current context snapshot. Provider usage.total can include cached input, output, and multiple tool-loop model calls, so it is useful for cost and telemetry but can overstate the live context window. Context displays and diagnostics use the latest prompt snapshot (promptTokens, or the last model call when no prompt snapshot is available) for context.used.

Native Codex turn usage sums the reported counts from each unique completed model response, including responses before a retry or cancellation. Missing response counts stay unknown; they do not erase already observed usage. A missing final response snapshot leaves context usage unavailable.

Cost estimation (when shown)

Costs are estimated from your model pricing config:

models.providers.<provider>.models[].cost

These are USD per 1M tokens for input, output, cacheRead, and cacheWrite. If both pricing and a recorded amount are missing, /usage full omits cost; use /usage tokens or a custom messages.usageTemplate when you need token/cache details in every reply. Cost display is not limited to API-key auth: non-API-key providers such as aws-sdk can show estimated cost when their configured model entry includes local pricing and the provider returns usage metadata.

When a model publishes tieredPricing, each request selects one tier using its total prompt input: uncached input plus cache reads and cache writes. Output tokens do not select the tier. The selected rates apply to every token bucket in that request, rather than only to tokens above a threshold. Ranges are half-open [start, end); an open-ended final range uses [start].

Turn totals sum costs calculated for each model request, retaining tier and model boundaries across tool loops and retries. If an older or external runtime provides only aggregate tokens, flat-rate estimates remain available. A tiered aggregate without complete per-request costs omits the cost instead of treating the summed tokens as one large request. Provider-billed totals, including zero, take precedence over catalog estimates and remain visible even when token counts are unavailable. Unknown token counts are not inferred from a billed amount.

Transcript reports preserve valid recorded per-call totals and allocations, including priority/flex adjustments, and use current catalog pricing only for missing costs or unknown-price zero placeholders. Anthropic fast-mode estimates multiply base and tier rates alike, preserving tier thresholds and mixed 5-minute/1-hour cache-write pricing.

Omitting cost, or setting it to {}, inherits the catalog pricing schedule. Explicit flat or all-zero model prices do not inherit a catalog tier schedule. Omitted flat-rate fields can still inherit catalog defaults. An explicit tieredPricing schedule takes precedence over the catalog schedule.

Pricing updates ship in the hosted model catalog alongside model metadata. Its publisher reads public pricing sources, including OpenCode's official catalog and Venice's public model API when the provider declares the native source. Base rates and context tiers come from the same source; usage rendering makes no network requests. Hosted updates activate after the next Gateway restart. Set models.catalogRefresh.enabled: false to disable hosted catalog traffic on offline or restricted networks; bundled pricing still works. Agent-local models.json prices take precedence over explicit models.providers.*.models[].cost entries, and both override catalog estimates, including explicit flat and zero rates.

When the Gateway writes updated agent-local models.json prices, subsequent local estimates use those rates without a restart. Recorded per-call costs keep their original amounts.

OpenRouter :nitro and :floor routing shortcuts use the base model's catalog estimate when the exact shortcut has no price. Recorded costs and explicit prices keep their precedence. Private endpoints and other model variants do not use this fallback. Priority and flex billing can differ from the base estimate.

Cache TTL and pruning impact

Provider prompt caching only applies within the cache TTL window. OpenClaw can optionally run cache-ttl pruning: it prunes the session once the cache TTL has expired, then resets the cache window so subsequent requests re-use the freshly cached context instead of re-caching the full history. This keeps cache write costs lower when a session goes idle past the TTL.

Configure it in Gateway configuration and see the behavior details in Session pruning.

Heartbeat can keep the cache warm across idle gaps. If your model cache TTL is 1h, setting the heartbeat interval just under that (e.g., 55m) can avoid re-caching the full prompt, reducing cache write costs.

In multi-agent setups, you can keep one shared model config and tune cache behavior per agent with agents.entries.*.params.cacheRetention.

For a full knob-by-knob guide, see Prompt Caching.

For Anthropic API pricing, cache reads are significantly cheaper than input tokens, while cache writes are billed at a higher multiplier. See Anthropic's prompt caching pricing for the latest rates and TTL multipliers: https://platform.claude.com/docs/en/build-with-claude/prompt-caching

Example: keep 1h cache warm with heartbeat

agents:
  defaults:
    model:
      primary: "anthropic/claude-opus-4-6"
    models:
      "anthropic/claude-opus-4-6":
        params:
          cacheRetention: "long"
    heartbeat:
      every: "55m"

Example: mixed traffic with per-agent cache strategy

agents:
  defaults:
    model:
      primary: "anthropic/claude-opus-4-6"
    models:
      "anthropic/claude-opus-4-6":
        params:
          cacheRetention: "long" # default baseline for most agents
  list:
    - id: "research"
      default: true
      heartbeat:
        every: "55m" # keep long cache warm for deep sessions
    - id: "alerts"
      params:
        cacheRetention: "none" # avoid cache writes for bursty notifications

agents.entries.*.params merges on top of the selected model's params, so you can override only cacheRetention and inherit other model defaults unchanged.

Anthropic 1M context

OpenClaw sizes GA-capable Claude 4.x models such as Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6 with Anthropic's 1M context window. You do not need params.context1m: true for those models.

agents:
  defaults:
    models:
      "anthropic/claude-opus-4-6":
        alias: opus

Older configs can keep context1m: true, but OpenClaw no longer sends Anthropic's retired context-1m-2025-08-07 beta header for this setting and does not expand unsupported older Claude models to 1M.

Requirement: the credential must be eligible for long-context usage. If not, Anthropic responds with a provider-side rate limit error for that request.

If you authenticate Anthropic with OAuth/subscription tokens (sk-ant-oat-*), OpenClaw preserves the OAuth-required Anthropic beta headers while stripping the retired context-1m-* beta if it remains in older config.

Tips for reducing token pressure

  • Use /compact to summarize long sessions.
  • Trim large tool outputs in your workflows.
  • Lower agents.defaults.imageMaxDimensionPx for screenshot-heavy sessions.
  • Keep skill descriptions short (skill list is injected into the prompt).
  • Prefer smaller models for verbose, exploratory work.

See Skills for the exact skill list overhead formula.