* fix(codex): restore installed skills with catalog-backed models
* fix(codex): keep skill catalog refreshable on live incognito threads
Moving the installed-skill catalog into thread developer instructions made
every catalog refresh look like a generic policy change. On a live incognito
thread the warm-reuse guard compared the whole creation string, so editing a
skill description and sending another message threw
CodexIncognitoPolicyChangeError before turn/start.
The catalog is now a separate lifecycle input (`skillsInstructions`). The
request builder joins it onto the wire developer carrier after the generic
policy, so catalog-backed models still receive it and persistent threads keep
cold-resuming through the existing policy handoff. The retained ephemeral
record splits the immutable generic policy from the catalog last delivered;
incognito warm reuse compares only the generic policy (unchanged guard, still
fail-closed) and delivers a changed catalog with a `thread/inject_items`
developer message, tracked so an unchanged catalog is not re-sent.
Docs: the runtime guide, system-prompt concept, and token-use reference still
said the catalog rode turn-scoped collaboration instructions; they now match
the thread carrier and the incognito refresh.
Codex protocol behavior checked in sibling ../codex (main 459a79e):
core/src/context/world_state/collaboration_mode.rs:30-42 (catalog messages
replace caller collaboration instructions), app-server-protocol v2 thread.rs
(thread/settings/update has no developerInstructions), turn_processor.rs
thread_inject_items, context-fragments/src/additional_context.rs (1,000-token
cap rules out additionalContext as the catalog carrier).
* test(codex): cover catalog refresh on reused persistent threads
Adds the missing half of the reused-thread coverage. The incognito case
already proved an edited catalog reaches a live ephemeral thread through
thread/inject_items. This covers the persistent case: the catalog rides the
thread developer carrier, so it participates in the thread-config
fingerprint (thread-fingerprints.ts:121). An edited skill therefore
invalidates warm reuse and cold-resumes the SAME thread with the new
catalog, instead of either losing the conversation or answering from the
stale catalog.
Mutation-checked: dropping options.skillsInstructions from the join in
buildCodexThreadConfiguration fails both this test and the incognito
two-turn test, so neither assertion is vacuous.
Also adapts the native skills test to upstream's changed shutdown contract.
closeAndWait() returns CodexAppServerCloseResult rather than boolean since
bcef3b3526 (#139657); the assertion now uses the same
toMatchObject({ exited: true }) form as the other native tests
(thread-lifecycle.native.test.ts:142, settled-turn-finalizer.native.test.ts:57).
* fix(codex): break app-server import cycle and register skills test lanes
Extract joinPresentSections into a zero-import leaf module so
thread-requests.ts no longer imports run-attempt-state.ts, which closed a
10-file strongly connected component in extensions/codex/src/app-server.
The merge base is clean, so the cycle was introduced by this branch.
Register the two new skills test files in the vitest project that already
owns the comparable run-attempt.native-hook-relay test, so full-suite file
ownership is exactly one project per file.
* fix(codex): keep refreshed incognito skill catalogs across compaction
An incognito thread carries its skill catalog in creation-time native
developer instructions. A mid-session catalog change is delivered as a
client-authored developer message, but Codex remote compaction drops those
messages unless retain_client_developer_messages is enabled (off by default)
and rebuilds initial context from the unchanged native instructions. The host
had already recorded the refreshed catalog as delivered, so later turns
skipped reinjection and the conversation silently reverted to the catalog it
started with: edited skills kept stale descriptions and removed skills
stayed visible.
Track the creation-time catalog separately from the current one and
re-deliver the current catalog after every compaction, including compaction
inside a turn, via the existing onContextCompacted hook. An unchanged catalog
still sends nothing, because compaction rebuilds it natively.
* test(codex): cover catalog restoration across native compaction
Drive the packaged Codex app-server binary through a real compaction on a
live incognito thread: turn two refreshes the catalog, turn three exceeds the
resolved auto-compact limit, and turn four must still show the model the
current catalog.
The continuation request after summarization is asserted to carry the
creation-time catalog and not the refresh, which reproduces the reversion
natively. The creation-time catalog stays in the ephemeral thread's immutable
developer instructions and is re-emitted on every request, so restoration is
pinned by catalog ordering rather than absence.
Pin nativeSkillsInstructions in the ephemeral policy binding assertions so a
refresh is shown never to rewrite what compaction will restore.
* fix(codex): re-deliver the catalog after a failed post-compaction restore
A failed or aborted restore was logged and dropped while the retained record
still named the refreshed catalog as delivered. Warm reuse compares that
record, so every later turn skipped the refresh and the conversation stayed
stranded on the creation-time catalog for the rest of its life.
Record what compaction actually restored when the injection fails, so the
next turn delivers the current catalog again.
* docs(codex): state catalog restore timing precisely
The restore is issued as soon as a compaction completes, but that turn's own
continuation request may already have been built, so the guaranteed floor is
the following request rather than the remainder of the compacting turn.
* fix(codex): keep incognito creation policy across standalone compaction
Standalone compaction re-retained the live thread without its ephemeral
policy, so the record lost the creation-time generic policy and the catalog
it had delivered. The next turn then read the live incognito thread as policy
drift, and the discarded catalog refresh was never re-sent.
Carry the ephemeral policy back through the re-retain and record the catalog
compaction actually restored, so an edited or withdrawn skill is delivered
again on the next turn instead of reverting for the rest of the conversation.
* fix(codex): revert the catalog on the incognito compaction path
Standalone compaction only claims and re-retains a thread for non-incognito
session keys, so the previous reversion never ran for the conversations it was
written for. An incognito thread keeps its own live subscription, leaving the
retained record still naming the discarded refresh as delivered.
Correct the retained record in place after a completed compaction, and cover
both session kinds so an incognito key can no longer be masked by an ordinary
one in the regression.
* fix(codex): preserve trajectory helper import after rebase
* test(codex): budget the compaction history read in skill catalog tests
The three post-compaction catalog re-delivery tests never observed
`thread/inject_items`, which read as the repair being broken. It was not.
Compaction runs the real `before_compaction` hook, which performs an actual
mirrored-session history read; its first worker start costs several seconds.
`createParams` defaults `timeoutMs` to 5s, so the attempt execution budget
expired inside that read, aborted projection, and the `onContextCompacted`
restore never ran -- reproducing precisely the symptom these tests assert
against.
Give these turns the same budget the sibling compaction tests already use in
`run-attempt.skills.native.test.ts` and `run-attempt.hooks.test.ts`. Every
assertion is unchanged.
* fix(codex): preserve parent-only skills on managed connections
* test(codex): validate managed instructions before collection
---------
Co-authored-by: Rook <rook@openclaw.invalid>
Co-authored-by: Colton Harris <287090614+Colton-Harris@users.noreply.github.com>
Co-authored-by: jalehman <550978+jalehman@users.noreply.github.com>
16 KiB
summary, read_when, title
| summary | read_when | title | ||
|---|---|---|---|---|
| How OpenClaw builds prompt context and reports token usage + costs |
|
Token use and costs |
OpenClaw tracks tokens, not characters. Tokens are model-specific, but most OpenAI-style models average ~4 characters per token for English text.
How the system prompt is built
OpenClaw assembles its own system prompt on every run. It includes:
- Tool list + short descriptions
- Skills list (metadata only; instructions load on demand with
read). Native Codex turns on the managed bundled app-server get the compact skills block in parent-local model request instructions. Connections without that relay use thread developer instructions; other harnesses get it in the normal prompt surface. Bounded byskills.limits.maxSkillsPromptChars, with optional per-agent override atagents.entries.*.skillsLimits.maxSkillsPromptChars. - Self-update instructions
- Workspace + bootstrap files (
AGENTS.md,SOUL.md,IDENTITY.md,USER.md,BOOTSTRAP.mdwhen new, plusMEMORY.mdwhen present). Large injected files are truncated byagents.defaults.bootstrapMaxChars(default:20000); total bootstrap injection is capped byagents.defaults.bootstrapTotalMaxChars(default:60000).- Native Codex turns do not paste raw
MEMORY.mdwhen memory tools are available for that workspace; they get a small memory pointer in parent-local request instructions instead and use memory tools on demand. If tools are disabled, memory search is unavailable, or the active workspace differs from the agent memory workspace,MEMORY.mdfalls back to the normal bounded turn-context path. - Lowercase root
memory.mdis never injected. It is legacy repair input foropenclaw doctor --fix, which migrates it intoMEMORY.md. memory/*.mddaily files are not part of the normal bootstrap prompt; they stay on-demand via memory tools on ordinary turns. Reset/startup model runs can prepend a one-shot startup-context block with recent daily memory for that first turn, controlled byagents.defaults.startupContext. Bare chat/newand/resetare acknowledged without invoking the model.- Post-compaction
AGENTS.mdexcerpts require explicitagents.defaults.compaction.postCompactionSectionsopt-in; plugins can add other context throughbefore_prompt_build.
- Native Codex turns do not paste raw
- Time (UTC + user timezone)
- Reply tags + heartbeat behavior
- Runtime metadata (host/OS/model/thinking)
See the full breakdown in System Prompt.
When documenting credentials or auth snippets, use the Secret Placeholder Conventions to avoid secret-scanner false positives in docs-only changes.
What counts in the context window
Everything the model receives counts toward the context limit:
- System prompt (all sections above)
- Conversation history (user + assistant messages)
- Tool calls and tool results
- Attachments/transcripts (images, audio, files)
- Compaction summaries and pruning artifacts
- Provider wrappers or safety headers (not visible, but still counted)
Runtime-heavy surfaces have their own explicit caps under
agents.defaults.contextLimits (per-agent overrides under
agents.entries.*.contextLimits):
| Key | Purpose |
|---|---|
memoryGetMaxChars |
Max characters memory_get returns before truncation. |
postCompactionMaxChars |
Max characters retained from AGENTS.md during post-compaction refresh. |
These are bounded runtime excerpts and injected runtime-owned blocks, separate from bootstrap limits, startup-context limits, and skills prompt limits.
OpenClaw derives the live tool-result cap from the effective model context
window: 16000 chars below
100K tokens, 32000 chars at 100K+ tokens, 64000 chars at 200K+ tokens.
The runtime context-share guard also caps a single tool result at 30% of the
context window.
Large provider windows are not enabled automatically when they materially
change cost or latency. For example, direct OpenAI GPT-5.5 and GPT-5.6 models
publish a 1050000 token total window, but OpenClaw defaults their active
runtime budget to 272000 tokens. The opt-in 922000 input budget reserves the
full 128000 output allowance, and OpenAI applies higher long-context pricing
to the entire request once input exceeds 272000 tokens. See
OpenAI context window defaults.
For images, OpenClaw downscales transcript/tool image payloads before
provider calls. Tune with agents.defaults.imageMaxDimensionPx (default:
1200):
- Lower values reduce vision-token usage and payload size.
- Higher values preserve more visual detail for OCR/UI-heavy screenshots.
For a practical breakdown (per injected file, tools, skills, and system
prompt size), use /context list or /context detail. See
Context.
How to see current token usage
In chat:
/status-> emoji-rich status card with the session model, context usage, last response input/output tokens, and cost from recorded billing or local pricing for the active model./usage off|tokens|full-> appends a per-response usage footer to every reply. Persists per session (stored asresponseUsage)./usage reset(aliases:inherit,clear,default) clears the session override so it re-inherits the configured default./usage tokensshows turn token/cache details./usage fullshows compact model/context/cost details. Cost comes from a recorded amount or usage metadata with local pricing for the active model. Custommessages.usageTemplatelayouts can include token/cache fields.
/usage cost-> local cost summary from OpenClaw session logs.
Other surfaces:
- Control UI: the working indicator and completed-run recap show cumulative output tokens for that run, including its model calls across tool use and retries. Counts update when the runtime reports completed-response usage, not on every streamed text fragment. Reloading an active run restores its latest count. This counter excludes input tokens and is separate from the composer context-window meter and persisted billing summaries.
- TUI/Web TUI:
/statusand/usageare supported. - CLI:
openclaw status --usageandopenclaw channels listshow normalized provider quota windows (X% left, not per-response costs). Usage-window providers, checked against 2026.9.3: Claude (Anthropic), ClawRouter, Copilot (GitHub), DeepSeek, MiniMax, OpenAI, OpenRouter, Venice, xAI, Xiaomi, Xiaomi Token Plan, and z.ai. Provider plugins supply these snapshots, so an installed plugin can add one.
Usage surfaces normalize common provider-native field aliases before
display. For OpenAI-family Responses traffic, that includes both
input_tokens/output_tokens and prompt_tokens/completion_tokens, so
transport-specific field names do not change /status, /usage, or session
summaries. Gemini CLI usage is normalized too: the default stream-json
parser reads assistant message events, and stats.cached maps to
cacheRead, with stats.input_tokens - stats.cached used when the CLI omits
an explicit stats.input field. Legacy JSON overrides still read reply text
from response.
For native OpenAI-family Responses traffic, WebSocket/SSE usage aliases
normalize the same way, and totals fall back to normalized input + output
when total_tokens is missing or 0.
When the current session snapshot is sparse, /status and session_status
can recover token/cache counters and the active runtime model label from the
most recent transcript usage log. Existing nonzero live values still take
precedence over transcript fallback values, and larger prompt-oriented
transcript totals can win when stored totals are missing or smaller.
Usage auth for provider quota windows comes from provider-specific hooks first; if a provider has no hook (or the hook does not resolve a token), OpenClaw falls back to matching OAuth/API-key credentials from auth profiles, env, or config.
Assistant transcript entries persist the same normalized usage shape,
including usage.cost when the runtime calculates an estimate or the provider
reports a billed amount. This gives /usage cost and
transcript-backed session status a stable source even after the live
runtime state is gone.
OpenClaw keeps provider usage accounting separate from the current context
snapshot. Provider usage.total can include cached input, output, and
multiple tool-loop model calls, so it is useful for cost and telemetry but
can overstate the live context window. Context displays and diagnostics use
the latest prompt snapshot (promptTokens, or the last model call when no
prompt snapshot is available) for context.used.
Native Codex turn usage sums the reported counts from each unique completed model response, including responses before a retry or cancellation. Missing response counts stay unknown; they do not erase already observed usage. A missing final response snapshot leaves context usage unavailable.
Cost estimation (when shown)
Costs are estimated from your model pricing config:
models.providers.<provider>.models[].cost
These are USD per 1M tokens for input, output, cacheRead, and
cacheWrite. If both pricing and a recorded amount are missing, /usage full
omits cost; use /usage tokens or a custom messages.usageTemplate when you need
token/cache details in every reply. Cost display is not limited to API-key
auth: non-API-key providers such as aws-sdk can show estimated cost when
their configured model entry includes local pricing and the provider
returns usage metadata.
When a model publishes tieredPricing, each request selects one tier using its
total prompt input: uncached input plus cache reads and cache writes. Output
tokens do not select the tier. The selected rates apply to every token bucket
in that request, rather than only to tokens above a threshold. Ranges are
half-open [start, end); an open-ended final range uses [start].
Turn totals sum costs calculated for each model request, retaining tier and model boundaries across tool loops and retries. If an older or external runtime provides only aggregate tokens, flat-rate estimates remain available. A tiered aggregate without complete per-request costs omits the cost instead of treating the summed tokens as one large request. Provider-billed totals, including zero, take precedence over catalog estimates and remain visible even when token counts are unavailable. Unknown token counts are not inferred from a billed amount.
Transcript reports preserve valid recorded per-call totals and allocations, including priority/flex adjustments, and use current catalog pricing only for missing costs or unknown-price zero placeholders. Anthropic fast-mode estimates multiply base and tier rates alike, preserving tier thresholds and mixed 5-minute/1-hour cache-write pricing.
Omitting cost, or setting it to {}, inherits the catalog pricing schedule.
Explicit flat or all-zero model prices do not inherit a catalog tier schedule.
Omitted flat-rate fields can still inherit catalog defaults. An explicit
tieredPricing schedule takes precedence over the catalog schedule.
Pricing updates ship in the hosted model catalog alongside model metadata. Its
publisher reads public pricing sources, including OpenCode's official catalog
and Venice's public model API when the provider declares the native source.
Base rates and context tiers come from the same source; usage rendering makes no
network requests. Hosted updates activate after
the next Gateway restart. Set models.catalogRefresh.enabled: false to disable
hosted catalog traffic on offline or restricted networks; bundled pricing still
works. Agent-local models.json prices take precedence over explicit
models.providers.*.models[].cost entries, and both override catalog estimates,
including explicit flat and zero rates.
When the Gateway writes updated agent-local models.json prices, subsequent
local estimates use those rates without a restart. Recorded per-call costs keep
their original amounts.
OpenRouter :nitro and :floor routing shortcuts use the base model's catalog
estimate when the exact shortcut has no price. Recorded costs and explicit
prices keep their precedence. Private endpoints and other model variants do not
use this fallback. Priority and flex billing can differ from the base estimate.
Cache TTL and pruning impact
Provider prompt caching only applies within the cache TTL window. OpenClaw can optionally run cache-ttl pruning: it prunes the session once the cache TTL has expired, then resets the cache window so subsequent requests re-use the freshly cached context instead of re-caching the full history. This keeps cache write costs lower when a session goes idle past the TTL.
Configure it in Gateway configuration and see the behavior details in Session pruning.
Heartbeat can keep the cache warm across idle gaps. If your model cache
TTL is 1h, setting the heartbeat interval just under that (e.g., 55m) can
avoid re-caching the full prompt, reducing cache write costs.
In multi-agent setups, you can keep one shared model config and tune cache
behavior per agent with agents.entries.*.params.cacheRetention.
For a full knob-by-knob guide, see Prompt Caching.
For Anthropic API pricing, cache reads are significantly cheaper than input tokens, while cache writes are billed at a higher multiplier. See Anthropic's prompt caching pricing for the latest rates and TTL multipliers: https://platform.claude.com/docs/en/build-with-claude/prompt-caching
Example: keep 1h cache warm with heartbeat
agents:
defaults:
model:
primary: "anthropic/claude-opus-4-6"
models:
"anthropic/claude-opus-4-6":
params:
cacheRetention: "long"
heartbeat:
every: "55m"
Example: mixed traffic with per-agent cache strategy
agents:
defaults:
model:
primary: "anthropic/claude-opus-4-6"
models:
"anthropic/claude-opus-4-6":
params:
cacheRetention: "long" # default baseline for most agents
list:
- id: "research"
default: true
heartbeat:
every: "55m" # keep long cache warm for deep sessions
- id: "alerts"
params:
cacheRetention: "none" # avoid cache writes for bursty notifications
agents.entries.*.params merges on top of the selected model's params, so you
can override only cacheRetention and inherit other model defaults
unchanged.
Anthropic 1M context
OpenClaw sizes GA-capable Claude 4.x models such as Opus 4.8, Opus 4.7, Opus
4.6, and Sonnet 4.6 with Anthropic's 1M context window. You do not need
params.context1m: true for those models.
agents:
defaults:
models:
"anthropic/claude-opus-4-6":
alias: opus
Older configs can keep context1m: true, but OpenClaw no longer sends
Anthropic's retired context-1m-2025-08-07 beta header for this setting and
does not expand unsupported older Claude models to 1M.
Requirement: the credential must be eligible for long-context usage. If not, Anthropic responds with a provider-side rate limit error for that request.
If you authenticate Anthropic with OAuth/subscription tokens
(sk-ant-oat-*), OpenClaw preserves the OAuth-required Anthropic beta
headers while stripping the retired context-1m-* beta if it remains in
older config.
Tips for reducing token pressure
- Use
/compactto summarize long sessions. - Trim large tool outputs in your workflows.
- Lower
agents.defaults.imageMaxDimensionPxfor screenshot-heavy sessions. - Keep skill descriptions short (skill list is injected into the prompt).
- Prefer smaller models for verbose, exploratory work.
See Skills for the exact skill list overhead formula.