Commit Graph
492 Commits
Author SHA1 Message Date
Zayn Jarvisandtangtao e757dd3749 fix(openclaw): support turn retention and consistent pending tokens (#4972)
Reuse the server turn-budget planner for opt-in auto-commit retention, preserve manual full compaction and setup configuration, and handle zero retained turns when rebuilding pending tokens.

Refs #4415 (items 3 and 6).

Co-authored-by: tangtao <1024583279@qq.com>
2026-09-13 17:51:41 +08:00
Hao Zhe 84bbf79711 feat(parser): add LlamaParse v2 bridge (#4824)
* feat(examples): add LlamaParse understanding bridge

* docs(parser): document long-running live test behavior

* feat(parser): support LlamaParse image and audio URLs

* fix(parser): harden LlamaParse bridge results

* refactor(examples): simplify LlamaParse bridge
2026-09-12 19:49:18 +08:00
NanHan Qing 7c88d5c5ad fix(pi-plugin): preserve camelCase tool result messages (#4940)
* fix(pi-plugin): preserve camelCase tool result messages

* test(pi-plugin): cover native tool result capture and incremental sync

Refactor tests for extractBranchCapturePayloads to handle various tool result scenarios and ensure correct payload extraction.
2026-09-12 01:27:39 +08:00
Zayn Jarvis b1c0bef43a feat(reset-context): on session commit clear, reset context without changing session identity (#4937)
* fix(openclaw): clear reset context without changing session identity

* fix(sessions): unblock commits after reset boundary write failure

* refactor(sessions): read reset boundary from .done and skip duplicate empty resets

- _is_context_reset_archive reads the context_reset property directly instead
  of inferring it from a missing overview
- reset on an already-reset empty session returns confirmation without
  appending another empty archive directory

* refactor(sessions): reset boundary archive holds only .done

Terminal archives never have messages.jsonl read and a missing overview
already reads as empty, so the empty placeholder files were unused.
2026-09-11 19:38:44 +08:00
t0saki bf8c5e9d34 feat(examples): experimental pi extension where the agent manages its own context windows (#4941)
* feat(pi-experimental): fork pi extension into experimental context-management skeleton

* feat(pi-experimental): add context-window core, OV archive client methods and contextWindow config

- lib/context-window-core.mjs: pure, io-injected core for agent-driven context
  windows (tool-call-anchored cut, frozen window header, blocking reset
  pipeline with a single deadline, pending-overview refresh, reminder ladder,
  pi-compaction fallback rules)
- client.ts: sessionRootUri, readArchiveOverview/readArchiveMessages via
  content/read, listSessionArchives, grepSessionArchives, getTask; the
  /sessions/{id}/archives route is not used (blocked on the target gateway)
- config.ts/config.json: contextWindow block with clamps and env overrides;
  dead captureMode key removed

* feat(pi-experimental): wire agent-driven context windows into pi

- context-window.ts adapter binds the core to SyncManager/OVClient/pi
- tools.ts: new_context, history (list_windows/list_items/read_item/
  search_contents) and get_context_remaining with Codex-style names
- index.ts: offline restore before the health check, cut before recall,
  per-prompt [context-status] message, one-shot reminders, pi-compaction
  fallback through the core
- core follow-ups: abort checks before the handoff post, flush budget floor,
  syncBranch in handleBeforeCompact, lazy session validation, persisted
  previous overview
- docs: README, CONTEXT-WINDOW.md, agent-integrations pages; CI test glob
- scripts/e2e-window.mjs live gate replaces the takeover-era e2e-live

* fix(pi-experimental): apply three-lens review findings

- coexistence guard keyed on tool sourceInfo.path and re-checked on
  start/before_agent_start/turn_end
- reminders: ignore aborted/error turns as first observation, gate the idle
  note on the gap the user just returned from, threshold-free guidance
- core: persisted archives ledger for window/archive mapping, stale reminder
  detection for custom-role messages, task-aware overview wait, retracted
  handoff on post-handoff refusals, notes cleared on compaction fallback
- tools: refusal instructions, history fails closed on unreadable listings,
  grep line/item index kept stable, viking_search scope text without ~
- config: recentResetGuardMs exposed, dead keys removed; docs synced

* fix(pi-experimental): status line metrics from the branch on a fresh process, archive line index stability, wording nits

* test(pi-experimental): reasoning level, deterministic tool output and keepable OV session in the e2e gate

- E2E_LLM_REASONING=off|minimal|low|medium|high|xhigh marks the model as
  reasoning-capable and sets pi's defaultThinkingLevel; a custom relay also
  needs compat.supportsReasoningEffort, which URL auto-detection cannot infer
- T1 now reads a seeded release.md, so the archive reliably carries a
  [tool-result ...] entry instead of depending on the model reaching for a tool
- E2E_KEEP_OV_SESSION=1 skips the session delete so the archives stay readable
  for a demo

* test(pi-experimental): add a long-context scenario where the agent resets under pressure

E2E_WINDOW_LONG=1 seeds the extension's own sources (22 files, ~105k tokens of
material) into the workspace and asks for a file-by-file inventory, without ever
mentioning the context tools. The gate then checks what the agent did on its
own: how full the window got, whether it reset, whether the cut was legal, and
whether the work continued across the boundary.

- peak pressure is measured from the provider payloads, not only from the
  per-prompt [context-status] line, which undersamples a tool-heavy turn
- runTurn takes a timeout; the long turns get 25 minutes instead of 10
- judgement-dependent checks warn, harness behaviour still fails the gate

Observed on doubao-seed-2-1-pro with reasoning high: 96 requests, peak 47% of a
128k window, three self-initiated resets, 22/22 files inventoried.

* docs(pi-experimental): publish redacted demo evidence, drop the hardcoded relay

- demo-evidence/pi-ctxwin-demo.zip: two real runs against a live OpenViking
  server and a live model, with transcripts, the provider payloads either side
  of every reset and the archives pulled back off the server. Secrets file
  absent, relay hostname and operator username replaced; REDACTIONS.md inside
  the archive lists every substitution
- e2e-window.mjs no longer defaults E2E_LLM_BASE_URL / E2E_LLM_MODEL to a
  private relay: both are now required, so no endpoint of anyone's is baked in
- CONTEXT-WINDOW.md documents the reasoning, long-context and keep-session
  knobs added with those scenarios

* docs(pi-experimental): name the real endpoint and model in the demo evidence

The published archive now says what the runs actually used — Volcengine's
Doubao 2.1 Pro (doubao-seed-2-1-pro-260628) on Ark at
https://ark.cn-beijing.volces.com/api/v3 — instead of an anonymous relay
placeholder. The runs went through a private proxy in front of Ark; the same
key reaches the official endpoint directly, and REDACTIONS.md says so.

* docs(pi-experimental): say which repo-level hooks this directory deliberately does not touch

The extension keeps its diff inside its own folder, so the CI test glob and the
shared-file sync TARGETS entry are not added. Record both, plus the local
.gitignore that re-includes lib/, so a maintainer promoting this out of
experimental status knows what to wire up.
2026-09-11 17:19:15 +08:00
Jiahui ZhouandTRAE CLI f2832b7b9b Feat/configurable executor workers (#4913)
* feat(server): configure asyncio executor workers

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* refactor: clarify executor thread configuration

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-11 11:02:33 +08:00
Hao Zhe f55cb4170f fix(openclaw): recall from the incoming assembly prompt (#4906)
* fix(openclaw): recall from the incoming assembly prompt

* fix(openclaw): allow fresh recall after transient injection
2026-09-10 20:49:02 +08:00
Zayn Jarvis 6b9d54900f fix(openclaw): preserve capture across host turn protocols (#4900) 2026-09-10 19:19:07 +08:00
Nick Jamesandpc.yu 3841e6f299 fix(plugins): drain the dsh pending queue in-process so a transient write failure self-heals (#4779)
* fix(plugins): drain the pending queue in-process so a transient write failure self-heals

The dsh memory plugin latches capture and commit on the first retryable
write failure (hasPendingWrites) and only reset the latch at session
init, so the long-lived dsh process stayed stuck until restart.

Add a per-process single-flight drainer (default 60s, env
OPENVIKING_PENDING_DRAIN_INTERVAL_MS) that follows the session-start
flow: probe health, replay the queue without consuming retry budgets,
then re-derive every session's latch from the queue. replayPending gains
an optional consumeRetries flag (default true, byte-compatible):
drainers release a failed claim back to its original filename instead of
incrementing the retry count, so the session-start path keeps owning all
retry accounting and D4 deletions. Latch and health transitions are
logged once per flip for observability.

* test(plugins): cover the drainer and non-consuming replay mode

Add pending-queue coverage for consumeRetries:false (retryable failures
stay retryable and ordered, non-retryable and exhausted entries still
delete, commitSession failures keep the run going, default mode
unchanged) and runtime drainer coverage (recovery clears the latch,
outages keep it and leave entries retryable, empty queue means zero
HTTP, commit resumes after the drain, single-flight, per-session latch
isolation, interval wiring with env fallback).

---------

Co-authored-by: pc.yu <nick@fourieralpha.com>
2026-09-10 13:26:46 +08:00
Qin Haojie a386269d38 fix(resources): 明确等待超时语义并改用任务轮询示例 (#4888)
等待超时返回任务 ID 和状态查询提示,避免误认为后台导入失败;主要文档和示例改为提交后查询任务状态。
2026-09-10 11:52:14 +08:00
t0saki 708dba60bd feat(plugins): configurable regex input filters for recall queries and captured turns (#4858)
* feat(plugins): ordered regex input filters for plugin input

The memory plugins feed the raw user prompt into recall and the raw
transcript into capture. Nothing between the harness and the server lets
an operator shape that text: a thinking-keyword prefix (`ultrathink ...`)
goes into the search query verbatim, a slash command becomes a memory,
and a token pasted into a prompt is stored as it was typed. The only
user-configurable pattern today is `bypassSessionPatterns`, which matches
a session id or a cwd, never the text; every text rule -- `ACK_RE`,
`SLASH_COMMAND_RE`, `sanitizeCapturedText` -- is hard-coded.

`lib/input-filters.mjs` adds the missing layer: an ordered list of
sed-style rules, `s` to substitute, `d` to drop the text on a match and
`k` to keep it only on a match, each optionally scoped to one role. The
rules are plain strings so they survive the existing string-list config
coercion (env CSV, `ovcli.conf` array) without a new config type. Nothing
in the module throws: an unparsable rule, an unknown flag or a pattern
RegExp rejects becomes an entry in `errors` and is skipped, because a
typo in a config file must not take a hook down. Compilation is memoised
per rule list so a long-lived host compiles once, and rule count and
pattern length are capped. The time budget is deliberately advisory --
`applyInputFilters` reports `slow`/`elapsedMs` and runs every rule
anyway, since skipping the tail of the list would skip exactly the
redaction rule an operator added.

`capture-utils.mjs` grows the capture-side seam. `filterCaptureParts`
prunes blank text parts as before, then takes one drop verdict per turn
on the aggregate of its text parts and rewrites the individual parts with
the substitutions only, so a rule can never both keep and drop the same
turn; tool call/result payloads are carried through unmatched, though a
turn dropped on its text takes them with it. `shouldCaptureText` gains
the same step for the text path, behind a `filters` option that
`extractCaptureTurns` turns off when parts are on the wire -- one
application per stored string, and `turn.text` stays faithful for callers
that scan it for trigger words.

No loader supplies either knob yet, so this commit changes no behaviour:
with `captureFilters` unset every path is the identity it was.

* feat(claude-code): recallQueryFilters / captureFilters

Wires the shared filter engine into the Claude Code plugin. Both knobs
are string lists shaped exactly like `bypassSessionPatterns`: an env CSV
replaces the configured array wholesale, entries are trimmed and
non-strings dropped, and a rule that needs a literal comma has to come
from the `ovcli.conf` array because the env value is split on one.

Recall filters run in `auto-recall.mjs` immediately above the
`minQueryLength` gate, so a prompt whose only content was a stripped
prefix reports `short_query` rather than searching for the empty string,
and a dropped prompt gets its own `last-recall.json` reason,
`query_filtered` -- distinct from `filtered_out`, which is the score
threshold. The statusline only tests for `ok`, so it needs no change.

Capture filters run at the send site rather than in the extractor,
because the cursor CC persists is an index into the extracted turn list:
dropping a turn earlier would leave `capturedTurnCount` describing a list
the next run does not reproduce. `sanitizePartsForSend` in
`auto-capture.mjs` and the equivalent in `subagent-stop.mjs` are the two
places every payload passes through, and both now delegate their blank-
part pruning to `filterCaptureParts`, which keeps that behaviour when no
rules are configured.

`ov-memory-doctor` lists the active rules per knob and warns, with the
exact parse or RegExp error, about the ones it could not compile --
config errors have to be visible somewhere other than a debug log,
because a skipped rule is silent by design.

The README gains an "Input filters" section with the grammar, six worked
examples (every one of them asserted in the shared engine test), an
ovcli.conf sample, and the caveats that actually bite: the env comma
split, that drops win over substitutions, that filters run before the
built-in ack/slash heuristics, that a mid-session drop rule shortens the
turn list and reads as a transcript rewrite, and that nothing already
stored is rewritten.

* feat(codex): recallQueryFilters / captureFilters

The same two knobs on the Codex side. The loader gains the string-list
helper the harness did not have yet -- an env CSV replaces the configured
array wholesale, both paths trim and drop non-strings -- and reads
`cx.recallQueryFilters` / `cx.captureFilters`, so `ovcli.conf`
`plugin.codex`, `plugin` and the legacy `ov.conf` `codex` block all work
without further change.

Recall filters go in `auto-recall.mjs` immediately above the
`minQueryLength` gate: a dropped prompt emits the empty envelope before
the health check, so a filtered turn costs no request at all.

Capture needs no edit here. Codex's four write hooks funnel through
`catchUpTurns` into the shared `extractCaptureTurns`, which is where the
filters already run -- and it has to be there rather than at the send
site, because this cursor advances by payloads durably sent: dropping a
turn later would leave the cursor short and replay the surrounding turns
on the next hook. The test covers exactly that, running the same
transcript twice and asserting the second pass finds nothing new.

`ov-memory-doctor` reports the active rules and the compile errors, and
the README documents the grammar, both env examples, the ovcli.conf array
for rules that need a literal comma, and the ordering and mid-session
caveats.

* docs(agent-integrations): input filter env vars

The Claude Code integration guide's env table is where most people look
first, so the two new knobs belong there next to the bypass patterns they
resemble. The grammar and the worked examples stay in the plugin README,
which the table already links to.

* refactor(plugins): drop machinery the input filters do not need

A review pass against the actual requirement — let an operator configure
a regex over plugin input — found five constructions that solve problems
this feature does not have. All of them go:

- The compile memo cache, its FIFO eviction and the test-only reset. It
  was justified as "a long-lived host compiles once", but the only hosts
  that supply these knobs are the Claude Code and Codex hooks, which are
  short-lived subprocesses; the per-turn capture path is the one repeat
  caller, and compiling five rules measures 3 microseconds. The cache
  cost more code than it saved work.
- `maxRules` / `maxPatternLength` as injectable options nothing passed,
  and the caps themselves. Rules come from the operator's own config, so
  a 33rd rule or a 600-character pattern is self-inflicted and harmless,
  and `bypassSessionPatterns` — the same shape of knob, also operator
  regexes — has never capped anything.
- The advisory time budget: `slow` / `elapsedMs` and the injectable
  `now` whose only non-default caller was the test that made it look
  slow. The pattern that actually hurts is one that backtracks, and that
  hangs rather than reporting a number.
- The try/catch around rule application. Every hook already ends in
  `main().catch(...)` that logs and approves, so an exception here was
  never going to take a session down. The try/catch around `new RegExp`
  stays: that one parses operator input, and the doctors render its
  message.
- The warnings channel, which existed to carry one advisory string.
  Stripping `g` from a `d`/`k` rule is load-bearing — `.test()` on a
  global regex advances `lastIndex` and would alternate between calls —
  but it needs no announcement, since the rule then means exactly what
  it looked like it meant.

`filterCaptureParts` also returned `ruleIndex` / `op` / `slow` that no
caller read, and `extractCaptureTurns` tested `decision.reason ===
"filtered"` inside a condition where that disjunct could never be the
deciding one: `filters` is passed as `shaped.parts.length === 0`, so the
reason can only appear when the other half is already true.

Docs lose the retired caps and a duplicated note; the Claude Code recall
test loses the case Codex's suite already covers. Engine 303 -> 218
lines, behaviour identical for every rule an operator can write.

* docs(agent-integrations): point the filter knobs at their grammar

The env-var rows added to the Claude Code guide named the two knobs but
said nothing about how to write a rule, and described only the
comma-separated environment form — which reads as if that were the only
way to configure them. Both rows now link to the plugin README's Input
filters section, and a note under each table says the knobs also live in
`ovcli.conf` under `plugin.<harness>.<key>` or the shared `plugin.<key>`
as a JSON array, which is the better form: the env vars are split on
commas, so a rule containing a literal comma can only be written in the
array. The grammar itself stays in the plugin READMEs, where these guides
already send readers for the full env-var list.

The Codex guide had no mention of the feature at all; it gets the same
two rows, the same note, and a link to its own README.

The capability reference carried counts this feature invalidated:
`memory-plugin-shared/lib/` is 24 modules rather than 23, the hook set
`sync.mjs` gives every hook-driven plugin is 13 rather than 12, and every
per-target total shifts with it (dsh 15, opencode 17, zcode 19,
claude-code and codex 22 each). pi's entry was already wrong before this
change — it takes `batch-send` as well as `setup-wizard`, so 15, not 13 —
and is corrected in the same sentence. `input-filters.mjs` joins the
module table, and the `capture-utils.mjs` row now says "built-in capture
heuristics" so the two rows do not both read as "capture filtering".

* docs(configuration): document the ovcli.conf plugin section

`docs/*/configuration/02-client.md` is titled "ovcli Configuration" and
covers connection, command behaviour, upload filters and the workspace
layers — but never described the `plugin` section itself. It appeared
only twice in passing: two rows of the workspace precedence table, and
one sentence about `peerSource`. The nearest thing to a general
explanation was buried in the agent-integrations overview, inside a
discussion of recall latency.

That gap is why the input filter note added in the previous commit had to
explain the section's shape inline for two knobs, which reads as an
oddity rather than a reference. The section is now documented once, where
someone editing `ovcli.conf` would look: what `plugin` and
`plugin.<harness>` mean, that keys are the camelCase counterparts of the
`OPENVIKING_*` tuning variables and that the mapping does not go both
ways, that list-valued knobs are JSON arrays here while their environment
counterparts are comma-separated, the full resolution order, when an edit
takes effect versus when the agent has to restart, that only Claude Code
and Codex read it today, and that `ov-memory-doctor` flags unrecognised
keys. The per-knob lists stay in the plugin READMEs, which the section
links.

The two integration guides now point at it in one line instead of
restating it, the workspace precedence table links the section it already
named, and the complete example grows a `plugin` block.
2026-09-09 17:12:16 +08:00
t0saki 3ae94afeea fix(skills): 让 Experience 技能兼容没有 viking://~ 的旧服务端 (#4857)
* fix(skills): resolve an explicit user root when the server has no viking://~

The Experience workflow tells the agent to scope its search by
`viking://~/memories/experiences`. `viking://~` is the home alias for
the caller's own space, added in v0.4.16 (#4167); a server older than
that has no branch for the `~` scope, so the URI never resolves and the
request comes back as `INVALID_URI` (HTTP 400) instead of a search. The
skill offered no second spelling, so on those servers the first step of
the workflow failed and Experience retrieval stopped there — reported
in #4828, where every read succeeded once the caller substituted its
own `viking://user/<user_id>/memories/...` root by hand.

The alias is the only spelling that works everywhere it exists, so the
skill keeps leading with it and now says what to do when it is
rejected: rebuild the root from a canonical `viking://user/<user_id>/`
URI already visible in the session, or repeat the query unscoped and
keep the hits whose URI contains `/memories/experiences/`, whose URIs
carry the canonical user id for the reads that follow. Both paths
resolve the id from evidence, which is what keeps the existing "never
hardcode `default`" rule intact — the reporter's server resolved to
`default`, another account's will not.

The uid-less `viking://user/memories/experiences` is called out as a
non-fallback: only pre-v0.4.17 servers expand it, only for USER and
ADMIN callers, and #4196 made current servers reject it, so an agent
reaching for the obvious shorthand would fail on both ends of the
version range.

The same note goes to the agent-plugins memory skill, which scopes
recall the same way, and a sync test pins the pairing so a later edit
cannot drop the fallback while keeping the alias.

* chore(plugins): bump the plugins that ship the experience skill

Claude Code and Codex cache a plugin under
`cache/<marketplace>/<plugin>/<version>/`, so a skill edit behind a
frozen manifest version never reaches an installed user: `plugin
update` sees the same version and does nothing. The three plugins whose
skills changed move one patch step — claude-code 0.4.5 -> 0.4.6 (its
`package.json` stays in lockstep with the manifest), codex 0.8.1 ->
0.8.2 and agent-plugins 0.1.0 -> 0.1.1. The openclaw plugin resolves
its version at release time, so its copy needs no manual bump.
2026-09-09 15:36:58 +08:00
bot-of-qin-ctxandqin-ctx 7eed9adf75 refactor(compile): 改用 instruction 并兼容 reason 参数 (#4852)
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-09-09 14:55:48 +08:00
Zayn Jarvis 26cb585cae fix(opencode): automatically version npm releases on main (#4838) 2026-09-09 12:05:39 +08:00
zgy 2eb36eabbd test(fs): align cp overwrite integration expectation (#4809)
* test(fs): align cp overwrite integration expectation

* test(plugin): sync openclaw memory shared copy
2026-09-08 17:13:13 +08:00
Xinmin Zeng 58bafa5ba1 fix(codex): reuse shared recall compressor (#4445)
* fix(codex): reuse shared recall compressor

* fix(plugins): harden recall compressor fallbacks
2026-09-08 15:45:13 +08:00
Qin Haojie b6af3d4bfc feat(compile): 由 OV 托管外部任务生命周期 (#4436)
* feat(compile): let OV own external task lifecycle

Keep durable task state, recovery, query, and cancellation in OpenViking while external providers execute the workload.

* feat(compile): align external session protocol

Persist external session state in OV and adapt VikingBot to the documented create, status, and cancel contract.

* fix(compile): infer external API from base URL

Remove the redundant enable switch so generated Base Server configuration activates Compile directly from its configured endpoint.

* refactor(compile): simplify task lifecycle controls

Remove client-facing runtime and wait controls, and retire legacy OV routes in favor of the generic Task API. Bound transient status polling failures so unavailable providers fail the owned task.

* fix(compile): restore server runtime deadline

* fix(compile): enforce timeout in OV task polling

* feat(compile): expose args in CLI and SDKs

* fix(compile): correct provider retry boundaries

* fix(compile): bound cancellation convergence

* fix(compile): retry submit without runtime cap

* refactor(compile): keep provider credentials minimal

* fix(compile): poll tasks every 30 seconds by default

* refactor(compile): standardize runtime task protocol

* fix(compile): wait for remote cancellation to settle
2026-09-08 14:53:17 +08:00
now-ingandnow-ing 527139ad49 test(codex-plugin): give fake codex a startup-safe compressor budget (#4750)
The endpoint-compression cases set OPENVIKING_RECALL_TIMEOUT_MS=10000,
which leaves the compressor child only max(1000, recallTimeoutMs -
10000) = 1000ms before auto-recall SIGKILLs it. The fake codex is a
real node process behind an env shebang whose cold start intermittently
exceeds that budget on a loaded CI runner: the child is killed before
it appends its args log, so the base_url assertions see an empty log.

Set explicit generous budgets (recall 60000ms, compressor 30000ms) so
the timing cannot race process startup. Runtime behaviour under test
is unchanged.

Co-authored-by: now-ing <now-ing@users.noreply.github.com>
2026-09-08 12:34:35 +08:00
now-ingandmac 188ef4251c refactor(openclaw-plugin): remove dead recall-trace compat exports (#4732)
* refactor(openclaw-plugin): remove dead recall-trace compat exports

recall-trace.ts still exported the RecallResourceType type plus the
normalizeResourceTypes and resolveRecallSearchPlan helpers even though
registries/recall-resource-types.ts is the canonical home and every
production consumer already imports from the registry.

Delete the dead exports and wire the trace schema's internal uses to
the registry type. The facade union included a "session" member that
ALLOWED_RESOURCE_TYPES always rejected; it only existed so archive-grep
traces could label entry.resourceTypes with ["session"]. That single
use is now typed honestly at the field level
(Array<RecallResourceType | "session">), mirroring how searches/results
already extend the union with "archive". No trace payload or behavior
change; tsc now verifies the rest of the tree against the narrow union.

Duplicate describe blocks in tests/ut/recall-trace.test.ts go away;
their coverage already lives in tests/ut/recall-resource-types.test.ts
against the registry, and the session-rejection case is preserved there.

Clears the 'keeps dead recall-trace resource helper compatibility
exports removed' architecture-boundaries guard. Follow-up to the #4720
discussion, which tracks each remaining guard separately.

Verified: tsc -p tsconfig.json and tsconfig.build.json clean; full
vitest suite 760 tests -> 757 passed with the only failures being the
three boundary guards outside this PR's scope.

* chore: retrigger CI after unrelated pathlock flake in test_fs_cp

---------

Co-authored-by: mac <bishopapril850965@yahoo.com>
2026-09-08 12:31:56 +08:00
now-ingandmac 70027e68a0 refactor(openclaw-plugin): route setup CLI network probes through the probe service (#4731)
commands/setup.ts kept inline copies of probeApiKeyType and
checkServiceHealth with direct fetch/AbortController usage even though
services/setup/probe-service.ts already provides the same probes behind
an injectable transport seam (covered by tests/ut/setup-probe-service.test.ts).

Wire the setup CLI to createSetupNetworkProbes and delete the local
duplicates: same URLs, headers, 10s timeout, and status-code handling;
pluginVersion/compatRange/checkVersionCompatibility come from the same
module constants. Also drop the dangling probeApiKeyType entry from the
__test__ aggregate (no consumers).

Clears the 'keeps setup CLI free of direct network fetch logic'
architecture-boundaries guard. Follow-up to the #4720 discussion, which
tracks each remaining guard separately.

Verified: tsc -p tsconfig.json and tsconfig.build.json clean; full
vitest suite 766 tests -> 763 passed with the only failures being the
three boundary guards outside this PR's scope.

Co-authored-by: mac <bishopapril850965@yahoo.com>
2026-09-08 12:28:09 +08:00
ralf003 1cbc948243 refactor(openclaw-plugin): remove client-facade duplicates left by the seam refactor (#4720)
The routing/memory-uri and adapters/resource-packager seams already own
memory-URI classification and temp-upload/zip packaging respectively, but
client.ts still carried the pre-refactor copies, which keeps the
architecture-boundaries guard suite failing. Remove the superseded helpers
and their now-unused node:*/fflate imports; the facade keeps delegating to
the injected ResourcePackager and still emits viking://~/session URIs, so
there is no runtime or public-API change.

Also normalize one path.relative() result to POSIX separators in the guard
suite so its test-self exclusion matches on Windows as well (no-op on POSIX).

Two other red guards in the same file track separate, unfinished migrations
(root recall-trace.ts compat exports, commands/setup.ts direct fetch) and
are intentionally left for follow-up PRs.
2026-09-08 11:53:15 +08:00
chenjwandTRAE CLI a843ab6bf2 session: add restricted Python DSL extraction protocol and make it the default (#4581)
* session: add restricted Python DSL extraction protocol and make it the default

Introduce a restricted Python memory SDK output protocol as an alternative to
the JSON extraction protocol, and switch the default to python. Both protocols
share the same ResolvedOperations post-processing, schema rules, and patch-repair
path via a new ExtractionOutputProtocol abstraction.

- Add extraction_output_protocol/{base,json,python}.py; python compiles a
  restricted AST into the same operations model as json.
- Default memory.extraction_output_format flips json -> python.
- Surface the offending source line on Python syntax errors and add targeted
  triple-quote retry guidance for string-literal breaks.
- Preserve every distinct fact on canonical merges; remove hardcoded memory
  type names from prompts so custom memory_types render dynamically.
- Downgrade benign batch-delete link-inheritance read failures to WARNING.
- Add memory_organization A/B benchmark and message_format pretty-printer.

Tests: extraction protocol, config loader, memory react suites pass.
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: drop redundant entity split hint and duplicate abstractmethod

- entities.yaml: remove the size-triggered split hint; when to split/compact is
  decided at read time by memory_maintenance_notice, so the static schema
  description only keeps the identity semantics and fact-preservation rule.
- vlm/base.py: remove a duplicated @abstractmethod on get_completion_async.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* vlm: drop redundant *.vlm.call span decorators for trace parity

volcengine already dropped its @tracer("volcengine.vlm.call") wrapper to avoid
duplicate spans now that the request is logged via tracer.info(llm_input_messages=...).
Remove the symmetric litellm/openai decorators so all three backends behave the same.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* benchmark/locomo: fix commit_session kwarg for CLI AsyncHTTPClient

ov.AsyncHTTPClient resolves to openviking_cli.client._http_compat.AsyncHTTPClient,
whose commit_session takes a flat telemetry= kwarg and has no options= parameter.
Passing options={...} (the SDK-client shape) raised TypeError during import.
Use telemetry=True to match the CLI client, consistent with the other locomo
import scripts.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: don't parse error-target sentinels as URIs in extraction telemetry

The by-type extraction telemetry treated result.errors[].uri as a valid viking
URI and fell back to MemoryUpdater.memory_type_from_uri(), but that field is an
error *target* — it can be a sentinel like "unknown" or "events(page_id=100)".
VikingURI() then raised 'URI must start with viking://', turning a single
recorded extraction error into a crash of the whole long_term extraction step.
Count failed errors by the known uri->type map only, defaulting to "unknown".

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* bot: pass context_type via FindOptions after SDK find/search sync

The SDK find/search sync moved context_type from a top-level find() kwarg into
FindOptions. VikingBot still called client.find(context_type='memory') for peer
recall, so every per-turn type-quota recall raised 'unexpected keyword argument
context_type' and silently returned no memories. The answer agent then fell back
to manual multi-round search (iteration ~1.3 -> ~3.9) and accuracy dropped from
~83% to ~72-77%. Pass it via options={'context_type': 'memory'} instead.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* bot: adapt VikingClient.find to SDK FindOptions for context_type/filter

The SDK find/search sync moved context_type and filter out of top-level find()
kwargs into FindOptions. VikingClient.find still forwarded them as top-level
kwargs to the SDK client, so peer memory recall raised 'unexpected keyword
argument context_type' (and after the prior partial fix, 'options') and returned
no memories — the answer agent fell back to manual multi-round search, spiking
iteration ~1.3 -> ~4 and dropping accuracy ~83% -> ~76%.

Do the SDK adaptation once in VikingClient.find (pack context_type/filter into
options={...}); callers keep the stable VikingClient.find(context_type=...)
interface, so memory.py reverts to passing context_type= directly. Verified via
a single-question smoke: type_quota recall returns 13 memories, injection is
non-empty, iteration=1, answer correct.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: route event resolution repair through the output protocol

The event resolution-repair instruction was hardcoded to demand a JSON object,
but under the default Python protocol the repaired response is parsed by the
Python SDK compiler. When a first-pass event had out-of-bounds ranges, an
assistant-only span, or an ambiguous peer, the repair round returned JSON, the
compiler rejected it as an invalid program, retries were exhausted, and the
recoverable event memory was never written.

Add ExtractionOutputProtocol.render_resolution_repair(); JSON keeps the existing
JSON-object wording, Python asks for corrected sdk.create_events(...) calls.
_build_resolution_repair_instruction now delegates to the active protocol, like
patch-repair and the final instruction already do.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session,bot,benchmark: address PR review findings

- extract_loop: document 32768 as the extraction output floor (tuned for Doubao;
  lower-max models override via vlm.max_tokens) and extract
  _resolve_effective_max_output_tokens; ov.conf.example notes the override.
- python_protocol: alias non-identifier memory_type/field names on the Python DSL
  surface only (real names kept in URIs/storage/JSON); map aliases back when
  compiling, instead of hard-rejecting kebab-case custom schemas.
- run_full_eval.sh: move auto-commit + GIT_COMMIT_ID capture AFTER arg parsing so
  --auto-commit is honored and run metadata records the committed HEAD.
- litellm_vlm: strip Gemini cache_control from the already-sanitized messages so
  empty assistant turns are not reintroduced; sanitize_openai_messages passes
  through non-dict entries.
- Tests for each fix.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: reject Python DSL alias collisions instead of silently overwriting

_identifier_alias() is not one-to-one: memory_type 'project-notes' and
'project_notes' (or fields 'note-body'/'note_body') fold to the same DSL alias.
The alias->real dict comprehensions would silently drop one, making a schema/
field unreachable and routing writes to the wrong target. Add
_validate_alias_uniqueness(), invoked in render_contract() and the compiler
__init__ (so parse() paths without render are also guarded), which fails loudly
with a rename hint. Distinct-identifier names never collide, so real configs are
unaffected. Tests cover type, field, and no-render parse collisions.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: bound str.replace result size before allocation in Python DSL

The _MAX_EXPRESSION_SIZE guard covered * (repeat) and + (concat) but not the
whitelisted string methods: only join() had a projected-size check, so
('x'*1000).replace('x','y'*10000) could still allocate a >1MB result and bypass
the limit. Add _check_replace_size() that bounds source + occurrences*(len(new)
-len(old)) BEFORE calling str.replace (which builds the whole result in C), so
the oversized string is never allocated. replace is the only whitelisted method
that can materially inflate output (join already guarded; upper/lower/strip/
split/startswith/endswith do not grow). Tests cover an oversized replace being
rejected and a normal replace passing.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: close f-string width and str %-format inflation in Python DSL

The replace guard alone was insufficient: f-string format specs (f"{'x':>1000001}")
and str %-formatting ("%1000001s" % "x") also turn a small integer literal into an
arbitrarily large string with no repeat operator, bypassing _MAX_EXPRESSION_SIZE.
Neither has a legitimate use in memory content, so disallow them outright rather
than bounding width inflation: reject any non-empty f-string format spec and reject
str/bytes %-formatting (numeric % still allowed). Combined with the existing
*/+/join/replace pre-allocation checks, all small-input->large-output amplifiers
are now closed. Tests cover f-string width rejection, plain f-string, and str %.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-07 20:49:02 +08:00
Nick Jamesandpc.yu 98f24e1690 fix(plugins): treat camelCase isError as an error tool result (#4724)
* fix(plugins): treat camelCase isError as an error tool result

DSH emits tool-result blocks with a camelCase isError field
({type:"tool-result", toolCallId, content, isError}), but the shared
capture-utils toolStatus() only checks is_error / error / state.error.
Failed results were therefore labeled tool_status=completed, while
their error text still landed in tool_output. The mislabel does not
mislead LLM memory extraction (verified end-to-end), but it does
corrupt status-driven consumers: experience read lineage
(experience_lineage.py), usage reporting, working-memory formatting,
and rollout training artifacts.

Recognize block.isError alongside is_error. Vendored copies are
regenerated via examples/memory-plugin-shared/sync.mjs. Adds a
capture-utils regression test plus a DSH capture test for the real
tool-result wire shape.

* fix(plugins): register capture-utils test in CI and cover state.isError

Review follow-up: add examples/memory-plugin-shared/capture-utils.test.mjs
to the pr.yml plugin-tests list (it was author-local only), and recognize
block.state.isError alongside state.error in toolStatus with a matching
test case. Vendored copies regenerated via sync.mjs.

---------

Co-authored-by: pc.yu <nick@fourieralpha.com>
2026-09-07 17:22:02 +08:00
now-ingandmac f7c6e84386 fix(memory-plugins): report client-side MCP proxy timeouts as -32004 instead of unreachable (#4741)
A request aborted by the proxy's own timeout budget fell into mapError()'s
catch-all and surfaced as -32001 'check the URL / server reachable' even
while the server was healthy and still computing — rerank-inclusive
find/search legitimately runs tens of seconds past the default 15s budget,
so every such call was mislabeled as an outage (#4739).

Branch on AbortError before the catch-all and return a dedicated -32004
naming the elapsed budget, the endpoint, and the OPENVIKING_TIMEOUT_MS
knob; -32001 keeps its meaning of genuine connection failure. -32003 is
already taken by the empty-response error, so timeouts use -32004.

The change is applied to the shared lib and propagated to every generated
copy via sync.mjs; the new mcp-proxy-core.test.mjs covers both the abort
and the connection-failure contrast and is registered in the plugin-tests
list in pr.yml.

Signed-off-by: mac <bishopapril850965@yahoo.com>
Co-authored-by: mac <bishopapril850965@yahoo.com>
2026-09-07 17:20:39 +08:00
Abhay fd6b3f62c9 fix(codex): quote plugin hook script paths (#4707) 2026-09-07 14:02:30 +08:00
Xinmin Zeng 9c31b11076 fix(codex-doctor): accept [features] hooks = true in modern Codex (#4634) (#4635)
* fix(codex-doctor): accept [features] hooks = true and retain legacy plugin_hooks compatibility

* fix(codex-doctor): probe live CLI features for default-enabled hooks and normalize symlink run

* fix(codex-doctor): prioritize modern hooks over legacy flags
2026-09-07 14:01:35 +08:00
Zayn Jarvis e86642cac2 fix(openclaw-plugin): fall back to native compaction for bypassed sessions (#4704) 2026-09-05 10:28:52 +08:00
854ff4ceb0 fix(pi-extension): replay offline backlog through the batch endpoint (#4692)
* fix(pi-extension): replay offline backlog through the batch endpoint

After a server outage pi's local pending queue can hold hundreds of
addMessage entries. The takeover barrier requires that queue to be empty
for the session, but replayed it with the shared replayPending(): one
POST per entry, at most one replay window (50) per attempt. At the
~0.67s/entry measured in #4504 a 670-entry backlog needed 14 takeover
attempts of ~33s each, and pendingTokens kept growing meanwhile.

- flushForTakeover drains the session's backlog via
  /messages/batch (sendSessionMessages, BATCH_LIMIT=100), so the whole
  backlog clears in a handful of requests within one attempt. Entries
  are claimed one batch at a time; a failed batch costs a retry only for
  the entries it contained, the rest stay untouched.
- syncBranch sends the whole turn in one batch request instead of one
  POST per payload, with retryable failures queued as before.
- pi's shared dir now vendors batch-send.mjs (already used by the
  claude-code, codex, opencode and zcode plugins).

Fixes #4504

Co-Authored-By: ktz03 <2484593937@qq.com>
Co-Authored-By: jiale li <2946192893@qq.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Uued6mAGJ4jfTW4XRmNWwL

* fix(pi-extension): enqueue non-retryable batch failures and bound drain (#4702)

Address #4692 review nits from now-ing and jiale-li-orion:
- enqueueRemainder always queues on enqueue-on-failure (incl. 400/403) so the sync watermark advances
- drainSessionBacklog soft-bounded by OPENVIKING_PENDING_DRAIN_BUDGET_MS / MAX_BATCHES
- tests for watermark enqueue and maxBatches stop

* fix(pi-extension): keep batch-send drop policy, advance watermark in pi

#4702 made sendSessionMessages enqueue payloads after a non-retryable
rejection so pi's sync watermark would advance. That reverses a policy
pinned by the shared batch-send tests (poison payloads are dropped, not
queued) for every harness, and the other vendored copies were not
regenerated.

Keep the shared policy and fix the watermark where the need is: pi's
sendPayloads counts non-retryable drops as accepted, matching the
outcome replayPending() applies to such entries.

Co-Authored-By: ktz03 <2484593937@qq.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Uued6mAGJ4jfTW4XRmNWwL

---------

Co-authored-by: ktz03 <2484593937@qq.com>
Co-authored-by: jiale li <2946192893@qq.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-04 19:18:11 +08:00
Zayn Jarvis 30c5092672 fix(openclaw): strip dev dependencies from release artifacts (#4699)
OpenClaw installs extracted plugin artifacts as project roots, so npm still
resolves devDependencies under --omit=dev. Build first, then strip the field
from both npm and ClawHub packages to avoid Arborist peer-resolution crashes.
2026-09-04 18:22:11 +08:00
DuTao a320726655 增加session compile skill (#4697) 2026-09-04 17:53:55 +08:00
Zayn Jarvis db1fd7ccf6 fix(openclaw): restore package build after peer role rename (#4691)
* fix(openclaw): restore package build after peer role rename

* chore(openclaw): keep release fix minimal
2026-09-04 16:25:11 +08:00
t0saki 75be3bd0f6 fix(plugins): stop the installer aborting inside the credentials wizard (#4689)
* fix(plugins): stop the installer aborting inside the credentials wizard

The API key prompt advertises "enter = keep <masked key>", but taking it
up killed the installer before it wrote ovcli.conf or installed a single
plugin. `[ -z "$current_key" ] && WIZ_KEY=""` is prompt_connection's last
command, so a stored key makes the test false, the function returns 1,
and `set -Eeuo pipefail` unwinds the whole script from step 2 of 3. The
user sees the ERR trap fire on a line that reads like an internal
detail, an unchanged ovcli.conf, and no plugin anywhere.

Digit shortcuts fed that same branch. tui_menu and tui_choose_cli_format
confirmed on the digit itself and left the Enter pressed right after it
in the tty buffer, where the next prompt read it as an empty answer.
Picking "Volcengine OpenViking Cloud" with `2` + Enter therefore skipped
past the URL menu with its default, wrote the stray "2" as the server
URL, and hit the abort on the API key prompt -- with the key the user
then typed going nowhere. Digits now only move the cursor, which is what
tui_menu's own English hint ("1-9 jump . enter confirm") already
promised; Enter still confirms.

Both menus are drawn on /dev/tty and cannot be driven without a pty, so
the wizard tests stub tui_menu and feed fd 3 directly, and the key
handlers are pinned by reading the script.

* fix(plugins): keep an empty list element from aborting the installer

`split_csv_list` and `split_harnesses` end their loop body with
`[ -n "$item" ] && printf ...`, so an empty last element -- a trailing
comma is enough -- makes the loop, the pipeline under `pipefail`, and
the function itself exit 1. `normalize_bin_list` passes that status
straight to its caller, where it lands in an assignment and `set -e`
kills the run before the first step:

    $ OPENVIKING_CLAUDE_BIN="claude," bash install.sh --yes
    xx  OpenViking installer stopped unexpectedly.
        Exit status: 1
        Script line: 514
        Command: CLAUDE_BINS="$(normalize_bin_list "$CLAUDE_BINS_ARG" claude)"

`--claude-bin`, `--codex-bin` and their OPENVIKING_* env spellings all
reach it. `split_harnesses` has the same shape and was saved only by
every call site expanding it inside a heredoc, where the status is
discarded; give both the `if` form so neither depends on that.

Two smaller things in the same area. The ERR handler sets up
`>/dev/tty` before `2>/dev/null`, so on a machine with no controlling
terminal bash reports that redirection failing on its own line, ahead of
the diagnostic the handler exists to print. And the credentials step
tested for a url or key change but only ever printed the url pair, so
rotating just the key reported `url: <same> -> <same>`; it now names the
field that moved and masks both sides of the key.
2026-09-04 16:02:50 +08:00
Zayn Jarvis 58139b46ad fix(openclaw): make peer scope optional again and default to none (#4546)
* fix(openclaw): make peer scope optional again and default to none

Peer routing became an implicit default and then stopped being offered at
all. #2626 flipped the plugin default from peer_role=none to assistant, and
--peer-role example it documented. The flag still exists in the setup command
and in the installer, but nothing on the guided install path mentions it, so
a user installing through the skill gets assistant scoping with no visible
way to choose otherwise.

Restore the choice and put the default back to none:

- default peer_role is none again in config.ts, the setup command and the
  setup helper, so memories land under the OpenViking user unless the user
  asks for separation
- the install skill asks for the memory scope again and documents all three
  values, and its setup/installer invocations expose --peer-role
- reword the option everywhere in terms of what the user gets ("one shared
  memory" / "a separate memory per assistant" / "a separate memory per
  sender") instead of "peer identity mode"

peer_role=assistant and peer_role=person are unchanged for anyone who has
them configured; only the default and the wording move.

* fix(openclaw): rename person peer role to sender
2026-09-04 15:18:48 +08:00
t0saki 37ef554bb2 refactor(plugins): converge the harness forks back onto the shared library (#4594)
* refactor(plugins): ship each harness only the shared modules it imports

`sync.mjs` grouped its lists by how they had grown rather than by what each target imports, so three modules travelled to plugins that never load them: `setup-wizard.mjs` reached dsh and zcode, neither of which ships a setup entry point, and `async-writer.mjs` reached opencode, which its host imports in-process and so has no hook subprocess to detach a write from. Split the groups by capability — the hook set, the wizard, the stdio proxy pair, the batch sender, the async write path — and give pi its own list, so a target's entry says which capabilities it has.

`sync.test.mjs` kept its own copy of those lists, and the copy had drifted: it was missing `plugin-config`, `retryable`, `recall-compress-core` and `mcp-proxy-config`, so a stale vendored copy of any of them would have passed CI. Import the lists from `sync.mjs` instead, guarding the sync behind an entrypoint check, and add the check the duplicate could never make: every module in `lib/` is claimed by some target, and no target holds a banner-carrying file the sync no longer ships.

* refactor(zcode): capture and shape the proxy config through the shared modules

zcode sent every turn it parsed straight to the server: no length cap, no acknowledgement filter, no slash-command or injected-status guard — the four things `shouldCaptureText` does for every other harness. It re-derived the proxy config object by hand too, which is how it came to watch a narrower set of credential files than `buildMcpProxyConfig` watches.

Route both through the shared modules. The dedup key stays keyed on the raw turn, so raising the cap later never resends a turn the server already holds in truncated form.

* refactor(pi): log through the shared JSON Lines logger

pi carried two copies of a hand-written `debugLog` — one in `index.ts`, one in `sync.ts` — that appended `<ISO timestamp> <message>` lines, read `OV_DEBUG_LOG` directly, and could only be turned on through the environment. Both are the shared `debug-log.mjs` with the structure taken out: no stage field, no JSON payload, no config knob, and a spelling of the variable no other harness uses.

Use `createLogger` in both places, add a `debugLogPath` config key so the log can be turned on the way every other pi setting is, and read `OPENVIKING_DEBUG_LOG` with `OV_DEBUG_LOG` kept as a deprecated alias so existing setups keep logging.

* fix(dsh): honor syncTurns on every write path, not just capture

`syncTurns: false` gated `capture()` alone, so a read-only session still committed on `turn/end`, still committed again on dispose, and still replayed whatever an earlier session had queued. The toggle promised no writes and made three.

Gate the commit paths and the replay on it too, and document it — the README and the integration page never mentioned the key at all. A backlog queued while capture was on stays on the queue for a session that still writes.

* chore(plugins): bump the zcode and dsh plugin versions

Both changed behavior in this branch — zcode now filters and truncates what it captures and watches the full credential set, dsh now writes nothing when `syncTurns` is off — and installed copies are keyed by version.

* fix(plugins): make the installer and the plugin test matrix work in a git worktree

`resolve_self_checkout` looked for `.git` as a directory. A linked worktree keeps it as a file pointing at the real gitdir, so `CHECKOUT_DIR` stayed empty there: `--source dev` resolved the marketplace to `/examples` and failed outright, and every other path fell through to `remote`, which clones from GitHub. On this machine that turned six of the thirteen installer tests red and made `release-marketplace.test.mjs` hang for twenty minutes on the network — long enough that the CPU starvation failed an unrelated recall timing assertion too. Test for existence instead of for a directory.

Three test files were never run by CI, so nothing noticed that one of them had gone stale: the pi wiring assertion still required `new RecallManager(...)` to end at the session-id getter, which stopped being true when the recall ledger was added a fourth argument. Assert only the getter, and register all three files in `pr.yml` — the matrix now covers every `*.test.mjs` under `examples/`.
2026-09-04 13:59:29 +08:00
t0saki 1d89f8d465 feat(plugins): derive the workspace peer from git, and let a repository carry its own config (#4595)
* fix(plugins): stop dropping ovcli.conf's plugin section, and unrot the sync test

`ov config add|edit`, the config wizard and `ov config switch` all rebuild
ovcli.conf from the `Config` struct, which has no `plugin` field and no
catch-all — so every write silently deleted the whole `plugin` section the
memory plugins own. `write_config_file` now carries over the top-level keys
`Config` does not model, `save_edited_config` reads them from the old name on a
rename, and `activate_config` keeps the active file's. Modeled keys are
deliberately not carried over: one that is `None` was cleared on purpose.
`KNOWN_CONFIG_KEYS` decides what counts as modeled, guarded by a test that
fails when a struct field is added without listing it.

sync.test.mjs kept its own copies of the target lists and they had drifted —
dsh, opencode and agent-plugins were each missing modules the sync ships, so a
stale vendored file passed CI. It now imports the lists from sync.mjs (whose
`main()` moved behind an entrypoint guard) and additionally fails on a module
no target claims or a banner-carrying orphan no target lists.

Also in this hygiene pass:

- postRecall dropped `peer_scope` on any 400/422, silently widening recall from
  the caller's own peer to the whole user root. It now retries only on an
  unknown-field rejection, remembers the downgrade so every turn stops paying
  for a rejected request, and both doctors warn while that memo is live.
- The session peer pins (Claude Code's `ws-peer-*.json`, Codex's
  `workspacePeerId`) carry a version, so a pin written under one derivation
  rule cannot outlive it. Derivation is unchanged, so today they re-derive to
  the same value.
- recall-session-wiring.test.mjs was never registered in CI and had rotted
  against a fourth RecallManager argument; assertion fixed and registered.

* feat(plugins): layered workspace configuration

A workspace can now carry `<root>/.openviking/config.json`, which a team
commits, and `config.local.json`, which stays private; a per-machine registry
under `~/.openviking/workspaces/` sits above both so the user keeps the last
word over any repository they clone. All three share one schema and one merge,
and they slot in exactly where ovcli.conf's `plugin` section already does, so
`OPENVIKING_*` still wins over everything.

Three modules, synced to all seven plugin targets:

- workspace-identity.mjs finds the workspace root and reads git's own idea of
  what the repository is called, using only filesystem reads. No `git`
  subprocess: hooks are fresh Node processes on prompt-level paths with budgets
  as tight as Codex's 3s SessionEnd, and this keeps working where git is absent
  from PATH or would refuse the repo over dubious ownership. Worktrees converge
  through `commondir`; submodules stay separate; `$HOME` and `/` are never
  workspace roots.
- workspace-config.mjs discovers, parses, filters and merges the layers, and
  records per-key provenance — which layer won and what it covered up.
- workspace-registry.mjs keeps one file per workspace rather than one listing
  them all, so concurrent hooks cannot lose each other's writes, and treats a
  path whose git identity has changed as a miss rather than inheriting the
  previous repository's peer.

These files are trusted without a prompt, because a hook is non-interactive and
any approval gate degrades into "run one command per workspace first". What is
refused instead is structural and costs nobody anything: connection and
credential keys are stripped loudly, `${VAR}` is never expanded, and
`cli_config_profile` — which decides which credentials reach which server — is
registry-only and name-only. What a committed file switches off is announced in
doctor rather than blocked.

An adversarial review pass over these three modules found 18 defects, all fixed
here and each now covered by a test. The one that mattered: `JSON.parse` keeps
`__proto__` as an own property, so a 128-byte committed file could write
straight into `Object.prototype`, and since `process.env` reads through the
prototype chain and the environment outranks ovcli.conf, that set
`OPENVIKING_URL` and `OPENVIKING_API_KEY` for the whole process — silently
shipping the user's real API key to an attacker's host. Prototype keys are now
dropped with a warning in both the strip and the merge. The rest: unbounded
recursion (a 4KB file could take out every sibling layer), the identity cache
storing a remote's embedded token at 0644, worktrees under a directory named
`modules` misread as submodules, the registry's negative-evidence check being
inert in its only caller, `Number(null)` pinning cost knobs to a bound, and
provenance lying when two layers disagree about a key's type.

`.gitignore` no longer ignores all of `.openviking/`, which would have stopped
a team's config.json from ever being committed; doctor warns when a workspace
still does. The schema maps `capture.commit_token_threshold`, matching the knob
the loaders actually read — the RFC's example named a turn-based one that does
not exist.

* feat(plugins): derive the workspace peer from git, configurably

The peer a workspace writes its memories under was the working directory with
every non-alphanumeric byte turned into a dash. That made the identity an
accident of where the repository happened to sit: a clone on another machine, a
rename, a worktree, or simply `cd examples/` each minted a separate, empty
namespace, and there is no server-side rename or merge to recover from it.

The default is now git's own idea of the repository. `peer.source` decides the
rule and reads from every layer — `OPENVIKING_PEER_SOURCE`, ovcli.conf's
`plugin.peerSource`, or `peer.source` in a workspace file:

- `git` (new default) ≡ `["{git_remote}", "{git_root}", "{cwd}"]` — the
  normalized origin, else the repository root, else the working directory. No
  preset adds a prefix; a path-derived id already starts with `-` on POSIX, so
  it cannot collide with a remote-derived one.
- `cwd` — the old rule, byte for byte.
- `none` — send no peer. `OPENVIKING_WORKSPACE_PEER=0` still means this.
- Any template, or a list tried in order, over `{git_remote}` `{git_root}`
  `{cwd}` `{dir}`. Substitution is all-or-nothing: an empty variable falls
  through to the next template rather than leaving a half-formed shared id.

So `/Users/x/Dev/OpenViking/examples/codex-memory-plugin` with origin
`git@github.com:volcengine/OpenViking.git` is `github.com-volcengine-openviking`
from any subdirectory, worktree, machine or clone. Every clone of one repository
shares one peer; a fork has a different origin and stays separate, and
`gh pr checkout` of someone else's PR does not change origin, so reviewing does
not move a session's memory.

Nobody has to migrate. The pre-git id is always recomputable locally, so
`resolveEffectivePeerId` returns it alongside the effective one and recall still
reaches it: under the default `peer_scope: "all"` the server's cross-peer sweep
already covers it for free, and under `"actor"` — where that sweep is off by
definition — the plugin asks that peer separately, as itself, which is cheaper
and reaches more than a bare cross-peer read would. There is no deadline on
this. Wired through all five recall paths; doctor names the previous peer and
says which of the two is carrying it.

`source` keeps its three values because five call sites compare it against the
literal `"workspace"` to decide whether a session pin may be reused; the new
`origin` field names the template that actually produced the id, and doctor
prints it. Both session pins bump their version, so a pin frozen under the old
rule cannot outlive it.

* docs: the workspace peer comes from git, and a workspace can carry config

Every page that described the peer as the working directory with its
non-alphanumerics dashed now describes `peer.source` and the git default, with
the presets, the template variables, the clone-vs-fork identity semantics, and
why no migration is required. The capability reference gains the three new
shared modules; the client configuration page gains a Workspace Configuration
section covering the two workspace files, the per-machine registry, the
precedence table, the v1 schema and what a workspace file may not set; the
Claude Code and Codex integration pages, which never mentioned peer derivation
at all, each gain a short section. All zh mirrors follow.

`examples/schemas/workspace-config-v1.json` is the schema the `$schema` key in
a workspace file points at, and `examples/workspace-config.example.json` is a
file to copy. `ovcli.conf.example` shows `plugin.peerSource`.

Two facts worth stating plainly, both verified against the loaders rather than
assumed: `OPENVIKING_PEER_SOURCE` and the workspace-file layer are read only by
the Claude Code and Codex plugins today, so the other harnesses run on the
default and their pages document the config key rather than an env var that
would be inert; and pi and dsh compute their legacy id from the process cwd,
so their pages promise dual-read only under the default `peer_scope: "all"`.

* fix(mcp): stop the proxies from guessing a peer out of their launch directory

Three MCP proxies keyed the actor peer off `process.cwd()`, which for a
long-lived server started from a static MCP config is the directory the harness
happened to launch from — often the plugin's own. The Codex proxy already
refused this and had a test forbidding it; the rule now holds for all of them
through the same `resolveMcpActorPeerId`.

The plan called for the parent process to inject `OPENVIKING_PEER_ID` at launch
instead, following dsh's `mcp.mjs:27`. That only works for dsh: Claude Code and
agent-plugins are launched from a static `.mcp.json`/`mcp.json` with no
environment block, and OpenCode's `createOpenVikingMcpConfig` builds a command
and args with nowhere to put one. So the fix is to stop guessing rather than to
guess better — a proxy sends no actor peer, which is broad recall, the default.

`resolveMcpActorPeerId` now warns and widens where it used to throw. Refusing to
start took away every memory tool because a scope preference could not be
honoured, which costs the user far more than the wider search does; the warning
says which two settings would scope it.

* test(pi): follow the git-derived peer default rather than pinning the cwd id

* fix(plugins): warn on the camelCase spelling of a connection key too

The projection into harness knobs is an allowlist, so `apiKey` in a workspace
file could never take effect — but it vanished without a word, which reads as
acceptance. It is refused by name now, like its snake_case twin.

* feat(cli): ov workspace show, and ov peer link|migrate|forget-previous

`ov workspace show` answers "which layer actually set this" the way
`git config --show-origin --show-scope` does: the workspace root and how it was
found, the git remote, every template variable, the effective peer and the
template that produced it, each config layer with whether it applied, and per
key the effective value plus everything it shadowed.

That question matters here because three languages read this configuration and
each could drift. So the Rust reader is not a paraphrase of the JS one — the two
were run side by side over the identity helpers, the merge with full provenance
trees, the file-read rules and the registry's raw bytes, and made byte-identical.
That comparison paid for itself: it caught `serde_json::Map::remove` being a
swap remove under `preserve_order`, which reshuffled a registry file the JS half
reads on every rewrite.

It also caught the divergence that would have made the command a liar. ovcli.conf's
`plugin` section speaks the flat knob names a harness loader reads (`peerSource`,
`recallLimit`); a workspace file spells the same settings nested (`peer.source`,
`recall.max_items`). Both are one chain in `loadPluginSettings`, and the port had
merged the flat file into the nested tree, so `plugin.peerSource: "cwd"` in
ovcli.conf left `ov workspace show` reporting the git-derived peer while every
plugin sent the cwd-derived one.

`ov peer link <id>` pins a peer for this workspace in the registry — the way out
of a fork that should share the upstream's memory, or a legacy id worth keeping.
`ov peer migrate` moves a peer's memories and resources with the fs mv API,
reporting the plan by default and requiring `--apply`; the server has no merge
semantics, so a collision is refused with the colliding path rather than
overwritten, and a listing that fills its limit aborts rather than planning from
a truncated view that could hide one. `ov peer forget-previous` clears the
recorded ids.

`workspace show`, `peer link` and `peer forget-previous` are local and do not
require ovcli.conf; `peer migrate` talks to the server and does. Both config
gates and the hand-rendered help are registered, with a test pinning the gates
against each other.

A test now reads FORBIDDEN_KEYS, REGISTRY_ONLY_KEYS and FREE_FORM_SECTIONS out
of the JS module and compares them to the Rust constants, because those lists
are what someone fixing a bug in one language edits — and they had already
drifted once while this was being written.

* build(cli): record the sha2 dependency edge in Cargo.lock

Already vendored for other workspace members; ov_cli now uses it for the
registry slot hash.

* test(plugins): the fixtures the plan named that were still missing

A moved or renamed repository keeping its identity is the change's whole point
and had no test; a shallow clone was worth pinning because it is exactly what
the rejected root-commit scheme could not answer; and the registry's
read-modify-write window between two hooks of one session is now written down as
a test rather than only as a comment.

* fix: the defects a plan review turned up

An independent review against the plan this branch was built from found ten
real defects. Each was reproduced before being fixed and is now covered by a
test.

The four that broke a promise the feature makes:

- The workspace config layer was resolved from the hook process's own working
  directory, not from the `cwd` on its stdin payload — which is the
  authoritative one. A hook started in one repository while the session sits in
  another applied the wrong `.openviking/config.json`: its peer, its bypass
  patterns, its `capture.enabled`. `loadConfig` now takes the directory, and
  every hook that receives one re-resolves with it. Late re-resolution is safe
  precisely because connection and credential keys are structurally forbidden
  in a workspace file, so `baseUrl` and `apiKey` cannot move under an already
  built client — the gates that a workspace can switch off moved below the
  parse so they are decided on the right config too.
- Codex threw away a workspace file's `peer.id`: it returned the credential
  chain's peer verbatim, so `{"peer": {"id": "team-a"}}` did nothing. Claude
  Code had always honoured it.
- Claude Code's session pin returned only the id and source, dropping
  `legacyPeerId` — so from the second hook of a session onward, dual-read
  stopped asking the peer that holds everything written before the derivation
  changed. Silently, and exactly where it mattered.
- `ov peer migrate` read the source peer with the actor-peer header set, which
  the server refuses for another peer's path, and treated every `stat` error as
  "does not exist". The common case — an `actor_peer_id` in ovcli.conf — got a
  cheerful "Nothing to migrate" instead of a 403. It now uses a client with no
  actor peer and tells a real error apart from an empty source.

The rest:

- A directory outside any repository is a workspace again. It had no root at
  all, so a `.openviking/config.json` there was ignored entirely. `$HOME` and
  `/` are still never roots, now judged on the starting directory rather than
  on where the upward walk stops, and `git_root` stays empty outside a
  repository so the `git` preset still falls through to `{cwd}`.
- The registry slot is keyed on identity, not path. Two linked worktrees of one
  repository are one workspace — one peer, so one set of settings and one
  `ov peer link` — and keying on the checkout path split them in two. This also
  makes crossing two repositories physically impossible rather than merely
  detected. (Their `config.json` files still follow each checkout; those are
  files on a branch.)
- git folds section and key names to lower case, so `[Remote "origin"]` with
  `URL = …` is a remote `git config` reads and we did not. A quoted subsection
  stays case-sensitive.
- `min_client_version` warns instead of being silently kept as data, and still
  never blocks.
- `cli_config_profile` was validated and then never used. It now selects
  `~/.openviking/ovcli.conf.<name>` before credentials resolve — registry-only,
  name-only, and a hard error when the profile is missing, because quietly
  authenticating somewhere the user did not choose is the failure the key
  exists to prevent.
- `ov workspace show` is exempt from the language gate. It is a diagnostic and
  has to work on a machine where no language was ever chosen; `peer link`,
  `migrate` and `forget-previous` mutate state and still gate.
- `ov peer link` records the peer it replaced, so a later `migrate` with no
  `--from` finds it instead of falling back to a recomputed cwd id.

Two more the plan asked for that were missing: doctor now checks the knobs
*inside* ovcli.conf's `plugin` section — it was on the allowlist, so until now
`peerSorce` sat there doing nothing with no complaint — and suggests the key
you probably meant. A test derives the known-knob set from what the two loaders
actually read, so the list cannot rot into one that rejects a real knob; it
caught a missing entry the moment it was written.

The RFC is archived at docs/design/, with the three claims implementation
disproved corrected in place: the knob is `commit_token_threshold`, `__self`
and `ext-` are not reserved server-side, and a worktree converges its identity
rather than its config file.

* revert(cli): withdraw ov workspace and ov peer from this branch

The command surface these two files added was larger than the feature they
served: 4792 lines of Rust for `ov workspace show` and
`ov peer link|migrate|forget-previous`, against a change whose whole point is
what the hooks send. None of it had reached a user-facing document — only the
RFC named it — so it goes back out whole and the branch becomes a plugin
change plus one CLI bug fix.

Restored from the branch's merge-base rather than from origin/main, since main
has moved on since the branch was cut and those commits are not this branch's
to carry. `sha2` was pulled in only by `workspace.rs`, so its dependency edge
leaves with it.

What stays is `config.rs` and `config_wizard/store.rs`: `ov config add|edit`
dropped the whole `plugin` section because the wizard round-tripped the file
through a typed struct, and that fix has nothing to do with the withdrawn
commands.

The registry under `~/.openviking/workspaces/` stays too, as a layer the
plugins read. Nothing writes it for now; `ov-memory-doctor` prints the path it
expects, and the file is small enough to create by hand. A writer can come back
on its own merits.

* fix(plugins): derive a peer only inside a git repository

Codex desktop opens a directory per task — `~/Documents/Codex/<date>/<slug>/` —
and none of them is a repository. The `git` preset ended its fallback chain at
`{cwd}`, so every one-off task minted its own empty peer, and each new one
started with no memory. Nine such directories here, nine peers.

There is nothing app-specific to read: the state file that lists those threads
is Electron-private, a megabyte wide, desktop-only, and would have to be parsed
inside SessionEnd's 3-second budget. The signal that generalizes is structural
— the directory is not a repository, and nothing in it says it is a project.

So the default chain is now `["{git_remote}", "{git_root}"]` and stops there. A
directory that is neither a repository nor marked gets no peer at all, and what
is remembered in it goes to the user-level space, which is where it went before
peers existed. Deriving an identity from a bare path is what `peer.source:
"cwd"` is for, and it is a word away.

Naming such a directory is the other half. `findWorkspaceRoot` now also stops
at a directory holding `.openviking/config.json` or `config.local.json`, so a
marker file works from any depth below it, the way a repository does — and when
that marker sits inside a repository the git variables still resolve to the
enclosing repository, so marking a subdirectory of a monorepo does not split
the default peer. `{git_root}` is the repository's root, `{dir}` the workspace
root's name whichever made it one.

Nothing moves. When no template resolves, the pre-git id is still computed and
returned as `legacyPeerId`, so `peer_scope: "actor"` keeps asking for it and
`"all"` keeps sweeping it.

Two doctor bugs fell out of the same walk: `checkWorkspace` read `git.kind`
unconditionally and threw wherever there was no repository, and the peer block
warned "set peer.source to git" at a directory where `git` is exactly what is
already set and correctly resolves to nothing. It now says why no peer is sent,
and prints the file to create.

* docs(plugins): say how to give a directory its own peer, to users and to agents

The behavior change is only useful if the reader can act on it, and two kinds
of reader have to: the person whose scratch folder stopped having a memory, and
the coding agent they ask about it.

`docs/{en,zh}/configuration/02-client.md` is the one place that spells the rule
out, and everything else links to it. It gains "Give a Directory Its Own Peer",
which opens with the file to create and then the ladder above and below it;
"By Situation", eight rows from fork to throwaway folder; and "Recall
Isolation", which separates where memories are written from what is read back,
names the server's per-category penalties, and states the cost of sending no
peer outside a repository — a user-level memory is read at full weight in every
project afterwards.

The eight integration pages, both capability references, six plugin READMEs,
the changelog and the schema stop promising a fallback to the working
directory. Checking those claims against the loaders turned up one that was
never true: opencode, dsh and pi do not read workspace files at all, so a
`peer.id` written for them does nothing. Said plainly rather than left to be
discovered.

For agents, `openviking-memory/SKILL.md` gains ten lines on where memories are
filed — it is the skill that fires when someone asks why a folder has no
project memory, and it had nothing to say — and both `ov-memory-doctor`
references gain a table from what the user says to the exact key to write.
`ov-memory-doctor` prints the same snippet, so an agent that runs it needs no
further reading.

One snippet has to be identical in twenty places for any of this to hold, so
`WORKSPACE_PEER_HINT` is a constant the report builds its line from, and
`peer-guidance.test.mjs` asserts it appears verbatim wherever it is promised,
that no page still spells the retired chain or names a command this branch
withdrew, and that every variable the canonical page documents is one the code
substitutes. It asserts no prose: rewording a page must not turn it red.

* fix(plugins): reject an unrecognized peer.source instead of using it as an id

`peer.source` accepts a preset name, a template, or a list of templates, and
anything that is not a preset was treated as a template. A template with no
`{...}` in it renders to itself, so a typo became the peer: `"Git"` wrote every
memory under a peer literally named `Git`, and `"gti"` under `gti`. Silently —
the wrong namespace is indistinguishable from an empty one until someone
notices their project has no memory.

A bare string that is neither a preset nor contains `{` now warns and falls
back to the `git` default. A list is still taken at face value: writing one is
explicit enough that a typo inside it is a different kind of mistake.

The warning travels through an optional `onWarn` callback falling back to
stderr, matching `resolveMcpActorPeerId` in `mcp-proxy-config.mjs` — there is no
warnings array in reach, because `resolveEffectivePeerId` is called from the
hook runtime and from four harness config loaders, none of which thread one.

* docs(plugins): say what each harness can actually do with a peer

The peer documentation promised the same thing everywhere, but only the Claude
Code and Codex plugins read workspace configuration files. `loadPluginSettings`
is called from exactly two loaders; the other harnesses build their config from
their own file plus the environment. So a reader following the docs under pi,
dsh, opencode or cursor would create `.openviking/config.json` and watch it do
nothing.

The skill is the sharpest case: `openviking-memory/SKILL.md` is synced to
cursor and dsh as well, and it told an agent to write that file. An agent would
have done it, reported success, and changed nothing. It now names the two
harnesses that read it and points everyone else at `OPENVIKING_PEER_ID`.

The integration pages had started teaching the recipe and then retracting it in
the same sentence, which is worse than not mentioning it; they now carry the
one instruction that works there, and link to the canonical section for the
rest. The capability reference gains the same qualification, next to the
paragraph that already says only two harnesses read those layers.

Two smaller corrections. The doctor references had the same question answered
twice, once in the peer table and once in the troubleshooting table 140 lines
below; the troubleshooting row survives, since it carries a diagnostic column.
The RFC still archived implementation notes for the CLI this branch withdrew,
which would read as a description of commands that exist.

`peer-guidance.test.mjs` guards this alignment, and had two flaws of its own: it
swept the changelogs, which are generated from GitHub Releases and would go red
on a release note nobody wrote by hand, and it sliced a page between two
headings with `indexOf` without checking either was found — renaming the closing
heading would have silently scanned to end of file.

Also here, because it is the same kind of mismatch: the zcode MCP proxy sent an
actor peer under broad recall, where the other proxies leave the header unset.
It now routes through `resolveMcpActorPeerId` like they do. The dsh proxy
deliberately does not — its parent process resolves the peer per session and
injects it into the child environment, so it is not guessing at a launch
directory, and that reason is now recorded next to the line.

* fix(plugins): ship and install only the shared modules a plugin imports

Three new modules were fanned out to all seven plugin directories in one hunk,
and only two plugins call them. That left dead weight in five directories, and
it broke three installs.

The install is the part that mattered. `install.sh` copies a hand-written list
of shared files into `~/.openviking/agent-integrations/memory-plugin-shared/lib`,
where cursor, TRAE and TRAE CLI import from. The list names `workspace-peer.mjs`
but not `workspace-identity.mjs`, which `workspace-peer.mjs` imports — nor
`workspace-config.mjs`, which identity had come to import for three filename
constants. Copying exactly that list and importing the hook runtime fails with
ERR_MODULE_NOT_FOUND, so every hook of those three harnesses would have died on
startup. Nothing tested the list.

`CONFIG_DIR_NAME`, `TEAM_FILE` and `LOCAL_FILE` now live in
`workspace-identity.mjs`, which is where the walk that recognises a marked
directory needs them; `workspace-config.mjs` imports and re-exports them, so no
call site moves. Identity has no library-internal dependency left, which is what
makes the installed set closed at fifteen files instead of pulling the whole
configuration layer along behind it. `install-lib-closure.test.mjs` derives both
sides — the list parsed out of the shell script, and the transitive imports of
the three entrypoints — and fails in either direction, so neither a new
dependency nor a stale entry can go unnoticed again.

With identity standing alone, the fan-out can follow what is actually imported.
`sync.mjs` moves from arrays chained by spread — where the harness that does not
need a file is often the one the array is named after — to explicit per-target
lists. `plugin-config.mjs`, `workspace-config.mjs` and `workspace-registry.mjs`
leave dsh, pi, opencode and zcode, none of which import them; all four workspace
modules leave agent-plugins, whose only importer was deleted a half hour after
they arrived and whose peer is environment-only by design. That is about 4500
lines of vendored code that said something the code did not do.

The registry loses its write path in the same spirit. `writeEntry`,
`rememberPreviousPeer` and `listEntries` had no caller outside their own tests:
the CLI that would have written them was withdrawn from this branch. `readEntry`
and `entryPath` stay, because a hand-created entry is still read and the doctor
still prints where to put one. `cli_config_profile` goes with them — the whole
mechanism, down to the documentation that described it, since nothing ever
resolved a profile through it.

This is not a new policy. `HARNESS_KEYS` already carried the rule in a comment,
added the day after the same speculative fan-out happened in August: add a key
as its loader starts calling `loadPluginSettings`, not before, so the section
never promises a knob that silently does nothing.

* fix(dsh): thread peerSource into the per-session peer

The integration page documents `peerSource` in dsh's Cordis patch, but `stateFor` never passed it to `resolveEffectivePeerId`, so the key resolved to nothing and every dsh session ran on the default derivation. Pass it, and keep the pre-git id alongside so dual-read reaches memories written before the default changed.

* docs(plugins): correct the shared-layer counts after the distribution changed

The capability reference still described the pre-branch distribution: 18 library modules against 23, per-target counts from before each target stopped receiving the workspace configuration layer, and `workspace-config` / `workspace-registry` listed as reaching every JS harness when only claude-code and codex load them. It also still said `cli_config_profile` was registry-only, and that the registry is written for you.

* docs(rfc): lead with a TL;DR of the workspace config and peer source proposal

* feat(plugins): let peer.source name the harness with {harness}

The peer templates could describe where a checkout sits but never which
agent was running in it, so one repository could not keep a separate
memory per agent even when its user wanted that. The harness name was
already in every config, only baked into the User-Agent string.

No preset uses the new variable: sharing one project memory across
agents is the more useful default, so splitting stays opt-in via a
template such as "{git_remote}-{harness}". It is composed at render
time rather than in the workspace identity, whose result is cached
under a cwd-only key that two harnesses in one directory would share.

* docs(rfc): record git_branch and peer.command as directions, not deliverables

* chore(plugins): resync the openclaw vendored recall-core after the rebase

* test(opencode): follow resolveEffectivePeerId's widened return shape

Also mark the openclaw shared copies generated, the way every other sync
target already is.
2026-09-04 12:50:43 +08:00
now-ingandmac da94ac1afd fix(opencode-plugin): fall back to config peerId when shared credentials define none (#4632)
loadConfig() unconditionally assigned config.peerId from resolved shared
credentials, and applyLegacyConnection() (the only reader of the config
file's peerId) is skipped whenever those credentials exist. When ovcli.conf
or environment variables authenticate without actor_peer_id, the project
config's peerId was silently dropped, so sessions and memories landed in
the shared user tree instead of the peer-scoped tree.

Apply the file config's peerId as a fallback when shared credentials carry
no peer, keeping the documented precedence: shared credentials >
extension config peerId > workspace-derived peer. Same defect class as
#3649 (Pi extension, addressed by #3653 for Pi only).

Fixes #4487

Co-authored-by: mac <bishopapril850965@yahoo.com>
2026-09-04 12:05:24 +08:00
now-ingandmac dfb4e324d2 test(codex-plugin): accept both convergences of the Stop/SessionEnd marker race (#4668)
The race test asserted state.ovSessionId === null after running a Stop
worker and a session-end worker concurrently, but the system has two
legal outcomes. auto-capture clears end markers older than its own start
(resume semantics), so when the SessionEnd parent writes its marker just
before the Stop hook starts — the exit race this test spawns — the
session-end worker finds no marker under the lock and exits superseded,
leaving the Stop worker's live session for the SessionStart sweep's
idle-TTL pass instead of an inline commit.

Assert the real invariants instead: exactly-once sends, cursor
consistency, only the derived cx-* session may stay live, and no stale
end marker survives either convergence.

Co-authored-by: mac <bishopapril850965@yahoo.com>
2026-09-04 11:59:11 +08:00
ktz03 cf18dfb479 fix(dsh-plugin): run profile and recall in parallel on pre-step (#4643)
agent/pre-step currently gates user/message push in dsh-agent-loop. Overlap profileMessage and recallMessage to cut waterfall wall time.

Related to #4515.
2026-09-04 11:57:48 +08:00
John Roweandjr_blue_551 094b76f24b fix(pi): tolerate hosts without buildContextEntries (#4653)
Co-authored-by: jr_blue_551 <jr551@github.com>
2026-09-04 11:50:03 +08:00
ktz03 a7cbf93214 fix(opencode-plugin): isolate flushAll session failures (#4608)
Fixes #4490

Wrap each flushSession in try/catch so one failed session during dispose does not abort later sessions.
2026-09-03 11:25:29 +08:00
Pedro Perez 9c9ad8c694 fix(opencode-plugin): make messageID optional in prependSyntheticRecallPart (#4601)
The chat.message hook in opencode's plugin API declares messageID as
optional (messageID?: string in @opencode-ai/plugin/dist/index.d.ts).
When opencode invokes the hook without a messageID, the null check in
prependSyntheticRecallPart causes it to return false silently — the
recall block is discarded, no context is injected, and no error is
logged. This means the autoRecall feature silently fails for all
opencode versions that don't pass messageID.

Fix: generate a timestamp-based fallback for messageID when absent.
sessionID remains the only hard requirement (it is always provided
by the opencode API).

Verified against opencode 1.18.26: synthetic recall parts are now
injected on every user message (confirmed via opencode session DB).
2026-09-03 11:25:25 +08:00
starslittle 4e3770d4f2 feat(openclaw): assemble auto-recall context server-side (#4450)
* feat(openclaw): assemble auto-recall context server-side

* refactor(openclaw): reuse shared context search contract
2026-09-02 22:54:28 +08:00
t0saki cabd34929d fix(codex): stop a stale takeover handing the lock to two takers (#4591)
`withSessionLock` aged a lock by its directory mtime, but the taker that
won a takeover stamped the `owner` file before refreshing that mtime. A
racer that lost the stamp retried immediately, read the still-old mtime,
and `claimStaleLock` renamed the live winner's stamp aside to install its
own — two holders in the critical section at once. On Linux the retry
reliably lands inside that window, which is why `plugin-tests` failed on
essentially every PR.

Age the lock by its own stamp instead, and settle takeovers with a claim
directory derived from the state the taker read: everyone who saw the
same dead lock races one atomic `mkdir`, exactly one wins, and the losers
re-read a lock that now carries the winner's fresh stamp. Neither the
directory nor the stamp is ever momentarily absent, so no racer can slip
in alongside the taker.

Verified with 8 racers over 30 rounds: 4 concurrent holders before, 1
after.
2026-09-02 19:14:41 +08:00
t0saki 1f90903283 fix(skills): shorten over-long skill descriptions and guard the limit in sync tests (#4560)
* fix(skills): shorten skill descriptions over the 1024-char limit

The ov-memory-doctor (Claude Code + Codex) and install-openviking-memory
skill descriptions exceeded the 1024-character maximum, so the skills were
rejected at load time. Trim them while keeping the trigger phrases.

* test(skills): guard skill description length and sync ov-experience-memory

Add ov-experience-memory to the skill sync targets so its per-plugin copies
cannot drift, and assert every shipped SKILL.md keeps its frontmatter
description under the 1024-character loader limit.
2026-09-01 12:41:59 +08:00
Xinmin Zeng 96b3058769 fix(pi): respect peerId from extension config (#3653) 2026-09-01 11:21:06 +08:00
NoxOSandEl Che 241af176bf fix(pi-extension): don't block session_start on the OV server chain (#4506)
session_start awaits start(ctx), which runs a sequential chain against
the OpenViking server — health check, session ensure, pending replay,
and the session profile build (system status, fs ls, content read,
two recursive memory ls calls). Against a remote server this measured
~1.9-2.5s on every pi startup, all of it blocking the session from
becoming interactive.

start() is memoized via startPromise, so session_start can fire it and
let before_agent_start await the same in-flight chain: the first turn
still gets profile injection and recall, and behavior on failure is
unchanged (connected stays false; before_agent_start skips injection).
Startup on a real remote-server setup dropped from ~6.0s to ~2.4s.

Validated with a live turn: reply renders, turn_end sync lands, and
OV_DEBUG_LOG shows the chain completing in the background.

Co-authored-by: El Che <gaodes@gmail.com>
2026-09-01 11:14:38 +08:00
Zayn Jarvis 575a366f92 fix(openclaw-plugin): remove legacy provider auth metadata (#4511) 2026-08-31 17:36:00 +08:00
t0saki 7200cdb176 feat(codex): commit on SessionEnd hook, keep SessionStart sweep as fallback (#4429)
* feat(codex): commit on SessionEnd hook, keep SessionStart sweep as fallback

Codex ships a SessionEnd hook since rust-v0.145.0. The plugin now commits the OpenViking session from that hook (marker + detached worker within the 1s/3s budget, catching up turns the last Stop never sent) and reduces the SessionStart active-window heuristic to an ended/idle sweep. Adds a per-session lock so the Stop worker, PreCompact, SessionEnd worker and sweep no longer clobber each other's state writes, stops an unreadable transcript from resetting the capture cursor, and merges the three copies of the HTTP/transcript helpers into scripts/ov-session.mjs.

* fix(codex): guard SessionEnd partial catch-up, make the sweep catch up, token the end marker, own the lock

Review follow-up for #4429: SessionEnd no longer commits after an incomplete catch-up (tail turns were archived away); capture hooks record transcriptPath so the SessionStart sweep catches up before committing; the .ended marker's timestamp acts as a token that the SessionEnd worker and the sweep re-verify under the lock, and Stop/PreCompact/resume only clear markers older than their own start; withSessionLock stamps an owner file and takes over stale locks by atomic rename with an inode check so a taker never moves a lock a racer just created. Plugin 0.8.1.

* fix(codex): catch up ended sessions with no live id, guard unreadable transcripts, make the end marker and the lock takeover race-free

Four defects found reviewing the SessionEnd commit path:

- The SessionStart sweep cleared the `.ended` marker of a state with no live
  ovSessionId before taking the lock, so a session PreCompact had released and
  that then produced more turns lost its tail when its SessionEnd worker died.
  A marker now always enters the lock, and only a catch-up that finds nothing
  new may clear it.
- catchUpTurns reported an unreadable transcript as an empty one, so all three
  callers committed and archived a session whose tail they never saw. It now
  reports `unreadable` and they keep the session live for a later retry.
- clearEnded read the marker's timestamp and then removed the path, deleting a
  marker a concurrent SessionEnd had just written. The timestamp moved into the
  filename, so a conditional removal targets an immutable path.
- session-state.test.mjs was missing from the CI test list.

Also replaces the lock's stale takeover: renaming the directory aside left the
lock path momentarily absent, which let another racer's mkdir succeed next to
the taker. Takers now race for the `owner` file inside the directory instead.

* fix(codex): make the end marker's generation unique within a millisecond

Date.now() alone is not a generation: two SessionEnd hooks in the same
millisecond produced the same marker path, so the first one's conditional
removal took the second one's marker with it. The marker is now created
exclusively and its timestamp bumped until that succeeds; bumping rather than
randomizing keeps the names ordered, which is what the `before` cutoff
compares. The race test drops its artificial 2 ms gap and runs 500 iterations.
2026-08-31 12:23:33 +08:00
txfangandchrisfang 3123e8d8ff feat(ragfs): unify CacheRuntime and Redis-backed CacheFS/QueueFS (#4353)
* feat(ragfs): unify cache runtime providers

* refactor(ragfs): remove bundled external cache providers

* refactor(ragfs): unify cache runtime and redis backend

* chore: sync contributing guide with main

* fix(cache): harden Redis runtime compatibility

* refactor(cache): gate memory mock behind test feature

* fix(cache): harden Redis runtime integration

---------

Co-authored-by: chrisfang <chrisfang@noreply.gitcode.com>
2026-08-31 11:10:08 +08:00
Zayn JarvisandClaude Fable 5 2c88269d54 fix(openclaw-plugin): declare OpenClaw 2026.8.1 durable-turn contract (#4465)
OpenClaw >=2026.8.1 degrades any context engine to "legacy" every turn
unless info.transcriptSemantics declares currentTurnFence +
turnAdvancementIdempotency and the engine implements commitTurn.
Declare both and add an idempotent commitTurn ack keyed by
advancementKey; capture stays in afterTurn, which the host still calls.

Fixes #4103


Claude-Session: https://claude.ai/code/session_012TsbqPJYxEqGWrBqYeodWR

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-30 14:31:07 +08:00