15 Commits
Author SHA1 Message Date
t0saki b7d0415c24 feat(plugins): skill catalog and openviking-skills for the memory plugins (#5161)
* refactor(skills): install skills through one shared helper

POST /api/v1/skills kept its whole install loop (source resolution, per-skill
install, source metadata, list_only) inline in the route. Move it into
openviking/server/skill_ingest.py:install_skills so the MCP add_skill tool
and signed skill uploads can reuse the exact same code path. The REST
route's behavior is unchanged.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* feat(mcp): add an add_skill tool

MCP clients had no way to create a skill: write refuses the skills/
subtree (_USER_MANAGED_SUBTREES) and add_resource validates its target as
a resource. Agents that should keep skills in OpenViking could read them
but never add one.

add_skill takes either the full SKILL.md text (data) or a path. A Git or
GitHub tree URL installs through the same source resolution as REST, with
skills=[...] to pick from a multi-skill repository and list_only to
preview it. A local SKILL.md, directory, or zip gets the add_resource
treatment: the tool mints a one-time upload token, now tagged kind="skill"
with the target root, selection and list_only, and the signed temp_upload
installs the file as skills instead of ingesting it as a resource.
target_uri="viking://agent/skills" shares the skill with the account.

All three paths (REST, MCP inline/Git, signed upload) go through
skill_ingest.install_skills. The tool count in the server log, app
comment, docs, and the Codex plugin's REAL_MCP_TOOLS moves to 16.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(mcp): search shared skills in find(context_type="skill")

Without a target_uri, find resolved the generic default targets, which
stop at the caller's user root, so a skill search never reached the
account-shared viking://agent/skills. REST /skills/find and the context
search already cover both roots. When context_type resolves to skill only
and no target_uri is given, the MCP tool now targets
default_target_directories(ctx, context_type=SKILL): the user's own skills
plus viking://agent/skills. REST find semantics are unchanged.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(mcp): print directory abstracts in tree(include_abstract=true)

The tree tool skipped to the next entry right after printing a directory,
and only printed abstracts for files, but the storage layer only fills
abstracts for directories (files always come back empty). The flag
therefore never printed anything. Print the abstract after either kind
of entry and ask for up to 1024 characters, enough for a full skill
description, so tree(uri="viking://~/skills", level_limit=1,
include_abstract=true) lists every skill with its description.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(skills): honor node_limit in GET /api/v1/skills

list_skills declared node_limit but always listed each skill root with a
hardcoded 1000. Pass it through per root; 0 keeps the default so the CLI's
accepted range (-n 0) still lists everything.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(mcp): point skill hits in find/search at their SKILL.md

A skill is indexed through its directory's .abstract.md, so find and
list-mode search printed hits like viking://agent/skills/x/.abstract.md.
Following the "use the read tool to expand a URI" advice returned only
the frontmatter, and read_content inlined the same stub. Skill hits now
show <dir>/SKILL.md, and read_content reads that file.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(mcp): validate add_skill targets and sources before minting an upload

Review findings on the add_skill tool:

- target_uri passed the content-kind check for any path under a skills
  root (viking://~/skills/pdf) and, for ROOT, for another user's root,
  but the installer only accepts the caller's own skills root or
  viking://agent/skills. On the local-path branch the tool minted a
  one-time upload token anyway, and the upload failed with 400 after the
  token was spent. The target is now resolved with the installer's own
  rule first; shared subpaths map to viking://agent/skills, the rest fail
  at once, and the error names both allowed roots.
- Non-Git remote sources such as tos:// were treated as remote, then
  refused as "direct host filesystem paths". add_skill now decides Git
  with the same prefixes resolve_skill_source uses (shared as
  GIT_SKILL_SOURCE_PREFIXES) and reports other schemes as unsupported.
- With list_only, the upload instructions still said the skill would be
  installed and that no further call was needed; they now say the upload
  only lists the source's skills.
- The zip example packaged hidden files, so .git and .env files went
  into the stored skill. It now excludes VCS data, .env files,
  node_modules and .DS_Store, starting from a fresh archive.
- tree(include_abstract=true) printed the "abstract is not ready"
  placeholder for directories that never get an abstract, such as a
  skill's scripts/. Those placeholders are skipped.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* docs(mcp): say that write only refuses the user's own skills subtree

The capability reference claimed MCP write refuses every skill URI. It
refuses the user's own skills/ subtree, but under viking://agent/skills
it writes a plain file that skips skill installation. State that, and
point shared skills at add_skill as well.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* style(skills): format skill_processor.py

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* feat(plugins): inject an <available-skills> catalog at session start

Agents on every harness learned about memories at session start but had
no idea which skills OpenViking held, so a stored skill was only found if
a later recall happened to surface it.

buildProfileBlock() now takes the caller's resolved config as a fourth
argument and, when skillCatalog is on (default), adds <available-skills>
after <available-memories>: one GET /api/v1/skills lists the user's own
skills first, then account-shared ones, dropping a shared skill the user
shadows by name. Descriptions are cut to about 40 tokens and envelope
tags in them are escaped, since the shared root is written by anyone on
the account. The block has its own budget (skillCatalogTokenBudget,
default 1200) and degrades from descriptions to names to a one-line
count; with no skills, or a server without the endpoint, it is omitted.

The shared formatListing now gives entries back so its "+N more" tail
fits: a greedily filled listing never left room for it, so a cut listing
ended silently. This also applies to <available-memories>.

All six callers (claude-code, codex, opencode, dsh, pi, and the thin hook
runtime for cursor/trae/zcode) pass their config through.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* feat(plugins): recall account-shared skills and flag skill entries

The server context face already mixes both skill roots into per-prompt
recall, but the plugins' last-resort ranked find only searched
viking://~/memories and viking://~/skills, so on servers without the
context face a shared skill in viking://agent/skills could never be
recalled. Add it as a third source, and name each skill hit by its
directory rather than the .abstract.md it was indexed through, matching
the context face and the session-start catalog.

When an injected recall block carries a skill (a type="skills" entry, or
a [skill] line from the fallback), its header gains one line telling the
agent to read the skill's SKILL.md before following it. Turns without a
skill are unchanged.

The openclaw plugin's vendored recall-core copy is regenerated with it.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* feat(plugins): point skill writes at add_skill in the URI guard

A local Write or Edit aimed at viking://.../skills/... was denied with a
hint to use MCP write or edit instead, but the server refuses both under
the skills subtree, so the hint led straight into a second error. Hints
may now depend on the URI: for a skill URI (viking://~/skills,
viking://user/<id>/skills, viking://agent/skills) the default table and
dsh's bridged table name add_skill with an add_skill(data="<the full
SKILL.md text>") example. Other URIs keep their hints.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* feat(plugins): bundle an openviking-skills skill

Nothing told an agent how to work with skills stored in OpenViking: how
to load one from the catalog, run its helper files, create one through
add_skill, install from Git or a local folder, share it with the account,
or move the user's existing local skills over.

examples/skills/openviking-skills covers all of that, including a
user-triggered, one-time migration of ~/.claude/skills, ~/.agents/skills
and ~/.cursor/skills that keeps environment-bound skills local (shipped
by a plugin or marketplace, symlinked in by a CLI installer such as
lark-cli, or needing a local binary) and uploads only what the user
approves skill by skill. It passes strict server validation.

sync.mjs ships it wherever add_skill is a real tool and a bundled skill
loads: the codex, claude-code, cursor and dsh plugins. openviking-memory
now points to it for skill work. A new sync test keeps synced skills
flat, since copySkill copies a flat file list and a subdirectory would
crash it; the marketplace tests pin the packaged copies, and dsh's
provider test expects both bundled skills.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* chore(plugins): bump versions for the skill integration

claude-code 0.6.0, codex 0.10.0, agent-hook (cursor/trae/zcode) 0.4.0 and
dsh 0.5.0 gain the skill catalog and, except trae/zcode, the bundled
openviking-skills skill; opencode 0.3.3 and pi 0.3.3 gain the catalog.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* feat(codex): recall shared skills and flag skill entries in Codex too

Codex builds its recall block itself instead of through recall-core's
wrappers, so the previous commit's changes never reached it: its raw
search still skipped viking://agent/skills, its digest carried no skill
hint, and its post-processing kept only level-2 leaves, which silently
dropped every skill hit (skills are found through their directory's
level-0 abstract), including the viking://~/skills search it already ran.

recall-core now exports skillEntryHint() and skillHitUri(). The hint is
added when a block carries a type="skills" entry or cites any skill URI,
which also covers digests that only keep URIs. Codex searches the shared
skill root as a third bucket, labels skill hits "skills" under their
directory URI, lets them through post-processing, and adds the same hint
line to its envelope.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(cursor): install openviking-skills next to openviking-memory

sync.mjs puts openviking-skills into hosts/cursor/skills, but the
installer copied only openviking-memory into ~/.cursor/skills, so Cursor
never saw the new skill. Install, uninstall, the post-install check and
the doctor's file list now cover both skills. The install test also moves
to the agent-hook plugin's new 0.4.0 version string.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(dsh): keep PLUGIN_VERSION in step with package.json

The version bump moved package.json to 0.5.0 but left the PLUGIN_VERSION
constant at 0.4.3, which bundle.test.mjs and npm run check:version
compare against the manifest.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* docs(skills): match ov-skills commands to the real ov CLI flags

The ov-skills skill documented flags the CLI never had (--json, and a
--limit that is only a hidden alias), a raw-content "ov skills add -"
form that sends a literal "-", and ov resources subcommands that do not
exist. Every command line now follows the clap definitions: -o json for
JSON, -n/--node-limit, -p/--uri on read commands and
-p/--parent-auto-create on add, -s/--skill as a comma list, show
--format, and validate's --strict-only body-length warning. ov add-skill
is documented as the same command as ov skills add.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* docs(plugins): document the skill catalog, openviking-skills, and skill-aware recall

The integration pages (Claude Code, Codex, Cursor, TRAE, opencode, pi,
dsh), the capability reference, the plugin development guide, and the
plugin READMEs now describe the <available-skills> session-start block,
its skillCatalog / skillCatalogTokenBudget knobs, the bundled
openviking-skills skill where it ships, recall reaching
viking://agent/skills with the skill-entry hint, and the URI guard
sending skill writes to add_skill. en and zh pages carry the same facts.

Stale statements fixed on the lines touched: thin hook hosts use the
same 10000-token profile budget as the rest, the claude-code and codex
plugins ship four skills, dsh mounts the server MCP surface, and the
opencode install guide lists openviking_add_skill once.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(plugins): let the user's own skills use the catalog budget the shared group leaves

Review findings on <available-skills>:

- Each group got at most its even share of the listing budget, with
  unused tokens passing only forward, so the user's own skills (always
  first) never got more than half. Twenty own skills and one shared skill
  fell back to names only while most of the 1200 tokens went unused. A
  group now takes its even share or everything the later groups leave
  when listed in full, whichever is larger.
- When a group's share could not hold its header plus the "+N more"
  tail, formatListing gave back every entry and printed a bare header,
  which reads as an empty directory. It now prints the one-line
  "N entries, budget too tight" stub instead (memory listings too).
- A budget too small for even the one-line count now injects nothing
  rather than overrunning it.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(plugins): rank skill hits like memory leaves in the recall fallbacks

Naming a skill hit by its directory instead of its .abstract.md cost it
the 0.12 leaf boost, since the boost keyed on a ".md" URI, so a skill
that main would recall lost to ten slightly weaker memory leaves. In
Codex, skills were also never picked while enough memory leaves passed
the threshold (leaves are picked first), and a hit the server labeled
with another category lost its "skills" label. Skill hits now count as
leaves in ranking and picking in both recall-core and Codex, and Codex
always labels them "skills".

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(plugins): send shared-skill edits back to the shared root in the URI guard

The guard's add_skill example carried no target_uri, and add_skill
without one installs into the caller's own root. Fixing a shared skill
that way left the team copy unchanged and created a private copy that
the catalog then shows instead. For viking://agent/skills URIs the
example now passes target_uri="viking://agent/skills", and a helper file
(anything below a skill's SKILL.md) points to a folder upload through
add_skill(path=...) rather than SKILL.md text. addSkillExample() builds
the example for both the default table and dsh's.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* docs(skills): tighten openviking-skills where review found unsafe steps

- The catalog is a snapshot that drops descriptions or entries with many
  skills, so a name missing from it does not prove the name is free.
  Check <root>/<name>/SKILL.md, and confirm with the user before
  replacing an existing skill, since add_skill replaces silently.
- Updating a shared skill must pass target_uri="viking://agent/skills"
  after the user confirms; otherwise add_skill creates a private copy
  that shadows it.
- Every file in an uploaded folder is stored with the skill: zip without
  .git, .env files, node_modules and .DS_Store, and delete the archive
  afterwards. The migration now inspects the whole folder, hidden files
  included, for secrets.
- Migration flags frontmatter keys OpenViking drops (for example
  disable-model-invocation or context), and fixes a missing name or
  description in a temporary copy, never in the user's file.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* feat(plugins): keep the session-start block under the host's inline limit

Claude Code saves hook context over 10,000 characters to a file and
shows the model a 2 KB preview; Codex and trae-cli spill past about
10,000 bytes; ZCode drops stdout over 32 KB. The session-start block
(profile at a 10,000-token budget, memory index, skill catalog, and the
archive on resume) routinely ran 25-40 KB, so on these hosts the model
saw only the start of the profile, and the catalog appended at the end
never reached it.

sessionStartMaxBytes caps the whole block in UTF-8 bytes: 9500 for
claude-code and codex, 20000 for zcode, no cap elsewhere. Under the cap
buildProfileBlock shrinks its token budgets (about 4 bytes per estimated
token); if the block still does not fit it drops the memory index, then
the catalog. On resume/compact the archive takes up to half, truncated
on a line with a pointer to viking://~/sessions/<id>/history/.

On resume, claude-code and codex skip the profile block when it matches
the one this session already received, since the restored history holds
it; a changed block is injected again.

Claude-Session: https://claude.ai/code/session_01ECVkFufAU2LKe4g83cxt28

* fix(mcp): return one hit per skill package in find

#5045 made a skill index as a whole package, so an item-level find now
returns one hit per file inside it. Route skill-only find through
SearchService.find_skills, which keeps the best hit per package, and
resolve every skill hit to its package's SKILL.md instead of only
rewriting the two index sidecars.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* docs(mcp): point skill changes at add_skill in the tool descriptions

The server keeps accepting write/edit under viking://agent/skills, and
forget still removes a skill directory, so the constraint lives in the
tool descriptions: add_skill is the one entry point for creating and
updating a skill, and removal goes through ov skills remove or Studio.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* fix(mcp): describe a skill hit by its own abstract

A package hit can be any file inside the skill, whose abstract describes
that file and not the skill, so find would list a skill under a helper
script's summary. Read the package's abstract for those hits, the way
GET /skills/find already does. Keep a filter-only skill query on the
generic find, which find_skills does not serve.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* docs(mcp): document package-level skill retrieval in find

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* feat(agent-plugins): ship the openviking-skills skill

Agent Plugins has no hooks, so no session-start catalog: without this skill
the model never learns that the account's skills exist. The skill's own text
now reaches for find(context_type="skill") first and treats
<available-skills> as something only some harnesses inject, so one copy
reads correctly in both kinds of harness.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* fix(plugins): name the skill package a recall hit came from

Since a skill is indexed as a whole package, a hit can be any file inside
it, not just the two index sidecars the old rewrite stripped. Derive the
package root the way the server's skill_root_uri does, drop the internal
update backups, and keep one entry per package at its best score.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* fix(plugins): let the skill catalog use both roots' full listings

The server already caps each skill root at node_limit, so a second cap over
the merged list only bites once the private root alone fills it — and then
it drops the shared root whole while reporting nothing dropped. The token
budget is what should decide, and it already does.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* docs(plugins): say node_limit caps each skill root, not the merged list

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* fix(mcp): do not paste an unready abstract over a skill hit's own summary

fs.abstract returns a placeholder string rather than raising when a package
has no usable .abstract.md, so the substitution replaced a useful file
summary with a diagnostic line. Reject the same placeholders tree already
rejects, and bound the per-package reads the way read_content is bounded.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* docs(mcp): correct how search reports skill hits, and refresh the tool table

find and search both render one line per skill package, so the earlier
wording — that search returns several hits per package — contradicted the
code. Say what actually differs: search still spends a limit slot per
matching file and keeps that file's summary. Also point forget and
add_resource at add_skill where an agent would look for them, name the REST
delete alongside the CLI, and bring the capability table's line citations
back in step with mcp_endpoint.py.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* docs(agent-plugins): list add_skill among the tools the package exposes

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* refactor(mcp): drop guards and prose no caller can reach

install_skills only ever returns a dict, add_skill always fills root_uri,
and a source with no SKILL.md raises before it gets here, so the
isinstance, empty-list and missing-uri branches were unreachable.
fs.abstract only returns the directory placeholder. One skill package
resolves to one rendered item, so the pending map holds one each. In the
docstrings, drop what Args already says and the one removal path an MCP
caller cannot take.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* docs(plugins): drop a comment about a branch formatListing cannot take

A listing left with only its header returns the stub above, so it never
reaches the silent close the comment described. Also name
sessionStartMaxBytes in buildProfileBlock's options type, where the .d.mts
already has it.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* revert(plugins): drop the skill changes in the recall fallback

Reverts the recall half of the skill integration: 992215bfb, 553416200,
5e838a0a0 and 62a456c08. The seven files they touched go back to their
state at aa77061c1, the merge base with main.

Those four commits taught the local recall assembly about skills: a third
source for viking://agent/skills, skillHitUri naming a hit's package,
dedupeSkillHits keeping one entry per package, rankItem scoring a skill
like a memory leaf, and a header line telling the model to read SKILL.md.
Four of the five only ever ran in the raw-find fallback. recallForPeer
calls buildServerAssembledBlock first and returns as soon as it answers,
so searchAllSources, rankItem and the fallback block builder are reached
only on a deployment whose server has no context face. The fifth, the
header line in wrapContext, did run on the main path, but the server
already reports each entry's type and the URI it wants read, so the line
restates what the block carries.

Skills still reach the model on the path that runs: the context face
searches both skill roots, returns them as entries with type="skills",
and the session-start <available-skills> catalog lists every skill with
its description. Both stay.

The documentation that described the fallback behaviour goes with it. The
statements that survive are the ones the context face makes true on its
own, such as recall covering the skills shared with the account.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* chore(plugins): bump opencode and pi past main's releases

Both were 0.3.3 on this branch and main has since shipped 0.3.3 of its
own, so the version a host installs by no longer moves for the skill
catalog they now carry through the shared library.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW

* docs(mcp): say what list-mode search actually reports for a skill hit

The search row claimed a skill package's summary is the matching file's.
It is not: _format_search_result rewrites every skill hit onto the
package's SKILL.md, keeps the best-scored one per package, and
_describe_skills_by_package replaces the summary with the package's own
abstract. What is true is that limit applies during retrieval, before
that merge, so a package matching several files still spends several
slots and fewer than limit results come back.

Claude-Session: https://claude.ai/code/session_01CrZadR75kCueyZyoiGUFBW
2026-09-22 15:09:44 +08:00
t0sakiandTRAE CLI aa3ee2f431 refactor(plugins): resolve every harness's knobs from one declaration, and generate the packaged copies at pack time (#4773)
* fix(codex): stop dropping turns on an outage, and honour bypass patterns

Two capabilities every other JS harness has were wired in the shared library
but never reached codex.

Offline queue. `catchUpTurns` called `sendSessionMessages` without
`enqueueOnRetryable`, so a 5xx or a connection failure returned a count and the
turns were gone: codex hooks are short-lived subprocesses with no retry of
their own, and the `capturedTurnCount` cursor only compensates while the
process survives long enough for another Stop. The queue module was vendored
into `scripts/shared/` all along with no caller, and nothing ever replayed it.
The send now enqueues, and SessionStart drains the queue after its health check
— the one codex hook that runs against a known-healthy server.

Bypass. `isBypassed` had no codex caller, so `bypassSessionPatterns` did
nothing there: a directory a user had excluded still had its turns captured and
still got memories injected, and the `bypass.session_patterns` key a repository
commits in `.openviking/config.json` was silently inert for every codex user in
that repository. All five hooks now consult the shared matcher, and the loader
reads the same keys and the same `OPENVIKING_BYPASS_SESSION[_PATTERNS]` env
vars as claude-code.

SessionStart is deliberately only half-suppressed under bypass: it skips peer
registration and injection, but still replays the queue and still sweeps. Both
finish work recorded by sessions that were not bypassed, and this hook is the
only place codex runs either, so bailing out would strand that data for as long
as the user kept working in a bypassed repository.

Claude-Session: https://claude.ai/code/session_01GJouP1ghK4gwXw7sDjuZX6

* fix(pi): use the shared bypass matcher, and keep the pre-git peer reachable

`config.ts` kept only `.peerId` out of `resolveEffectivePeerId()` and dropped
`legacyPeerId`, so under `recallPeerScope: "actor"` — where the effective peer
is the only one asked — every memory written before the git-derived peer
replaced the path-derived one became unreachable. The shared recall builder has
read `options.legacyPeerId` for the dual read all along; pi just never passed
it.

Bypass was a hand-written `matchBypass()` under this extension's own
`bypassPatterns` key: it understood a leading or trailing `*` and nothing else,
so `**/scratch/**` did not work here while it did everywhere else. It now calls
the shared `isBypassed`, reads `bypassSessionPatterns` like every other
harness, and honours `OPENVIKING_BYPASS_SESSION[_PATTERNS]`. `bypassPatterns`
still reads.

That does change one behaviour for existing setups: a bare path used to match
its subdirectories as a prefix, and a glob does not. The README says to write
`"/tmp/scratch**"` where `"/tmp/scratch"` used to be enough.

Claude-Session: https://claude.ai/code/session_01GJouP1ghK4gwXw7sDjuZX6

* chore(plugins): remove the deprecated standalone TRAE CLI integration

`examples/trae-cli-memory-hooks` was added on 2026-08-17 (#4026) for TraeCode
CLI 1.0 and deprecated the next day by #4079, which routed `--harness trae-cli`
to the codex plugin alias instead. Since then `install_trae_cli` has not been in
the installer's dispatch block — `normalize_trae_cli_harness` rewrites the
harness to `codex` before it could run — so the directory, its 588 lines and its
CI test have been dead weight that the marketplace archive still shipped to
every user.

The uninstall path stays: `remove_legacy_trae_cli_integration` and
`agent_remove_trae_cli_configs` delete what an old install left on disk and
never read this directory, which is what the retention test in
`release-marketplace.test.mjs` claimed to protect. `install-agent-hooks.test.mjs`
covers that path and still passes.

Also drops the write half that only `install_trae_cli` called.

Claude-Session: https://claude.ai/code/session_01GJouP1ghK4gwXw7sDjuZX6

* fix(plugins): make a changed plugin prove it shipped

Claude Code resolves an installed plugin by the `version` string in its
manifest, not by the commit the marketplace ref points at. That string has read
0.4.5 since #4389 on 2026-08-27, with seven commits to the plugin and the shared
library behind it — so nobody running the marketplace build has received any of
them. Codex's manifest has been frozen at 0.8.1 for ten. Every vendored-copy
diff those commits paid for went to a payload no user received.

Both versions move, and a new PR check fails when files under a
marketplace-distributed plugin change without its version string moving with
them. The generated `shared/` copies live inside each plugin directory, so a
shared-library change reaches the check through them.

Also removes `claude-code-memory-plugin/package-lock.json`: 1174 lines locking
95 packages for a private manifest that declares no dependencies, still pinned
at 0.4.4.

Claude-Session: https://claude.ai/code/session_01GJouP1ghK4gwXw7sDjuZX6

* docs(agent-integrations): record what codex, pi and trae-cli now do

The capability matrices said codex has no on-disk queue and counted four hooks
where hooks.json declares five; both are now wrong. pi's bypass is no longer a
prefix match under its own key. And the removed TRAE CLI 1.0 integration is
described by what remains of it: the installer's uninstall path.

Claude-Session: https://claude.ai/code/session_01GJouP1ghK4gwXw7sDjuZX6

* refactor(plugins): derive each plugin's vendored set from what it imports

sync.mjs kept hand-written capability groups, and the groups got spread into
each other: the file's own header states the rule it could not enforce ("what a
plugin ships must equal what it imports"). Both directions had failed in
practice — a module a plugin imported but no list named became
ERR_MODULE_NOT_FOUND on the first hook of a fresh install, and modules nobody
imported sat vendored in plugin trees, landing in every review diff forever.

The lists are gone. Each target now says only where its code lives and where
its copies go; the file set is the transitive closure of the imports that code
actually resolves into the vendored directory — including re-export shims like
claude-code's scripts/lib/, which are the sole importer of several modules. The
sync reports a copy nothing imports and a lib module no plugin reaches, and
sync.test.mjs asserts what is on disk equals that closure in both directions.
The run is byte-identical for everything still shipped, and it named the two
copies nobody had noticed: claude-code's capture-utils (515 lines, its
auto-capture hand-rolls the same filter) and codex's uri-guard (78 lines, codex
has no PreToolUse hook).

ZCode stops vendoring altogether. Its only install path is
`assemble_agent_integration`, the same one cursor and trae use, which already
placed the canonical runtime beside it — it just imported its own committed
copy instead. Pointing its nine imports at `../../memory-plugin-shared/lib/`,
the path that resolves both in the repository and under
`$OV_HOME/agent-integrations/`, drops 4294 generated lines and leaves cursor,
trae and zcode reading one runtime. The installer's assembly list grows by the
three modules zcode needs, derived by the same closure code so it cannot drift
either, and the end-to-end archive install test covers it.

Claude-Session: https://claude.ai/code/session_01GJouP1ghK4gwXw7sDjuZX6

* fix(plugins): run cursor and trae capture through the shared filter

zcode got this in #4594; cursor and trae never did. Both sent whatever their
transcript parser produced straight to the extractor: `/compact`, a bare "ok",
a punctuation-only turn and a `[openviking-memory]` status line all became
memories, and a turn past `captureMaxLength` went out uncapped. Every other
harness has run `shouldCaptureText` on the way in for months.

The decision now lives beside the runtime the thin harnesses compose, as
`filterCaptureTurns`, so a harness gets it by calling rather than by carrying
its own copy — and the installer already assembles `capture-utils.mjs` for
these three since they share one runtime.

Both hooks keep hashing the raw turn rather than the filtered text, so raising
captureMaxLength never resends a turn the server already holds truncated, and a
dropped turn is recorded as handled so a re-read of the same transcript does not
re-evaluate it every run.

The TRAE hook test used `last_assistant_message: "done"` — an acknowledgement
the filter drops, and correctly so. Its fixture now carries a turn worth
remembering.

Claude-Session: https://claude.ai/code/session_01GJouP1ghK4gwXw7sDjuZX6

* fix(plugins): bump every plugin the shared change reached, and pair the manifests

The version check added earlier in this branch caught what a human review would
not have: `filterCaptureTurns` landed in the shared library, its generated copy
changed inside opencode and dsh, and neither manifest moved — so both hosts
would have kept installing the build without it. cursor, trae and pi move for
the same reason.

It also caught a mismatch this branch introduced: cursor carries the version in
both `.cursor-plugin/plugin.json` and `openviking.integration.json`, the host
installs by the first and the installer decides "nothing changed" by the second,
and bumping only one makes a plugin report upgraded while behaving as it did.
zcode had been sitting at 0.1.1 against 0.1.2 on main for the same reason. The
check now compares the pair, so neither can drift again.

Claude-Session: https://claude.ai/code/session_01GJouP1ghK4gwXw7sDjuZX6

* refactor(plugins): resolve every harness's knob from one declaration

A setting only stayed uniform across the harnesses when a shared module read
it on the hot path. `recallLimit`, `scoreThreshold` and `captureMaxLength`
agree everywhere because `recall-core` and `capture-utils` read them; the
switches, which nothing shared read, drifted into four spellings — a boolean
`autoRecall`, an `autoRecall` object in opencode, `syncTurns` in dsh and pi,
and no recall switch at all in the thin hook runtime. Three more inventories
had to be kept in step by hand: the doctor's known-knob set, the workspace
file's dotted-key map, and the implicit list inside `loadAgentHookConfig`.

`lib/config-schema.mjs` is now the one declaration, and the rest are
projections of it. `resolveSettings()` resolves any harness through the layers
the shared README has documented since the `plugin` section landed — env, the
workspace file and registry, `ovcli.conf` `plugin.<harness>`, `ovcli.conf`
`plugin`, ov.conf's harness section, defaults — which only Claude Code and
Codex implemented. cursor, trae, trae-cn and zcode read the environment and
nothing else before this, so a `plugin` entry named after them was inert and
`ov config switch` moved their credentials while leaving their behaviour
behind.

The two per-plugin config files are gone with it. pi's `config.json` shipped
with the extension holding exactly the code defaults, so nothing could tell an
operator's choice from the factory setting; opencode searched four paths for
`openviking-config.json`. Both harnesses read `ovcli.conf` now, as the others
do. The `.d.mts` stubs move to lib/ beside the modules they describe — pi and
openclaw each carried a different, both incomplete, declaration of
`recall-core` — and sync ships them to the targets written in TypeScript.

The switches move into the capability modules that own them:
`isRecallEnabled` and `isCaptureEnabled` accept every spelling and treat any
of them saying off as off, which is the one mechanism in this repository that
has actually held a name still.

* build(plugins): generate the packaged plugins' shared copies at pack time

`sync.mjs` had zero references under `.github/`: the generator ran on whoever
last remembered, and CI only checked the result byte for byte. So every change
to a 6,000-line library arrived as a 20,000-line diff, and the copies were
committed for plugins that never needed them committed.

The split is delivery, not taste. A host that installs Claude Code, Codex or
the Agent Plugins package points at a directory in this repository and can only
load what git holds, so those copies stay committed — and a push to main
regenerates them, which is what stops one going stale behind a merged pull
request. opencode, dsh and openclaw publish as npm packages and pi is tarred by
the installer, so those four build their copies on the way out: a `prepack`
script for the packages, the marketplace staging script for the archive the
installer reads, and `sync.mjs` in a source checkout for a direct `install.sh`
run. `npm pack` was verified to regenerate a deleted directory and ship the
same file set.

Two things had been holding those directories in place. Their release triggers
matched only `examples/<plugin>/**`, so a shared fix reached npm through the
vendored copy changing — the triggers name the library now. The version gate
found a changed plugin the same way, so it treats a change under
`memory-plugin-shared/lib` as a change to every plugin, and gains cursor and
trae, which are distributed the same way and were never gated. pi's own
`recall-ledger` moves to `lib/`, where the extension's other own modules live,
so the generated directory holds nothing but generated files.

* test(plugins): pin the layer order for the harnesses that had no config test

cursor, trae, trae-cn and zcode compose one shared loader and none of them owns
a config test, so nothing would have caught the `plugin` section going inert
again. This asserts the whole stack through that loader — the shared block, a
per-harness override under either spelling, the workspace file over ovcli.conf,
and the environment over all of it — and pins that a default never reports as
configured, since several fields reach the server only when the user asked for
them.

* docs(plugins): point opencode and pi at ovcli.conf, and say when copies are made

Both plugins documented a configuration file that no longer exists, down to the
four paths opencode searched for it and the nested blocks neither loader reads
any more. Their knobs live in `ovcli.conf`'s `plugin` section now, so the
examples are ovcli.conf examples, the flattened names are the ones the schema
declares, and each document says the resolution order and links to the one file
that declares every knob.

Two claims that had gone the other way: opencode and pi read the workspace
`.openviking/config.json` now, so a `peer.id` written there does apply, and the
shared README says which plugins keep their generated copies in git and which
build them at pack time, because a fresh checkout has to run the generator once
before their tests will resolve an import.

* fix(opencode): name the file the 401 hint should send the user to

* fix(plugins): let the canonical knob name win over its own alias

A file that carries both spellings — what a rename leaves behind — resolved to
whichever the layer loop reached last, which was the older name.

* fix(plugins): write each generated copy through a rename

The marketplace staging script runs the generator now, so it can rewrite a
vendored module while a test is byte-comparing it. A partial file read that way
fails as a drifted copy, which is a confusing way to say nothing drifted.

* fix(dsh,pi): honour the recall switch these two never read

Both gained a capture switch when `syncTurns` became an alias, but their recall
paths still retrieved on every prompt: nothing in either called the switch, so
`autoRecall: false` — settable from ovcli.conf, a workspace file or the
environment — was accepted everywhere and obeyed nowhere.

* docs(agent-integrations): say which harnesses read which configuration layer

The capability reference had cursor, trae, trae-cn and zcode down as
environment-variable-only, opencode reading a four-level `openviking-config.json`
search and pi its own `config.json`, and the workspace files and `plugin`
section scoped to Claude Code and Codex. All nine resolve through the same
loader now, so the tables, the quick-scope list and the profile cards say so,
and the module table gains the schema the layers project from.

The distribution paragraph was counting modules against a `HOOK_SHARED_FILES`
list that no longer exists and calling zcode the only harness that vendors the
hook runtime, which stopped being true when its copies were dropped. The counts
are the real ones, and the paragraph now says when each target's copies are
generated, since that is what decides whether git holds them.

* build(plugins): keep node_modules out of the marketplace archive

Three of the staged plugin trees carry installed development dependencies,
so the copy shipped 134MB nobody unpacks and spent most of the release test's
54 seconds on them. Copying through tar leaves them behind while still
carrying the generated shared copies, which a tracked-file listing would miss.

* build(plugins): read the installer's shared-module list from the sync

The installer carried a second, hand-written copy of the closure sync.mjs
already computes, and only a regex over the bash text held the two equal. The
sync writes that closure to lib/MANIFEST now and the installer copies what it
names, so the set has one source and the test compares two lists instead of
scraping a script. The installer cannot just run the sync: a marketplace
archive is flat, and the generator resolves an empty list there.

* build(plugins): derive what the marketplace archive must contain

The stage script hand-listed 56 required paths beside three generators that
already knew them, and it had drifted from all three: eleven of the twenty-two
shared modules the installer assembles, four of the thirty-eight copies the sync
generates, and neither marketplace manifest. The list is now computed from each
plugin's own manifests, the imports those entrypoints reach, and the sync's
target list, leaving the script with only the directories it ships.

* fix(plugins): resolve credentials from the calling harness's ov.conf section

The shared resolver read `ov.conf`'s `codex` block for every caller, so the api
key, account, user and peer of opencode, pi, dsh, cursor, trae, trae-cn and
zcode all came out of a block named after another harness, while the block
named after them did nothing. A user who wrote `opencode.apiKey` saw no effect
and a stale `codex` block quietly authenticated everything else.

`resolveOpenVikingCredentials` now takes the harness and reads the section named
after it, under either spelling. The layer keeps its place — after ovcli.conf,
ahead of `server.root_api_key` — and the ovcli-pinned mode still skips ov.conf
entirely, so codex, the caller that keeps the default, is unchanged.

For users this means the `<harness>` block finally configures that harness, and
anyone who was relying on the `codex` block to supply credentials to a different
harness has to copy them into their own block. dsh merges its section in
explicitly because the cordis patch occupies its legacy layer, and codex now
honours an api key set through `ovcli.conf`'s `plugin` section, which the
credential chain has never seen. cursor, trae, trae-cn and zcode honour it too:
the hook runtime spread the resolved credentials over the settings, so an empty
key from the chain used to erase the one `plugin.<harness>.apiKey` had supplied.
The doctor warns about all of these blocks now, since the server refuses to
start on any of them.

* fix(plugins): let the credential chain supply the peer on claude-code and dsh

Claude Code had no peer logic at all: it took only the user agent from the
shared credential module, so `actor_peer_id` in ovcli.conf and `peerId` in
ov.conf's `claude_code` block were read by every other harness and ignored
here. `ov config switch` moved the peer and Claude Code kept writing under the
one derived from the workspace. It now takes the peer from the same chain, and
a `plugin` entry or a workspace file still names the more specific answer for
that directory, which is the rule the thin harnesses already follow.

dsh had the opposite fault: it assigned the credential peer over whatever the
layers resolved, and that chain is the empty string whenever no actor peer is
named, so `plugin.dsh.peerId` was resolved and then thrown away. The credential
peer is a fallback now instead of an overwrite.

* fix(plugins): resolve the peer through one chain on every harness

Six loaders each wrote their own peer order, and two of them disagreed with the
rest: codex let a `peerId` written for the harness win, while opencode and pi
let `ovcli.conf`'s account-wide `actor_peer_id` win. One ovcli.conf therefore
produced two different peers depending on which host read it, and `ov config
switch` moved the peer for some harnesses and not others.

`resolvePluginPeerId()` in the shared library is now that order, once: a peer
the host named, then `OPENVIKING_PEER_ID` (suppressed when the credentials are
pinned to ovcli.conf, where the environment is meant not to apply), then the
workspace file, the registry and `ovcli.conf`'s `plugin` section, then
`ovcli.conf`'s `actor_peer_id`, and last `ov.conf`'s harness block, which keeps
the place it has always had. Nothing named means nothing explicit, which is
what makes `peer.source` derive one.

For opencode and pi users this is a precedence change: a `peerId` under
`plugin` or `plugin.<harness>` in ovcli.conf, or a `peer.id` in a workspace
file, now outranks that file's `actor_peer_id` instead of being outranked by
it. Anyone who wrote both and wanted the actor peer has to drop the plugin one.
It is one for claude-code too: `ov.conf`'s `claude_code.peerId` was the only
peer that harness read, and it now sits under `actor_peer_id` like every other
harness's block. Otherwise nothing moves: `ov.conf`'s harness block is ranked
here rather than left to the credential chain, which drops it whenever the
credentials are pinned to ovcli.conf, so that block names the peer in the same
place whether they are pinned or not.

* fix(plugins): shape the cursor and trae MCP proxies like every other one

Both proxies hand-built the config object the shared `buildMcpProxyConfig`
produces everywhere else, and had done so since before that helper existed, so
they never picked up the fixes it carries. The visible one is the actor peer:
they put whatever peer the config layers resolved straight onto
`X-OpenViking-Actor-Peer`, while the default `recallPeerScope` of `all` means
broad recall and every other proxy deliberately sends no such header. Cursor
and TRAE therefore recalled at a narrower scope than the rest on identical
configuration. Two smaller ones come with it: the watch list now includes the
default credential paths, so a proxy started before the first `ov login` picks
up the file it creates instead of never reloading, and the request timeout is
clamped rather than taken raw.

The guard is the file list in `mcp-proxy-config.test.mjs`, which named five
proxies by hand and so never covered the two that were wrong. It is a directory
scan now, pinned to a count so a rename cannot quietly empty it, and it asserts
that every proxy in the tree reaches the shared builder.

* fix(plugins): put one set of headers on the wire from every harness

Seven request builders each wrote their own header block and three of them
disagreed. Codex sent the api key twice, as `Authorization: Bearer` and again
as `X-API-Key`, which the open-source server prefers when both arrive — so a
gateway rewriting one of them changed which credential authenticated. Claude
Code read `cfg.accountId` where the other six read `cfg.account`. And only
Codex asked whether the server was in trusted mode before naming the operator:
everywhere else `X-OpenViking-Account` / `X-OpenViking-User` went out whenever
they resolved, including to `api_key` servers that read both out of the key,
ignore the headers, and leave the identity visible to every proxy on the path.
Both doctors have been warning users about that as if it were already fixed.

On the wire, from this commit:

- `X-API-Key` stops. Codex's `ov-session.mjs`, `auto-recall.mjs` and
  `session-start-commit.mjs` were the only senders; `Authorization: Bearer` is
  unchanged and remains the only credential header. A gateway that needs
  `X-API-Key` has to add it itself. openclaw keeps sending it and is untouched.
- `X-OpenViking-Account` and `X-OpenViking-User` are sent only under
  `sendIdentityHeaders`, which is `authMode === "trusted"`. Codex already
  behaved this way; claude-code, opencode, dsh, pi and the cursor / trae /
  trae-cn / zcode hook runtime now do too, and so do the two senders outside
  the hook stacks: every harness's stdio MCP proxy, which takes the switch
  through `buildMcpProxyConfig` the way it already takes the actor peer, and
  Claude Code's status-line server probe. Auth mode resolves through
  `resolveAuthMode()` in the shared credentials module: `plugin.<harness>`
  `authMode`, then `plugin.authMode`, then `ov.conf` `<harness>.authMode` on
  the three harnesses whose loader reads that section as a settings layer
  (claude-code, codex and dsh), then `ov.conf` `server.auth_mode`, then trusted
  whenever an account or a user resolved at all — the chain Codex already had,
  now shared. An operator whose server is in `api_key` mode and who has an
  account in ovcli.conf stops sending it; nothing else changes.
- `X-OpenViking-Actor-Peer`, `User-Agent` and `Content-Type` are unchanged.

Claude Code's config now also exposes the resolved identity as `account` /
`user` beside the existing `accountId` / `userId`, and its doctor and its
status-line probe use the headers the plugin would really send instead of
always sending both. The Agent Plugins package resolves an auth mode of its
own, so its proxy still names the operator to a trusted server.

The free half of the same theme: `session-start-commit.mjs` carried a private
`requestJSON` / `commitOvSession` pair that duplicated `ov-session.mjs`'s down
to the envelope. It imports them now, which is one header block fewer to keep
in step.

`wire-headers.test.mjs` drives six of the seven stacks through a stubbed fetch
and asserts one header map for all of them, in trusted mode and in api_key
mode; `auto-recall.test.mjs` covers the seventh, which only exists as a
subprocess, against a real server socket; and `mcp-proxy-config.test.mjs` puts
the proxy's two identity headers on the same switch.

* test(openclaw): run only the live test copies under tests/ut

Six files sat directly under tests/ as copies of suites that had since moved
on, and vitest's default include meant two of them still ran on every
`npm test` against an older shape of the code. What each one is:

- `tests/context-engine-assemble.test.ts` is an earlier cut of
  `tests/ut/context-engine-assemble.test.ts` — same imports one directory
  shallower, 360 lines against the live copy's 838. Five of its six cases are
  in the live copy verbatim; the sixth, "records senderId from runtimeContext
  in assemble diagnostics", is covered by the `senderIdFound` assertions in
  `tests/ut/context-engine-modules.test.ts` and
  `tests/ut/context-engine-afterTurn.test.ts`.
- `tests/context-bloat-730.test.ts` imports `memory-ranking.js`, `config.js`
  and `auto-recall.js`, and every symbol it exercises already has a home under
  `tests/ut`: `postProcessMemories` and `pickMemoriesForInjection` in
  `memory-ranking.test.ts`, `buildMemoryLinesWithBudget` and
  `estimateTokenCount` in `build-memory-lines.test.ts`,
  `recallScoreThreshold` and `recallMaxInjectedChars` in `config.test.ts` and
  `query-config.test.ts`.
- `tests/test-memory-chain.py` and `tests/test-tool-capture.py` are earlier
  drafts of `tests/e2e/test-memory-chain.py` and
  `tests/e2e/test-tool-capture.py`, which drive the same phases against a live
  gateway at greater length.
- `tests/demo-memory-ajie.py` and `tests/demo-memory-xiaomei.py` are one
  manual gateway demo under two personas, differing in the user name and the
  port, and named by no doc, script or manifest.

Nothing referenced them: the only prose pointer to an assemble test names the
`tests/ut` copy, and `.clawhubignore` plus `tsconfig.build.json` already kept
the whole directory out of the published package.

So that a file dropped beside the suite is not silently picked up again,
`vitest.config.ts` now pins `test.include` to `tests/ut/**/*.test.ts`. The
CI job's `--exclude` filters still apply on top of it — 42 files listed, 41
after the architecture-boundaries exclude. The one file left in `__tests__`
falls outside the pin until it moves into `tests/ut` with the rest.

`architecture-boundaries.test.ts` kept `tests/context-bloat-730.test.ts` in
the list of files it reads for index-facade imports, which would have thrown
on the missing path, so that entry goes too.

* test(openclaw): move the shouldBypassSession cases under tests/ut

`vitest.config.ts` now collects only `tests/ut`, and the six cases in
`__tests__/bypass-session-patterns.test.ts` were the sole behaviour coverage of
`compileSessionPatterns`, `matchesSessionPattern`, `shouldBypassSession` and the
`ingestReplyAssistIgnoreSessionPatterns` fallback in the config schema, so they
had stopped running. They move beside the other `text-utils.ts` tests, one
directory deeper, and the emptied `__tests__` directory goes with the two build
files that still had to name it.

* docs(openclaw): drop five reports that nothing points at any more

Four are dated one-off write-ups of manual runs, and three of them name the
repository and branch they were made against, neither of which is this one:

- `docs/openviking-install-real-scenario-verification-report.md` (2026-06-05)
  walks one machine's install of plugin `2026.6.2` from
  `iaasng/arkclaw-openviking-plugin`, against a hosted endpoint and a specific
  OpenClaw build.
- `docs/openviking-dynamic-query-config-test-report.md` (2026-06-04) is the
  acceptance run for the runtime query-config work in that same repository;
  `docs/openviking-runtime-query-config.md` is the reference that documents the
  feature and stays.
- `docs/workmemory-v2-test-report.md` (2026-05-02) is a `locomo10` benchmark of
  Working Memory v2, beside the design note that still describes it.
- `openclaw-multi-tenant-test-report.md` records a session driven through two
  ad-hoc remote ports and a live Feishu bot, reproducible by nobody.

The fifth, `docs/oc-resource-skill-import-design.md`, is an RFC whose own
opening paragraph says the implementation went the other way: no unified
`ov_import(kind=...)` tool and no `/ov-import` command, which is the shape the
rest of the document is about.

One line pointed at any of them — the Working Memory entry in the guide's
reference list, whose neighbour is the design note and stays. Nothing shipped
moves: `docs/` is outside the package's `files` list, the top-level report was
never in it, and the ClawHub release workflow names no document.

* test(openclaw): drop the unreferenced tool-result compression bench

tests/toolresult_compression_tests/ was manual measurement tooling: run-sccs-bench.mjs
spawns an external openclaw binary turn by turn and reports token usage, so it asserts
nothing and needs three repositories cloned to hardcoded /root paths plus a live server
before it can run at all. Nothing outside the directory named it — the two invocation
lines the plan cites were in its own README.

The behaviour it exercised is covered by tests/ut: tool-round-trip.test.ts asserts that
an externalized tool result survives conversion with its preview text, its
viking://session/<id>/tool-results/<id> ref and original_chars intact, and tools.test.ts
asserts the restore half through openviking_tool_result_read / _search / _list.

* test(openclaw): assert the prepack script that regenerates shared at pack time

The packaged plugin's prepack gained a `sync.mjs` prefix when the generated
shared copies stopped being committed, but the manifest contract still pinned
prepack to `npm run build` alone. The openclaw-tests job runs that file, so the
suite has been red on a stale expectation rather than on a real contract break.

* chore(claude-code): drop the two unwired debug scripts

Neither debug-recall.mjs nor debug-capture.mjs is reachable from hooks.json,
install.sh, or the sync target, so they drifted away from the hooks they were
meant to mirror; debug-capture also rewrites the live session's capture cursor
through an API flow the plugin no longer uses. The doctor skill, both READMEs
and the capability reference lose their pointers in the same change.

* docs(pi): rewrite DESIGN.md around the modules that exist

The spec was written in the future tense of a plan only partly executed: it
specified an `index_builder.ts` that was never built (the profile block from
shared/profile-inject.mjs folded into systemPrompt took its role), gave no
section at all to config.ts or takeover.ts, carried per-file line estimates
that went stale on every commit, and repeated its Knowledge Index sample
twice. Its Implementation Order and Testing Strategy described work the code
and tests/ have both already done.

It is now one section per module actually shipped, each written from that
module's current code, plus the event walkthrough, the design ancestry and
the two rationales the code cannot state for itself. TAKEOVER.md's model,
runtime flow, compaction and failure modes move into the takeover section so
the one harness capability that is unique to pi is documented beside the
module that implements it; its configuration table was already in README.md.

* docs(codex): fold VERIFICATION.md into a README Testing section

The 320-line manual SOP was pinned to plugin v0.6.0 and every step it
walked is now asserted by the node --test suite that CI runs. Only its
two live legs -- extraction landing in the user namespace and the
interactive Codex smoke test -- need a real server, so they stay as a
short appendix next to the command that actually runs the suite.

* chore(opencode): drop three unreferenced helpers from lib/utils.mjs

makeMultipartRequest was a near-copy of makeRequest that no caller ever
reached, and the wired URI check is lib/viking-uri-guard.mjs, not
validateVikingUri; ensureRemoteUrl had no caller either.

* chore(pi): drop six unreachable OVClient methods

createSession, addMessageParts, addMessagePayload and deleteSession have no
caller: the OV session id is derived locally and the server-side session is
created implicitly by the batch messages endpoint, messages go out as one
addMessage or through that batch endpoint, and nothing in the extension
deletes a session. resolveScopeSpace and resolveTargetUri were a URI-space
rewriter that no tool path ever reached, and they held the only readers of
RESERVED_USER, RESERVED_AGENT and the resolvedSpaces cache.

The sync-barrier stub kept an addMessagePayload key for the same reason; sync.ts
queues and replays through fetchJSON, so the key asserted nothing.

* chore(dsh): drop nine unreachable OpenVikingClient methods

health, ensureSession, getSessionArchive, find, read, list, stat, forget and
addResource had no caller: the runtime only ever uses fetchJSON, healthResult,
ensureSessionResult, getSession, addMessage and commitSession, and the tool
surface the read and write helpers look like they serve is bridged through the
stdio MCP proxy, so they could never be reached.

Only tests called them. The peer-override and live-recall cases move onto
ensureSessionResult and healthResult, which take the same arguments, and the
409/ALREADY_EXISTS tolerance moves to a runtime test: initializeState carries
its own copy of that check, so the behaviour stays pinned where it now lives.

* chore(codex): drop the ov-credentials pass-through

scripts/ov-credentials.mjs re-exported two names from shared/credentials.mjs
and repeated that module's CLI verb-for-verb. Nothing ran it as a command, so
the file only added a second path to the same resolver, one that every
credential fix since the shared library landed had to remember to keep in step.

Its three importers now reach the shared module directly. The test file keeps
its name because .github/workflows/pr.yml lists it by path, and it asserts the
shared resolver either way.

* chore(install): drop the pre-stdio layout migrations

The rc-wrapper stripper and the two legacy-marketplace migrations cleaned up
after installs that predate the stdio MCP proxy and the unified `openviking`
marketplace name, both of which landed in July 2026. Every install since then
writes the current layout, so the three helpers ran on every upgrade only to
find nothing, while keeping four constants and a TOML rewriter alive for a
shape the installer can no longer produce.

The installer-side cleanup is what goes; both doctors still detect a leftover
`openviking-plugins-local` install and tell the user how to remove it by hand.
The Codex README's step list loses the line that promised the migration, and
the installer test that asserted a Cursor-only run left the rc blocks alone
goes with the code it was pinning.

* chore(claude-code): drop the dead archive-abstract loop

SessionStart rendered up to five `<archive-abstract>` entries from the session
context's `pre_archive_abstracts`, but the server hard-codes that field to an
empty array: openviking/session/session.py returns `[]` from both context
builders and says so in a comment ("保留字段返回空数组,保持 API 向下兼容").
The loop is dead twice over, since the field's entries are objects while the
filter kept only strings.

What resumed and compacted sessions actually receive is unchanged: the
`<session-archive>` block still carries `latest_archive_overview`, which is the
one field the endpoint fills. The capability reference loses the "≤5
pre_archive_abstracts" claim it made for claude-code alone.

* refactor(plugins): put the four JS hook stacks on one HTTP module

Claude Code, Codex, dsh and the thin-harness runtime each carried their own
AbortController, header block and envelope parser for the same server, which is
how the wire drifted apart in the first place. lib/ov-http.mjs owns that shape
now — Bearer only, identity headers only when the config says the server is
trusted with the operator's name — so the next change to it lands once instead
of four times.

Also fixes the usage line in credentials.mjs, which still named the wrapper
that was deleted out from under it.

* refactor(plugins): move the Claude Code and Codex session stacks onto the runtime

Claude Code's scripts/lib/ov-session.mjs kept its own copy of the session
helpers the shared hook runtime already had — a fetch builder, the pending-queue
enqueue, add-message, commit, and the two session getters — so a fix to how a
failed write is parked had to be made twice, and the two copies had already
drifted: only Claude Code's set pendingQueued / pendingEnqueueFailed and warned
about a non-retryable failure.

The runtime carries that contract now. makeAgentFetchJSON takes the timeout and
the actor-peer getter its callers vary (Codex knows its peer only after loading
state under the session lock, and Claude Code names one per call), and resolves
the workspace peer on demand so a caller that never reads it pays nothing.
addAgentMessage and commitAgentSession report what became of a failed write the
way the capture hooks log it, and commitAgentSession takes the retention payload
Claude Code's threshold commit sends. Both harnesses' modules are wrappers over
it, keeping Codex's result-or-null fetchJSON beside the envelope fetchJSONRes.

That report now reaches further than Claude Code: the three thin harnesses on
this runtime — cursor, trae/trae-cn and zcode — name a write no retry can fix
on stderr too.

* refactor(plugins): put the last header blocks on the shared builder

opencode and pi still carried a whole fetch stack of their own, and the doctor,
the MCP proxy, the Claude Code statusline probe and the Codex recall hook still
spelled the header block out by hand. All six go through lib/ov-http.mjs now:
opencode's fetchJSON and makeRequest are a shell over it that keeps its throwing
contract and its two hints, pi's client keeps its typed methods over the same
envelope, and the four header-only callers ask buildOvHeaders.

A test asserts what that buys: X-OpenViking-Actor-Peer appears in the plugin
family's non-test sources only in lib/ov-http.mjs, so the next hand-rolled
header block fails before it can drift. Its scope comes from the sync targets,
so a harness added there is covered without touching the test.

Codex's recall hook stops overriding the per-call actor peer, which makes its
legacy-peer sweep ask for the legacy peer instead of asking twice for the
effective one.

* refactor(claude-code): put auto-recall back on the shared recall core

Claude Code's auto-recall.mjs carried a line-for-line copy of the recall core's
query profile, lexical overlap boost, ranking, dedup, user-space resolution,
multi-source search and budgeted block builder — 280 lines that only its own
statusline numbers kept it from calling. Those numbers now come from the shared
side: buildRecallBlockDetailed returns the block together with the counts the
fallback builder already computed, plus a stage that says which path produced
it, so a host can tell an empty recall the server had nothing to offer from one
the score threshold emptied.

buildRecallBlock stays the string-returning shell over it that opencode and pi
call. The hook is left with the Claude Code envelope and the last-recall.json
snapshot the statusline reads, and a test covers all eight reasons that file
records — including a guard that fails if the ranking pipeline is ever copied
back into this plugin.

* chore(claude-code): drop the dead per-message capture filter

auto-capture.mjs's shouldCapture has had no call site since the hook moved to
batches: its length bounds, slash-command, punctuation-only and question-only
rules each drop a whole multi-turn batch for what one turn in it looks like, and
the comment at the batch gate has said so ever since. That left four regexes
reachable only from a function nothing calls.

The batch gate keeps what it always did — skip an empty batch, and in keyword
mode require some user turn to carry a trigger phrase — and its comment stops
naming the function that is gone.

* refactor(claude-code): read the transcript through the shared capture utilities

Claude Code's transcript walk was written twice — once in auto-capture, once in
subagent-stop — and neither copy was the shared one every other harness uses, so
a fix to block handling reached nine harnesses and skipped this one.

Both now go through scripts/cc-transcript.mjs, which re-exports the shared
capture utilities and adds only what is genuinely Anthropic-shaped: a
tool_result nests its output in a content array and names its call by id, and
that output is dropped from the turn's text while travelling verbatim in the
turn's tool part. The kept turns, their order and their parts are unchanged, so
the capture cursor still means what it did.

The shared sanitizer now also strips <system-reminder> blocks and
[Subagent Context] lines; without them this switch would have started sending
Claude's own notes to itself to OpenViking as if the user had written them.

* test(codex): pin the digest URI repair the shared compressor performs

The recall compressor is a small model spawned through `codex exec`, and it
occasionally rewrites a long viking:// URI while rephrasing a bullet, which
leaves a dead link in the injected digest. Since codex went onto the shared
`compressRecallContext`, every citation is snapped back onto the URIs the
server actually returned before the digest is injected; this test pins that
a mangled URI comes back repaired, with the short-input short-circuit
disabled so the compressor is actually exercised.

* refactor(plugins): open every hook entry with the shared stage

The Claude Code and Codex hook entries each read stdin, re-resolved the config
for the payload's directory and answered the enabled and bypass gates in their
own words, and the copies had drifted: Codex's PreCompact never re-checked its
own switch after the reload, and the two auto-recalls asked the two gates in
opposite orders.

runHookStage is that opening. An entry hands it a config loader, its gates and
its envelope, and keeps only its own work, which now returns what the envelope
should carry instead of writing stdout from wherever it happens to stop. The
stage also carries the bypass verdict for the hook that must not stop on it —
Codex's SessionStart still sweeps and replays for other sessions inside a
bypassed repository — and a single-shot emit for the one that answers before its
worker starts.

The write-path preamble stays in the entry: a detaching hook has to spawn its
worker before stdin is consumed, and only the entry knows the response its host
expects. Claude Code's uri-guard keeps its own code as well; it reads no config
and answers no gates.

* refactor(plugins): fold the agent URI guard into the shared one

lib/agent-uri-guard.mjs was a second module over lib/uri-guard.mjs, holding one
more hint table and nothing else, and the three harnesses that imported it each
carried their own stdin read, entrypoint check and deny envelope — thirty-odd
lines apiece for the choice between two envelope shapes.

evaluateUriGuard is that module's body with the two things a host actually
differs on made arguments: the hint table, and the tool names that host guards
at all — Claude Code and opencode never see the shell in their matchers, pi
does, and the set was previously implied by whichever table a host happened to
import. The deny envelopes and the hook plumbing move in beside it, so cursor,
trae and zcode keep only the envelope they answer with.

* refactor(plugins): let the last four URI guards call the shared evaluator

Claude Code, opencode, dsh and pi each kept a private copy of the guard loop —
normalize the tool name, look it up in a local table, sweep the arguments for a
viking:// URI, format the message — because their tables name different
replacement tools. The loop is the same everywhere; only the table is host data,
and it has to stay host data: read(uris="..."), openviking_read(uris=[...]) and
viking_read(uri=..., level=...) cannot be spelled from one set of strings.

So each of the four now hands its table to evaluateUriGuard and keeps only the
decision its host answers with. Claude Code's table was character-identical to
the shared one, so it passes the guarded set instead — read, glob and grep,
leaving Bash to the model as its test has always pinned — and its stdin read,
entrypoint check and deny envelope go the way the other hook guards' did.

* refactor(plugins): run both doctors from one shared entrypoint

The Claude Code and Codex doctors were forked from one commit and still carry
the same run: parseArgs and tryJson byte for byte, expandHome inlined, the same
eight sections in the same order, the same --json envelope and exit code. The
copies had drifted apart in the details — Codex never learned to say which node
PATH resolves to when it differs from the one running the script.

runDoctor owns that run now. A harness hands it its name, its CLI, the three
sections only it can produce (install, config, activity) and the spelling it
uses for the configured account and user; it gets the argument parsing, the
section order, the envelope and the exit code back. Codex keeps assessHooksFeature,
parseFeaturesList and the isDirectRun guard its test imports, and its auth-mode
verdict arrives through onSummary.

Both wrappers also stop scanning shell rc files for the pre-2026-07 wrapper
blocks: the installer that wrote them is gone, so only the OPENVIKING_* exports
those files may still carry are worth reporting.

* refactor(plugins): compose the doctor configuration section from shared segments

Both doctors printed the same section from two copies that had drifted, so a
check written for one host was invisible on the other. The section is now six
shared segments a host calls in order, and the three checks that only one of
them had — Codex's hook-budget warning, Claude Code's extra_headers X-API-Key
and NODE_TLS_REJECT_UNAUTHORIZED warnings — run on both.

* refactor(plugins): declare which recall knobs the request omits by default

recall-core sends limit, max_tokens, query_expansion and rewrite_max_bullets
only when a layer actually set them, so the server's own defaults stand where
the user expressed no preference. That decision lived in five hand-written
`<name>Configured` projections and nowhere in the schema, so a loader had no
way to know which knobs it owed a flag. The schema marks the four now, and a
test ties the marked set to the fields recall-core actually gates.

* refactor(plugins): assemble every harness configuration in one place

Six loaders spelled out the same sequence — read the credential files,
resolve the knobs, resolve the peer, derive the timeouts, the log path,
the user agent and the send-only-when-configured flags — and each spelled
a slightly different subset of it. ov.conf's `<harness>.authMode` reached
only the three that passed a legacy layer, two of the four knobs the
server defaults for itself were reported as configured on two harnesses,
and Claude Code resolved its api key on a chain of its own.

buildPluginConfig() is that sequence, once. What is left in a loader is
what only that harness knows: Claude Code's four-valued credential source,
Codex's on/off reading of the digest switch, opencode's three section
knobs, pi's older debug-log variable, dsh's cordis input.

The api key now resolves the same way everywhere, with ovcli.conf's
`plugin` section ranked where the file it lives in ranks — under that
file's own `api_key`, over ov.conf. Codex read it last, behind
`server.root_api_key`; its doctor, its reference table and the capability
reference said so, and now say what the code does.

* refactor(plugins): resolve the portable bundle's connection from the shared loader

The Agent Plugins package kept its own copy of the credential chain, the
User-Agent, the timeout and the debug logger, so every fix to the shared
resolution had to be mirrored by hand and the proxy fabricated a credential
source out of whichever config file happened to load. Importing the full
loader is not the answer either: it would pull the knob schema and the
workspace layers into a bundle that has no hooks to run them, and those
copies are committed. `buildProxyConnection()` is the connection half of
that loader, which is all a stdio proxy needs.

The api_key still ends at ov.conf's `server.root_api_key` even when
ovcli.conf pins the chain to itself: this package ships without an
installer, so an install that names only a url there has nobody to migrate
its key.

* refactor(plugins): run every thin-harness hook from one entry

Cursor, TRAE and ZCode kept eleven shim scripts between them whose whole
body was an event assignment and an import of the harness dispatcher, and
each spelled the plugin root its own way — `${CURSOR_PLUGIN_ROOT}`,
`__OPENVIKING_TRAE_ROOT__`, `${ZCODE_PLUGIN_ROOT}` — so the installer
carried one substitution branch per spelling and a template that used the
wrong one failed only at the first hook of a fresh install. The shared
entry takes the event and the client from argv, keeping TRAE's trailing
client-id form, and one placeholder leaves one regex to render it.

The installed hooks.json is now checked against the tree it was rendered
for: a command naming a script the install did not put on disk used to
surface only when the host first ran it.

* refactor(plugins): merge the three thin harnesses into one plugin

Cursor, TRAE and ZCode were three marketplace directories running the same
state machine — the same 2000ms session-start debounce, the same prompt
dedup by event id and 500ms window, the same recall cache and cross-process
lock — around four things that genuinely differ: the event vocabulary, the
response envelope, how a prompt is read out of the payload, and how a
finished turn is captured. Everything else was triplicated, so a fix landed
in whichever copy the author happened to open: the URI guard existed three
times over one shared evaluator, the MCP proxy three times over one shared
builder, and the guard that pins how many proxies exist counted eight.

`agent-hook-plugin` keeps one dispatcher on stage 4's `runHookStage`, one
URI guard, one MCP proxy and one doctor, and puts the four differences in
`hosts/<id>.mjs`. The adapters sit at that depth because
`../../memory-plugin-shared/lib` has to resolve both here and in an
installed `agent-integrations/<client>/`, so `hosts/<id>/` holds only what a
host reads as configuration. The installer copies per host and reads the
templates from there, and the thin harnesses get the doctor entry the full
plugins have had since stage 5.

The archive guard was reaching none of this: the three plugins' hooks.json
named the shared entry across the plugin boundary, so nothing required
their own dispatchers to be in the release archive. hooks.json names the
plugin's own `scripts/hook.mjs` again, which puts every adapter back in the
derived requirement set, and the checker is handed the directories the
staging script declares instead of listing whatever the stage happens to
hold.

* fix(install): reclaim the URI-guard hook entries on uninstall

`--uninstall` recognised an entry as OpenViking's only by the hook script it
named, and the URI guard was not on that list. Uninstalling Cursor therefore
left beforeReadFile and beforeShellExecution in ~/.cursor/hooks.json, and TRAE
kept its PreToolUse entry, all of them running a uri-guard.mjs under
agent-integrations/<client>/ that the same uninstall had just deleted: the host
reported a failing hook on every file read and shell command until the user
edited the file by hand.

The uninstall filter now reclaims anything the installer wrote, by the
OPENVIKING_INTEGRATION_ID the rendered command carries, the way the install-time
filter and the TRAE CLI uninstall already did.

* ci(plugins): select the plugin test files by glob

The step named 54 paths by hand, so a new test file was only covered once
someone remembered to add it there too. Seven globs select the same set: the
union and the old list differ in neither direction (82 files each), and no
generated shared/ copy holds a test for a glob to pick up by accident.

  examples/*/scripts/*.test.mjs            19
  examples/*/scripts/lib/*.test.mjs         2
  examples/*/servers/*.test.mjs             2
  examples/*/tests/*.test.mjs              20
  examples/*-plugin/*.test.mjs             11
  examples/memory-plugin-shared/*.test.mjs 27
  agent-plugins/*.test.mjs                  1

Two things the job could not see before. A stale lib/MANIFEST or committed
shared copy passed CI, because nothing looked at the tree after the generator
ran; `git diff --exit-code` does now. And the two installer tests ran beside
the other 80: staging the marketplace regenerates the shared copies those
files import, and both fork whole installs, which is the load the two
unexplained single-test failures appeared under during this work. They now run
in a step of their own at --test-concurrency=1, and
install-agent-hooks.test.mjs copies examples/ into a tmpdir so `--source dev`
no longer installs out of the tree everything else is reading.

Neither failure reproduced in five full runs here. The only failure five runs
did find was deterministic and local: a gitignored openclaw-plugin/shared left
behind by work that has moved to another branch, which the generator reports
as stale but never deletes.

The openclaw job keeps its exclude; only its counts change, measured on this
branch: 42 files, 675 tests, and 2 pre-existing architecture-boundaries
failures rather than 4.

* ci(plugins): bring pi into the plugin version gate

pi's version was a hand-written constant in config.ts, which put it out of
reach of check-plugin-version-bumps.sh: the extension is copied wholesale by
the installer and reports that constant as the build talking on the wire, so
a shipped change under a frozen version had nothing to catch it. The
extension now carries a package.json, EXTENSION_VERSION reads it through the
shared readManifestVersion, and the gate watches that file the way it
watches the other five manifests.

The manifest changes nothing about how the extension loads or ships. pi's
loader reads a package.json only for a pi.extensions field and falls back to
index.ts without one; pi install parses a local path as a directory rather
than an npm package, and only installs dependencies for git-cloned sources;
the installer's tar copies the directory whole, so the file is in the
marketplace archive and the installed copy resolves the same 0.3.0 the
constant held. "type": "module" states the format node was already detecting
per load, which is what the warning naming this file asked for.

Against origin/main the gate still reports four plugins: pi's manifest is
new on this branch, so it takes the same no-baseline skip agent-hook-plugin
takes, and is enforced from the next base onward.

* ci(plugins): merge the two npm plugin releases into one matrix

The dsh and opencode release workflows were the same 81 lines with seven
values swapped, so every fix to the publish flow had to be made twice and one
copy could drift unnoticed. plugin-npm-release.yml carries the flow once and
lists the two packages as matrix entries.

Every way the two files differed, and how the matrix expresses it:

  workflow name    one "Plugin npm Release"; the per-package "Publish
                   @openviking/..." survives as the job name
  job name         name: Publish ${{ matrix.package }}
  work directory   defaults.run.working-directory: ${{ matrix.directory }}
  concurrency      moved from the workflow to the job and keyed by
                   matrix.directory, so the two packages still queue
                   independently rather than behind each other
  push paths       the union of both plugin directories plus this filename;
                   the shared lib entry was already identical in both
  install command  matrix.install: npm ci for dsh, npm install for opencode,
                   which ships no lockfile
  name assertion   compared against matrix.package through an env var
                   instead of a literal inside the node script

The union filter means a dsh-only push also starts the opencode leg. That leg
validates and then stops at the npm view already-published check, the same
check that already made a shared-lib push a no-op for whichever package was
not bumped. Everything else is byte-identical to what both files ran:
permissions, checkout, node 24, the validate step, and the publish step with
its NPM_TOKEN-or-OIDC fallback and --provenance. actionlint reports nothing
on the new file.

RELEASE.md named the opencode workflow by filename, so its bullet now names
the merged workflow and both packages.

* refactor(plugins): give each skill copy its own sync target

SKILL_TARGETS was the one list in the generator with its own shape — a skill
plus an array of directories — so the delivery flag every other target carries
had nowhere to live, and nothing asserted that git holds the skill copies a
host installs by path. One entry per copy, `{ skill, dir, committed }`, and the
test that decides which generated copies belong in git reaches skills as well
as vendored modules.

Every skill copy is committed and stays that way: .gitignore covers the
vendored shared/ directories only, and a host installs a skill by copying its
path out of this repository, so a copy git does not hold ships nothing.
Nothing else moves — the same two skills reach the same six directories and the
generator writes the same bytes. openclaw-plugin and agent-plugins stay out on
purpose: their skills are different files over different tool surfaces, not
copies of these.

* refactor(install): move the installer's JavaScript into lib/install

install.sh carried a 328-line JSONC editor and three near-identical hooks/mcp
merges as node heredocs. Nothing could exercise them except running the whole
installer against a scratch HOME, and they had already drifted: the uninstall
copy of the ownership test learned to reclaim the URI-guard hook entries and
the install copy never did, so reinstalling left a stale guard behind.

jsonc-edit.mjs is the editor, moved verbatim behind one entry point.
host-json-config.mjs is the merge, once: a single read that refuses to
overwrite what it could not parse, a single atomic write, a single answer to
"is this hook entry ours", and three commands over them — write, remove and
merge-zcode. The ownership list is the uninstall side's, so an install now
replaces a stale uri-guard.mjs entry instead of appending beside it, and the
zcode fold reclaims by the same rule as the other three hosts rather than by a
substring of its own.

lib/install/ is the installer's code, not the plugins'. No shared module
imports it, so it enters no vendoring closure and no lib/MANIFEST; one
`./install/...` import from a module that hooks do import would put it in both,
and install-lib-closure.test.mjs now fails on exactly that.

Finding it is the one thing that is not a straight move. `--uninstall` runs
before any source is resolved and the documented uninstall pipes this script
from a URL, where there is no sibling directory to read — reaching for the
checkout there would clone a repository just to remove hooks. So the assembled
runtime keeps a copy of the directory beside the manifest modules, and
install_lib_dir looks next to the running script, then there, then at the
checkout.

install-opencode-jsonc.test.mjs is six in-process cases over the editor —
comments wherever they sit, trailing commas, a single-quoted value holding a
brace and a `//`, an existing nested mcp object, idempotence, and a server the
user disabled — plus the one end-to-end install that proves install.sh still
hands the module the right arguments. The capability reference's install.sh
line count goes with them; it was already wrong and this makes it wronger.

* refactor(tests): share the mock server and the hook runner

Seven test files carried their own copy of the same three helpers and the
copies had drifted: two writeJson signatures, and two withMockOpenViking
contracts, one of which called .catch() on the handler's return value and so
could not take a synchronous handler at all. Three more carried a hook runner
that turned a non-zero exit into a rejected promise — a result no test could
assert on, which is why the async ZCode test spawns the hook by hand where it
wants to read the exit code.

testing/support.mjs already held buildConfigForTest and is the right home for
the rest: it sits outside lib/, so sync.mjs vendors none of it into a shipped
plugin, and a helper imported from a *.test.mjs file would have registered
that file's own tests a second time.

Two things the shared versions do that no copy did. The mock logs every
request it saw and hands the log to the callback, which the three tests that
only recorded pathnames now assert against instead of keeping an array of
their own. And runHookScript never rejects: the exit code comes back like
stdout does, and the call sites say expectExit where a clean exit is part of
the expectation.

writeJson takes the status last and defaults it to 200, so the two-argument
callers read unchanged and claude-code's auto-capture moves its status to the
end. claude-code's auto-recall is not one of the seven — it never carried
readRequestBody — and keeps its two local copies.

* refactor(tests): make the config and credential contracts shared

The credential chain is one resolver every harness calls, but only codex
tested it, so a change to the chain broke six harnesses and one suite
reported it. That file moves to memory-plugin-shared, where the glob picks
it up as the shared contract it always was.

The five per-harness config suites had the same problem in reverse: each one
re-tested the layer stack, the peer order and the cwd rule that
plugin-config.test.mjs already holds every loader to, and buried the one or
two cases only that harness answers. Three contract tests finish that file —
a knob walked through every layer, a default told apart from a choice, and
the workspace layer read from the cwd the loader is handed rather than the
process's — and the per-harness suites keep what is theirs.

Trimmed: claude-code and codex scripts/config.test.mjs, opencode and pi
tests/config.test.mjs. Every case deleted from them is answered by a shared
one in plugin-config.test.mjs, workspace-peer.test.mjs or the credentials
contract this commit moves. dsh loses nothing: the cordis input is a layer no
other harness has.

Two things only a harness knows had no test at all, and now do: claude-code's
tri-state digest mode, which it reads from the shared on/off switch under
either env spelling, and codex's boolean reading of that same switch, where a
compressor counts as configured only once it has been told what to run.

* refactor(tests): test the MCP proxy contract once

The proxy is one shared module, but its protocol contract was tested through
two harness entrypoints — twelve cases under codex, one of them again under
opencode — and both entrypoints re-exported the factory they import purely so
a test could reach it. A change to the core meant editing two suites, and the
opencode case differed from codex's only in which User-Agent string its
fixture made up.

The twelve cases and their fixture move to mcp-proxy-core.test.mjs, beside the
two transport-error cases already there, and the file now has one proxy
builder: it takes a fetchImpl, so an injected fetch and a real loopback
upstream are the same harness. Neither host file keeps a case of its own —
every one of the twelve exercises the shared core through the harness's config
shape, which mcp-proxy-config.test.mjs already covers per harness.

One case the split never had: a proxy with no local tool provider. Most
harnesses run it that way, so the local-tool branches have to stay invisible —
the listing is the upstream's, and a call named like a local tool is still the
upstream's to answer.

`examples/*/servers/*.test.mjs` now matches nothing. The pattern stays, for the
next case that really is one host's, under nullglob so it contributes nothing
rather than reaching node as a literal; opencode's test script drops it.
opencode keeps its servers/mcp-proxy.mjs export map entry, which describes the
entrypoint the package exposes, not the deleted test.

* refactor(tests): test the shared runtime where it lives

Three suites tested shared code from a harness directory, so a second
harness reusing that code was covered only by accident. The recall
timeout case, the queue cases that never touch cc's ov-session, and
findLastHumanTurnIndex now sit beside the module they exercise.

findLastHumanTurnIndex moves with its cases: it is a generic helper on
the shared turn shape, and codex's capture-utils.mjs already re-exports
everything shared, so its one caller needs no edit.

* refactor(plugins): label each credential from the chain that resolved it

Three doctors each re-walked ovcli.conf and ov.conf to say where the url, the
key and the identity came from — the walk `resolveOpenVikingCredentials` had
just finished. They were not the same walk. Claude Code and Codex rebuilt the
key's origin by hand and ignored the `apiKeySource`/`credentialPath` every
config now carries, so a key the chain took from one layer could be reported as
coming from another; the thin-harness doctor read `apiKeySource` but then let
ov.conf's harness block name the account under a pin the chain stops above.

`credentialSources(cfg, cliConf, ovConf, { section })` lives in doctor-core now,
with the section defaulting to the calling harness. The key's layer comes off
`cfg.apiKeySource`; the files are asked only which field inside that layer holds
the value, which is the one thing the layer cannot say — `plugin.<harness>.
apiKey` from `api_key`, a harness block from `server.root_api_key`. The identity
follows the chain step for step, plugin section included.

The stage-5 guard could not see the three copies because the core never exported
the name. It does now, so a fourth one fails that test. The two harness test
copies move to doctor-core.test.mjs, beside the function they describe.

* ci(plugins): generate OpenClaw shared runtime before tests

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* docs(plugins): define hook and MCP development standard

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* docs(plugins): publish bilingual development guide

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* fix(plugins): restore runtime and distribution contracts

Generate shared dependencies before validation and source installs. Keep Codex retries owned by the transcript cursor and defer threshold commits until the tail is delivered. Enforce the shared hook enable gate, restore auth-mode environment precedence, and preserve pi request options.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* refactor(plugins): use a host-neutral hook manifest

Move the shared hook integration metadata out of the Claude-specific directory and update version checks, diagnostics, installation verification, archive validation, tests, and documentation.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-15 18:28:20 +08:00
t0saki 1d89f8d465 feat(plugins): derive the workspace peer from git, and let a repository carry its own config (#4595)
* fix(plugins): stop dropping ovcli.conf's plugin section, and unrot the sync test

`ov config add|edit`, the config wizard and `ov config switch` all rebuild
ovcli.conf from the `Config` struct, which has no `plugin` field and no
catch-all — so every write silently deleted the whole `plugin` section the
memory plugins own. `write_config_file` now carries over the top-level keys
`Config` does not model, `save_edited_config` reads them from the old name on a
rename, and `activate_config` keeps the active file's. Modeled keys are
deliberately not carried over: one that is `None` was cleared on purpose.
`KNOWN_CONFIG_KEYS` decides what counts as modeled, guarded by a test that
fails when a struct field is added without listing it.

sync.test.mjs kept its own copies of the target lists and they had drifted —
dsh, opencode and agent-plugins were each missing modules the sync ships, so a
stale vendored file passed CI. It now imports the lists from sync.mjs (whose
`main()` moved behind an entrypoint guard) and additionally fails on a module
no target claims or a banner-carrying orphan no target lists.

Also in this hygiene pass:

- postRecall dropped `peer_scope` on any 400/422, silently widening recall from
  the caller's own peer to the whole user root. It now retries only on an
  unknown-field rejection, remembers the downgrade so every turn stops paying
  for a rejected request, and both doctors warn while that memo is live.
- The session peer pins (Claude Code's `ws-peer-*.json`, Codex's
  `workspacePeerId`) carry a version, so a pin written under one derivation
  rule cannot outlive it. Derivation is unchanged, so today they re-derive to
  the same value.
- recall-session-wiring.test.mjs was never registered in CI and had rotted
  against a fourth RecallManager argument; assertion fixed and registered.

* feat(plugins): layered workspace configuration

A workspace can now carry `<root>/.openviking/config.json`, which a team
commits, and `config.local.json`, which stays private; a per-machine registry
under `~/.openviking/workspaces/` sits above both so the user keeps the last
word over any repository they clone. All three share one schema and one merge,
and they slot in exactly where ovcli.conf's `plugin` section already does, so
`OPENVIKING_*` still wins over everything.

Three modules, synced to all seven plugin targets:

- workspace-identity.mjs finds the workspace root and reads git's own idea of
  what the repository is called, using only filesystem reads. No `git`
  subprocess: hooks are fresh Node processes on prompt-level paths with budgets
  as tight as Codex's 3s SessionEnd, and this keeps working where git is absent
  from PATH or would refuse the repo over dubious ownership. Worktrees converge
  through `commondir`; submodules stay separate; `$HOME` and `/` are never
  workspace roots.
- workspace-config.mjs discovers, parses, filters and merges the layers, and
  records per-key provenance — which layer won and what it covered up.
- workspace-registry.mjs keeps one file per workspace rather than one listing
  them all, so concurrent hooks cannot lose each other's writes, and treats a
  path whose git identity has changed as a miss rather than inheriting the
  previous repository's peer.

These files are trusted without a prompt, because a hook is non-interactive and
any approval gate degrades into "run one command per workspace first". What is
refused instead is structural and costs nobody anything: connection and
credential keys are stripped loudly, `${VAR}` is never expanded, and
`cli_config_profile` — which decides which credentials reach which server — is
registry-only and name-only. What a committed file switches off is announced in
doctor rather than blocked.

An adversarial review pass over these three modules found 18 defects, all fixed
here and each now covered by a test. The one that mattered: `JSON.parse` keeps
`__proto__` as an own property, so a 128-byte committed file could write
straight into `Object.prototype`, and since `process.env` reads through the
prototype chain and the environment outranks ovcli.conf, that set
`OPENVIKING_URL` and `OPENVIKING_API_KEY` for the whole process — silently
shipping the user's real API key to an attacker's host. Prototype keys are now
dropped with a warning in both the strip and the merge. The rest: unbounded
recursion (a 4KB file could take out every sibling layer), the identity cache
storing a remote's embedded token at 0644, worktrees under a directory named
`modules` misread as submodules, the registry's negative-evidence check being
inert in its only caller, `Number(null)` pinning cost knobs to a bound, and
provenance lying when two layers disagree about a key's type.

`.gitignore` no longer ignores all of `.openviking/`, which would have stopped
a team's config.json from ever being committed; doctor warns when a workspace
still does. The schema maps `capture.commit_token_threshold`, matching the knob
the loaders actually read — the RFC's example named a turn-based one that does
not exist.

* feat(plugins): derive the workspace peer from git, configurably

The peer a workspace writes its memories under was the working directory with
every non-alphanumeric byte turned into a dash. That made the identity an
accident of where the repository happened to sit: a clone on another machine, a
rename, a worktree, or simply `cd examples/` each minted a separate, empty
namespace, and there is no server-side rename or merge to recover from it.

The default is now git's own idea of the repository. `peer.source` decides the
rule and reads from every layer — `OPENVIKING_PEER_SOURCE`, ovcli.conf's
`plugin.peerSource`, or `peer.source` in a workspace file:

- `git` (new default) ≡ `["{git_remote}", "{git_root}", "{cwd}"]` — the
  normalized origin, else the repository root, else the working directory. No
  preset adds a prefix; a path-derived id already starts with `-` on POSIX, so
  it cannot collide with a remote-derived one.
- `cwd` — the old rule, byte for byte.
- `none` — send no peer. `OPENVIKING_WORKSPACE_PEER=0` still means this.
- Any template, or a list tried in order, over `{git_remote}` `{git_root}`
  `{cwd}` `{dir}`. Substitution is all-or-nothing: an empty variable falls
  through to the next template rather than leaving a half-formed shared id.

So `/Users/x/Dev/OpenViking/examples/codex-memory-plugin` with origin
`git@github.com:volcengine/OpenViking.git` is `github.com-volcengine-openviking`
from any subdirectory, worktree, machine or clone. Every clone of one repository
shares one peer; a fork has a different origin and stays separate, and
`gh pr checkout` of someone else's PR does not change origin, so reviewing does
not move a session's memory.

Nobody has to migrate. The pre-git id is always recomputable locally, so
`resolveEffectivePeerId` returns it alongside the effective one and recall still
reaches it: under the default `peer_scope: "all"` the server's cross-peer sweep
already covers it for free, and under `"actor"` — where that sweep is off by
definition — the plugin asks that peer separately, as itself, which is cheaper
and reaches more than a bare cross-peer read would. There is no deadline on
this. Wired through all five recall paths; doctor names the previous peer and
says which of the two is carrying it.

`source` keeps its three values because five call sites compare it against the
literal `"workspace"` to decide whether a session pin may be reused; the new
`origin` field names the template that actually produced the id, and doctor
prints it. Both session pins bump their version, so a pin frozen under the old
rule cannot outlive it.

* docs: the workspace peer comes from git, and a workspace can carry config

Every page that described the peer as the working directory with its
non-alphanumerics dashed now describes `peer.source` and the git default, with
the presets, the template variables, the clone-vs-fork identity semantics, and
why no migration is required. The capability reference gains the three new
shared modules; the client configuration page gains a Workspace Configuration
section covering the two workspace files, the per-machine registry, the
precedence table, the v1 schema and what a workspace file may not set; the
Claude Code and Codex integration pages, which never mentioned peer derivation
at all, each gain a short section. All zh mirrors follow.

`examples/schemas/workspace-config-v1.json` is the schema the `$schema` key in
a workspace file points at, and `examples/workspace-config.example.json` is a
file to copy. `ovcli.conf.example` shows `plugin.peerSource`.

Two facts worth stating plainly, both verified against the loaders rather than
assumed: `OPENVIKING_PEER_SOURCE` and the workspace-file layer are read only by
the Claude Code and Codex plugins today, so the other harnesses run on the
default and their pages document the config key rather than an env var that
would be inert; and pi and dsh compute their legacy id from the process cwd,
so their pages promise dual-read only under the default `peer_scope: "all"`.

* fix(mcp): stop the proxies from guessing a peer out of their launch directory

Three MCP proxies keyed the actor peer off `process.cwd()`, which for a
long-lived server started from a static MCP config is the directory the harness
happened to launch from — often the plugin's own. The Codex proxy already
refused this and had a test forbidding it; the rule now holds for all of them
through the same `resolveMcpActorPeerId`.

The plan called for the parent process to inject `OPENVIKING_PEER_ID` at launch
instead, following dsh's `mcp.mjs:27`. That only works for dsh: Claude Code and
agent-plugins are launched from a static `.mcp.json`/`mcp.json` with no
environment block, and OpenCode's `createOpenVikingMcpConfig` builds a command
and args with nowhere to put one. So the fix is to stop guessing rather than to
guess better — a proxy sends no actor peer, which is broad recall, the default.

`resolveMcpActorPeerId` now warns and widens where it used to throw. Refusing to
start took away every memory tool because a scope preference could not be
honoured, which costs the user far more than the wider search does; the warning
says which two settings would scope it.

* test(pi): follow the git-derived peer default rather than pinning the cwd id

* fix(plugins): warn on the camelCase spelling of a connection key too

The projection into harness knobs is an allowlist, so `apiKey` in a workspace
file could never take effect — but it vanished without a word, which reads as
acceptance. It is refused by name now, like its snake_case twin.

* feat(cli): ov workspace show, and ov peer link|migrate|forget-previous

`ov workspace show` answers "which layer actually set this" the way
`git config --show-origin --show-scope` does: the workspace root and how it was
found, the git remote, every template variable, the effective peer and the
template that produced it, each config layer with whether it applied, and per
key the effective value plus everything it shadowed.

That question matters here because three languages read this configuration and
each could drift. So the Rust reader is not a paraphrase of the JS one — the two
were run side by side over the identity helpers, the merge with full provenance
trees, the file-read rules and the registry's raw bytes, and made byte-identical.
That comparison paid for itself: it caught `serde_json::Map::remove` being a
swap remove under `preserve_order`, which reshuffled a registry file the JS half
reads on every rewrite.

It also caught the divergence that would have made the command a liar. ovcli.conf's
`plugin` section speaks the flat knob names a harness loader reads (`peerSource`,
`recallLimit`); a workspace file spells the same settings nested (`peer.source`,
`recall.max_items`). Both are one chain in `loadPluginSettings`, and the port had
merged the flat file into the nested tree, so `plugin.peerSource: "cwd"` in
ovcli.conf left `ov workspace show` reporting the git-derived peer while every
plugin sent the cwd-derived one.

`ov peer link <id>` pins a peer for this workspace in the registry — the way out
of a fork that should share the upstream's memory, or a legacy id worth keeping.
`ov peer migrate` moves a peer's memories and resources with the fs mv API,
reporting the plan by default and requiring `--apply`; the server has no merge
semantics, so a collision is refused with the colliding path rather than
overwritten, and a listing that fills its limit aborts rather than planning from
a truncated view that could hide one. `ov peer forget-previous` clears the
recorded ids.

`workspace show`, `peer link` and `peer forget-previous` are local and do not
require ovcli.conf; `peer migrate` talks to the server and does. Both config
gates and the hand-rendered help are registered, with a test pinning the gates
against each other.

A test now reads FORBIDDEN_KEYS, REGISTRY_ONLY_KEYS and FREE_FORM_SECTIONS out
of the JS module and compares them to the Rust constants, because those lists
are what someone fixing a bug in one language edits — and they had already
drifted once while this was being written.

* build(cli): record the sha2 dependency edge in Cargo.lock

Already vendored for other workspace members; ov_cli now uses it for the
registry slot hash.

* test(plugins): the fixtures the plan named that were still missing

A moved or renamed repository keeping its identity is the change's whole point
and had no test; a shallow clone was worth pinning because it is exactly what
the rejected root-commit scheme could not answer; and the registry's
read-modify-write window between two hooks of one session is now written down as
a test rather than only as a comment.

* fix: the defects a plan review turned up

An independent review against the plan this branch was built from found ten
real defects. Each was reproduced before being fixed and is now covered by a
test.

The four that broke a promise the feature makes:

- The workspace config layer was resolved from the hook process's own working
  directory, not from the `cwd` on its stdin payload — which is the
  authoritative one. A hook started in one repository while the session sits in
  another applied the wrong `.openviking/config.json`: its peer, its bypass
  patterns, its `capture.enabled`. `loadConfig` now takes the directory, and
  every hook that receives one re-resolves with it. Late re-resolution is safe
  precisely because connection and credential keys are structurally forbidden
  in a workspace file, so `baseUrl` and `apiKey` cannot move under an already
  built client — the gates that a workspace can switch off moved below the
  parse so they are decided on the right config too.
- Codex threw away a workspace file's `peer.id`: it returned the credential
  chain's peer verbatim, so `{"peer": {"id": "team-a"}}` did nothing. Claude
  Code had always honoured it.
- Claude Code's session pin returned only the id and source, dropping
  `legacyPeerId` — so from the second hook of a session onward, dual-read
  stopped asking the peer that holds everything written before the derivation
  changed. Silently, and exactly where it mattered.
- `ov peer migrate` read the source peer with the actor-peer header set, which
  the server refuses for another peer's path, and treated every `stat` error as
  "does not exist". The common case — an `actor_peer_id` in ovcli.conf — got a
  cheerful "Nothing to migrate" instead of a 403. It now uses a client with no
  actor peer and tells a real error apart from an empty source.

The rest:

- A directory outside any repository is a workspace again. It had no root at
  all, so a `.openviking/config.json` there was ignored entirely. `$HOME` and
  `/` are still never roots, now judged on the starting directory rather than
  on where the upward walk stops, and `git_root` stays empty outside a
  repository so the `git` preset still falls through to `{cwd}`.
- The registry slot is keyed on identity, not path. Two linked worktrees of one
  repository are one workspace — one peer, so one set of settings and one
  `ov peer link` — and keying on the checkout path split them in two. This also
  makes crossing two repositories physically impossible rather than merely
  detected. (Their `config.json` files still follow each checkout; those are
  files on a branch.)
- git folds section and key names to lower case, so `[Remote "origin"]` with
  `URL = …` is a remote `git config` reads and we did not. A quoted subsection
  stays case-sensitive.
- `min_client_version` warns instead of being silently kept as data, and still
  never blocks.
- `cli_config_profile` was validated and then never used. It now selects
  `~/.openviking/ovcli.conf.<name>` before credentials resolve — registry-only,
  name-only, and a hard error when the profile is missing, because quietly
  authenticating somewhere the user did not choose is the failure the key
  exists to prevent.
- `ov workspace show` is exempt from the language gate. It is a diagnostic and
  has to work on a machine where no language was ever chosen; `peer link`,
  `migrate` and `forget-previous` mutate state and still gate.
- `ov peer link` records the peer it replaced, so a later `migrate` with no
  `--from` finds it instead of falling back to a recomputed cwd id.

Two more the plan asked for that were missing: doctor now checks the knobs
*inside* ovcli.conf's `plugin` section — it was on the allowlist, so until now
`peerSorce` sat there doing nothing with no complaint — and suggests the key
you probably meant. A test derives the known-knob set from what the two loaders
actually read, so the list cannot rot into one that rejects a real knob; it
caught a missing entry the moment it was written.

The RFC is archived at docs/design/, with the three claims implementation
disproved corrected in place: the knob is `commit_token_threshold`, `__self`
and `ext-` are not reserved server-side, and a worktree converges its identity
rather than its config file.

* revert(cli): withdraw ov workspace and ov peer from this branch

The command surface these two files added was larger than the feature they
served: 4792 lines of Rust for `ov workspace show` and
`ov peer link|migrate|forget-previous`, against a change whose whole point is
what the hooks send. None of it had reached a user-facing document — only the
RFC named it — so it goes back out whole and the branch becomes a plugin
change plus one CLI bug fix.

Restored from the branch's merge-base rather than from origin/main, since main
has moved on since the branch was cut and those commits are not this branch's
to carry. `sha2` was pulled in only by `workspace.rs`, so its dependency edge
leaves with it.

What stays is `config.rs` and `config_wizard/store.rs`: `ov config add|edit`
dropped the whole `plugin` section because the wizard round-tripped the file
through a typed struct, and that fix has nothing to do with the withdrawn
commands.

The registry under `~/.openviking/workspaces/` stays too, as a layer the
plugins read. Nothing writes it for now; `ov-memory-doctor` prints the path it
expects, and the file is small enough to create by hand. A writer can come back
on its own merits.

* fix(plugins): derive a peer only inside a git repository

Codex desktop opens a directory per task — `~/Documents/Codex/<date>/<slug>/` —
and none of them is a repository. The `git` preset ended its fallback chain at
`{cwd}`, so every one-off task minted its own empty peer, and each new one
started with no memory. Nine such directories here, nine peers.

There is nothing app-specific to read: the state file that lists those threads
is Electron-private, a megabyte wide, desktop-only, and would have to be parsed
inside SessionEnd's 3-second budget. The signal that generalizes is structural
— the directory is not a repository, and nothing in it says it is a project.

So the default chain is now `["{git_remote}", "{git_root}"]` and stops there. A
directory that is neither a repository nor marked gets no peer at all, and what
is remembered in it goes to the user-level space, which is where it went before
peers existed. Deriving an identity from a bare path is what `peer.source:
"cwd"` is for, and it is a word away.

Naming such a directory is the other half. `findWorkspaceRoot` now also stops
at a directory holding `.openviking/config.json` or `config.local.json`, so a
marker file works from any depth below it, the way a repository does — and when
that marker sits inside a repository the git variables still resolve to the
enclosing repository, so marking a subdirectory of a monorepo does not split
the default peer. `{git_root}` is the repository's root, `{dir}` the workspace
root's name whichever made it one.

Nothing moves. When no template resolves, the pre-git id is still computed and
returned as `legacyPeerId`, so `peer_scope: "actor"` keeps asking for it and
`"all"` keeps sweeping it.

Two doctor bugs fell out of the same walk: `checkWorkspace` read `git.kind`
unconditionally and threw wherever there was no repository, and the peer block
warned "set peer.source to git" at a directory where `git` is exactly what is
already set and correctly resolves to nothing. It now says why no peer is sent,
and prints the file to create.

* docs(plugins): say how to give a directory its own peer, to users and to agents

The behavior change is only useful if the reader can act on it, and two kinds
of reader have to: the person whose scratch folder stopped having a memory, and
the coding agent they ask about it.

`docs/{en,zh}/configuration/02-client.md` is the one place that spells the rule
out, and everything else links to it. It gains "Give a Directory Its Own Peer",
which opens with the file to create and then the ladder above and below it;
"By Situation", eight rows from fork to throwaway folder; and "Recall
Isolation", which separates where memories are written from what is read back,
names the server's per-category penalties, and states the cost of sending no
peer outside a repository — a user-level memory is read at full weight in every
project afterwards.

The eight integration pages, both capability references, six plugin READMEs,
the changelog and the schema stop promising a fallback to the working
directory. Checking those claims against the loaders turned up one that was
never true: opencode, dsh and pi do not read workspace files at all, so a
`peer.id` written for them does nothing. Said plainly rather than left to be
discovered.

For agents, `openviking-memory/SKILL.md` gains ten lines on where memories are
filed — it is the skill that fires when someone asks why a folder has no
project memory, and it had nothing to say — and both `ov-memory-doctor`
references gain a table from what the user says to the exact key to write.
`ov-memory-doctor` prints the same snippet, so an agent that runs it needs no
further reading.

One snippet has to be identical in twenty places for any of this to hold, so
`WORKSPACE_PEER_HINT` is a constant the report builds its line from, and
`peer-guidance.test.mjs` asserts it appears verbatim wherever it is promised,
that no page still spells the retired chain or names a command this branch
withdrew, and that every variable the canonical page documents is one the code
substitutes. It asserts no prose: rewording a page must not turn it red.

* fix(plugins): reject an unrecognized peer.source instead of using it as an id

`peer.source` accepts a preset name, a template, or a list of templates, and
anything that is not a preset was treated as a template. A template with no
`{...}` in it renders to itself, so a typo became the peer: `"Git"` wrote every
memory under a peer literally named `Git`, and `"gti"` under `gti`. Silently —
the wrong namespace is indistinguishable from an empty one until someone
notices their project has no memory.

A bare string that is neither a preset nor contains `{` now warns and falls
back to the `git` default. A list is still taken at face value: writing one is
explicit enough that a typo inside it is a different kind of mistake.

The warning travels through an optional `onWarn` callback falling back to
stderr, matching `resolveMcpActorPeerId` in `mcp-proxy-config.mjs` — there is no
warnings array in reach, because `resolveEffectivePeerId` is called from the
hook runtime and from four harness config loaders, none of which thread one.

* docs(plugins): say what each harness can actually do with a peer

The peer documentation promised the same thing everywhere, but only the Claude
Code and Codex plugins read workspace configuration files. `loadPluginSettings`
is called from exactly two loaders; the other harnesses build their config from
their own file plus the environment. So a reader following the docs under pi,
dsh, opencode or cursor would create `.openviking/config.json` and watch it do
nothing.

The skill is the sharpest case: `openviking-memory/SKILL.md` is synced to
cursor and dsh as well, and it told an agent to write that file. An agent would
have done it, reported success, and changed nothing. It now names the two
harnesses that read it and points everyone else at `OPENVIKING_PEER_ID`.

The integration pages had started teaching the recipe and then retracting it in
the same sentence, which is worse than not mentioning it; they now carry the
one instruction that works there, and link to the canonical section for the
rest. The capability reference gains the same qualification, next to the
paragraph that already says only two harnesses read those layers.

Two smaller corrections. The doctor references had the same question answered
twice, once in the peer table and once in the troubleshooting table 140 lines
below; the troubleshooting row survives, since it carries a diagnostic column.
The RFC still archived implementation notes for the CLI this branch withdrew,
which would read as a description of commands that exist.

`peer-guidance.test.mjs` guards this alignment, and had two flaws of its own: it
swept the changelogs, which are generated from GitHub Releases and would go red
on a release note nobody wrote by hand, and it sliced a page between two
headings with `indexOf` without checking either was found — renaming the closing
heading would have silently scanned to end of file.

Also here, because it is the same kind of mismatch: the zcode MCP proxy sent an
actor peer under broad recall, where the other proxies leave the header unset.
It now routes through `resolveMcpActorPeerId` like they do. The dsh proxy
deliberately does not — its parent process resolves the peer per session and
injects it into the child environment, so it is not guessing at a launch
directory, and that reason is now recorded next to the line.

* fix(plugins): ship and install only the shared modules a plugin imports

Three new modules were fanned out to all seven plugin directories in one hunk,
and only two plugins call them. That left dead weight in five directories, and
it broke three installs.

The install is the part that mattered. `install.sh` copies a hand-written list
of shared files into `~/.openviking/agent-integrations/memory-plugin-shared/lib`,
where cursor, TRAE and TRAE CLI import from. The list names `workspace-peer.mjs`
but not `workspace-identity.mjs`, which `workspace-peer.mjs` imports — nor
`workspace-config.mjs`, which identity had come to import for three filename
constants. Copying exactly that list and importing the hook runtime fails with
ERR_MODULE_NOT_FOUND, so every hook of those three harnesses would have died on
startup. Nothing tested the list.

`CONFIG_DIR_NAME`, `TEAM_FILE` and `LOCAL_FILE` now live in
`workspace-identity.mjs`, which is where the walk that recognises a marked
directory needs them; `workspace-config.mjs` imports and re-exports them, so no
call site moves. Identity has no library-internal dependency left, which is what
makes the installed set closed at fifteen files instead of pulling the whole
configuration layer along behind it. `install-lib-closure.test.mjs` derives both
sides — the list parsed out of the shell script, and the transitive imports of
the three entrypoints — and fails in either direction, so neither a new
dependency nor a stale entry can go unnoticed again.

With identity standing alone, the fan-out can follow what is actually imported.
`sync.mjs` moves from arrays chained by spread — where the harness that does not
need a file is often the one the array is named after — to explicit per-target
lists. `plugin-config.mjs`, `workspace-config.mjs` and `workspace-registry.mjs`
leave dsh, pi, opencode and zcode, none of which import them; all four workspace
modules leave agent-plugins, whose only importer was deleted a half hour after
they arrived and whose peer is environment-only by design. That is about 4500
lines of vendored code that said something the code did not do.

The registry loses its write path in the same spirit. `writeEntry`,
`rememberPreviousPeer` and `listEntries` had no caller outside their own tests:
the CLI that would have written them was withdrawn from this branch. `readEntry`
and `entryPath` stay, because a hand-created entry is still read and the doctor
still prints where to put one. `cli_config_profile` goes with them — the whole
mechanism, down to the documentation that described it, since nothing ever
resolved a profile through it.

This is not a new policy. `HARNESS_KEYS` already carried the rule in a comment,
added the day after the same speculative fan-out happened in August: add a key
as its loader starts calling `loadPluginSettings`, not before, so the section
never promises a knob that silently does nothing.

* fix(dsh): thread peerSource into the per-session peer

The integration page documents `peerSource` in dsh's Cordis patch, but `stateFor` never passed it to `resolveEffectivePeerId`, so the key resolved to nothing and every dsh session ran on the default derivation. Pass it, and keep the pre-git id alongside so dual-read reaches memories written before the default changed.

* docs(plugins): correct the shared-layer counts after the distribution changed

The capability reference still described the pre-branch distribution: 18 library modules against 23, per-target counts from before each target stopped receiving the workspace configuration layer, and `workspace-config` / `workspace-registry` listed as reaching every JS harness when only claude-code and codex load them. It also still said `cli_config_profile` was registry-only, and that the registry is written for you.

* docs(rfc): lead with a TL;DR of the workspace config and peer source proposal

* feat(plugins): let peer.source name the harness with {harness}

The peer templates could describe where a checkout sits but never which
agent was running in it, so one repository could not keep a separate
memory per agent even when its user wanted that. The harness name was
already in every config, only baked into the User-Agent string.

No preset uses the new variable: sharing one project memory across
agents is the more useful default, so splitting stays opt-in via a
template such as "{git_remote}-{harness}". It is composed at render
time rather than in the workspace identity, whose result is cached
under a cwd-only key that two harnesses in one directory would share.

* docs(rfc): record git_branch and peer.command as directions, not deliverables

* chore(plugins): resync the openclaw vendored recall-core after the rebase

* test(opencode): follow resolveEffectivePeerId's widened return shape

Also mark the openclaw shared copies generated, the way every other sync
target already is.
2026-09-04 12:50:43 +08:00
t0saki 7200cdb176 feat(codex): commit on SessionEnd hook, keep SessionStart sweep as fallback (#4429)
* feat(codex): commit on SessionEnd hook, keep SessionStart sweep as fallback

Codex ships a SessionEnd hook since rust-v0.145.0. The plugin now commits the OpenViking session from that hook (marker + detached worker within the 1s/3s budget, catching up turns the last Stop never sent) and reduces the SessionStart active-window heuristic to an ended/idle sweep. Adds a per-session lock so the Stop worker, PreCompact, SessionEnd worker and sweep no longer clobber each other's state writes, stops an unreadable transcript from resetting the capture cursor, and merges the three copies of the HTTP/transcript helpers into scripts/ov-session.mjs.

* fix(codex): guard SessionEnd partial catch-up, make the sweep catch up, token the end marker, own the lock

Review follow-up for #4429: SessionEnd no longer commits after an incomplete catch-up (tail turns were archived away); capture hooks record transcriptPath so the SessionStart sweep catches up before committing; the .ended marker's timestamp acts as a token that the SessionEnd worker and the sweep re-verify under the lock, and Stop/PreCompact/resume only clear markers older than their own start; withSessionLock stamps an owner file and takes over stale locks by atomic rename with an inode check so a taker never moves a lock a racer just created. Plugin 0.8.1.

* fix(codex): catch up ended sessions with no live id, guard unreadable transcripts, make the end marker and the lock takeover race-free

Four defects found reviewing the SessionEnd commit path:

- The SessionStart sweep cleared the `.ended` marker of a state with no live
  ovSessionId before taking the lock, so a session PreCompact had released and
  that then produced more turns lost its tail when its SessionEnd worker died.
  A marker now always enters the lock, and only a catch-up that finds nothing
  new may clear it.
- catchUpTurns reported an unreadable transcript as an empty one, so all three
  callers committed and archived a session whose tail they never saw. It now
  reports `unreadable` and they keep the session live for a later retry.
- clearEnded read the marker's timestamp and then removed the path, deleting a
  marker a concurrent SessionEnd had just written. The timestamp moved into the
  filename, so a conditional removal targets an immutable path.
- session-state.test.mjs was missing from the CI test list.

Also replaces the lock's stale takeover: renaming the directory aside left the
lock path momentarily absent, which let another racer's mkdir succeed next to
the taker. Takers now race for the `owner` file inside the directory instead.

* fix(codex): make the end marker's generation unique within a millisecond

Date.now() alone is not a generation: two SessionEnd hooks in the same
millisecond produced the same marker path, so the first one's conditional
removal took the second one's marker with it. The marker is now created
exclusively and its timestamp bumped until that succeeds; bumping rather than
randomizing keeps the names ordered, which is what the `before` cutoff
compares. The race test drops its artificial 2 ms gap and runs 500 iterations.
2026-08-31 12:23:33 +08:00
t0sakiand7487 a61c57bf59 fix(codex): preserve transcript cursors across commit and compaction (#4191)
* fix(#4058): [Bug]: Codex memory plugin replays historical turns after resume or transcript compaction

Fixes #4058

Ref: https://github.com/volcengine/OpenViking/issues/4058

* fix(codex): retire committed cursors and keep activity-based concurrency

Preserving the transcript cursor after a commit stops the replay, but it also
means nothing deletes state files any more: clearState() lost its last caller,
so every codex session — including ones that never captured a turn — leaves a
file behind, and listStates() reads all of them on every SessionStart.

The sweep now retires cursor-only states in the same pass: a real cursor is
kept for resume until OPENVIKING_CODEX_COMMITTED_TTL_MS (default 30 days, past
the life of the codex rollout it indexes), and a state that never captured
anything goes on the idle schedule, which is what the old sweep did with it.

Releasing ovSessionId also wrote lastUpdatedAt, making a committed session look
freshly active; saveState() takes touch:false so the field keeps meaning "last
transcript activity" for both the active window and retention.

Requiring a live ovSessionId to count as recently-active made the heuristic
miss sessions PreCompact had just committed, which can still be running: the
count is back on activity alone, and only a state with a live session is
committed.

Also name the shrink predicate: role === "user" covers tool results too
(normalizeCaptureRole maps them onto the user role), so findLastHumanTurnIndex
requires a text part, and the no-human-turn fallback to a full replay is now
visible in the log instead of silent.

---------

Co-authored-by: 7487 <1042653432@qq.com>
2026-08-26 11:44:40 +08:00
t0saki a83b81715b feat(uri)!: remove uid-less current-user shorthand in favor of viking://~ (#4196)
* feat(uri)!: reject uid-less current-user shorthand in favor of viking://~

viking://user/<segment> (memories/resources/skills/peers/privacy/sessions
without a user id) was ambiguous with a user literally named after the
segment, and a user actually named e.g. "memories" was unreachable for
USER/ADMIN callers. Now that the viking://~ home alias (#4167) covers the
same need unambiguously, the shorthand fails closed at the request
boundary instead of expanding:

- resolve_current_user_uri raises NamespaceShapeError with a corrective
  hint naming both viking://~/<rest> and the explicit-uid form. Silently
  parsing the reserved segment as a peer user id would misdirect reads
  and writes, so rejection is the only safe removal.
- Bare viking://user falls through to the canonical parser and keeps
  container semantics (a user key listing it sees only its own space).
- The self-id escape stays: a caller whose user_id equals a reserved
  name keeps viking://user/<own-id> as their canonical root. ROOT-role
  literal parsing and the legacy viking://session alias are unchanged.
- AddTargetsConfig normalizes stored legacy config spellings
  (viking://user/resources|skills) to the viking://~ form at validation
  so existing ov.conf/user_config deployments keep working; the accepted
  per-user spelling is now viking://~/resources and viking://~/skills.
- usage_reporter keeps canonicalizing the historical shorthand found in
  old transcripts and additionally recognizes viking://~/memories/.

BREAKING CHANGE: requests using the uid-less viking://user/<segment>
spelling now fail with 400; use viking://~/<segment> or an explicit
viking://user/{user_id}/<segment> URI.

* refactor(clients): migrate first-party emitters to the viking://~ home alias

Every in-repo client that emitted the removed uid-less current-user
shorthand now sends viking://~/... instead: vikingbot fallbacks and
default sentinels, the LangChain store/tools defaults, the shared
recall-core.mjs (all synced plugin copies), the codex/claude-code/
openclaw/openwebui/dsh/zcode/pi plugin emitters, quick-app examples,
Go SDK example, tau2 benchmark targets, and the eval golden dataset.

Compat kept where legacy strings live in stored user configs: bot and
ov_dream sentinels accept both spellings while emitting only ~, and
recall-core still rewrites legacy viking://user/<reserved> config values
client-side. langchain_openviking._uri now classifies viking://~ with
the explicit-user shape so canonicalized server responses keep matching
a ~ root. Plugin READMEs note the server requirement for the alias.

* docs: replace current-user shorthand guidance with the viking://~ home alias

Rewrite every EN/ZH doc and model-facing prompt that advertised the
uid-less viking://user/<segment> spelling: URI concept catalogue,
context-types/storage/extraction/retrieval/session/privacy concepts,
configuration guide (with the legacy add_targets auto-normalization
note), resources/skills/sessions/retrieval/admin API references, FAQ,
capability reference, and the openviking-memory / ov-experience-memory /
openclaw / ov-resources skills. The stale MCP viking://user/<path>
dialect passage in the MCP guide is replaced by ~ guidance, and bare
viking://user is documented as the container of user spaces.

* test(api): migrate live API session-used tests off the removed shorthand

tests/api_test/sessions sent uid-less viking://user/skills/... URIs to
record_used, which the request boundary now rejects with 400 (caught by
the API & CLI Integration Tests CI job; these tests need a live server
and are not part of the local suites). The api_test client authenticates
as an admin-role user key, so the viking://~ home alias expands for it.
tests/api_test/common/test_edge_cases.py is left as is: it asserts a 400
for a non-resource add target, which still holds.
2026-08-21 19:00:19 +08:00
t0saki 33043cb1b8 feat(plugins): expose session commit trace IDs (#3977)
Preserve result.trace_id across plugin HTTP wrappers, include it in commit success and failure logs, and surface it in user-visible commit confirmations where supported.
2026-08-13 23:29:41 +08:00
2cc96e393e feat(retrieval): assemble auto-recall context server-side via /search mode="context" (#3534)
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"

Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.

This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.

New assembly kernel under openviking/retrieve/context_assembler/:

- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
  readable floor, then overview, then full for high-scoring entries. An
  oversized tier falls back to the previous one instead of being truncated,
  bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
  Summary section, code files reuse code_outline signatures, long documents use
  a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
  directories carry no stored abstract; their full tier stays capped at
  overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
  presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
  every harness inherits cross-turn dedup; exclude_uris remains as the
  stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
  element per entry. Every tier carries its URI, so the model can always drill
  down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
  timeout fuses, and a failed rewrite still returns the unrewritten block.
  Retrieval failures are counted into stats rather than silently yielding an
  empty block.

Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.

* refactor(retrieval): give context tiers a per-category default

The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.

Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.

Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".

Assembly fixes found alongside:

- Removing the abstract cap exemption would turn an oversized abstract into
  a bare URI, so it now falls back to overview first — for memory that is a
  cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
  `asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
  such attribute; usage was structurally always null. It now reads the model
  instance's tracker and reports only when the call count moved by exactly
  one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
  fail, and the file was never rewritten, so it could not heal. Records are
  now coerced on read and dropped on the next write, along with records left
  ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
  to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
  forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
  `viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
  bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
  started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
  resolved a different profile than `POST /recall`; an unknown `detail` value
  raised `KeyError` through the whole call instead of degrading.

* feat(codex): inject profile context on session start

Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(retrieval): raise rewrite timeout default to 30s

* docs(agents): document low-latency recall settings

* fix(codex): prefer luna as recall compressor fallback

* refactor(plugins): unify recall compression setting

* feat(plugins): enable recall compression by default

* docs(agents): use absolute links in image docs

* fix(retrieval): address context assembly review feedback

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* test: trim redundant context assembly coverage

* fix(retrieval): address second-round context assembly review

- Drop the backticked `/search` from the deprecated-recall row in both API
  overviews. The reference checker scans the whole row after the method cell
  for backticked paths, so it read the description as a route named
  `POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
  belongs to the Rust CLI, which writes `root_api_key`, `output`,
  `echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
  two Python readers had drifted into stricter subsets, so the shipped example
  already failed to load in both. Adding the new `plugin` section to a working
  ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
  Retrieval validates query and image_url before searching, and the gather
  fuse swallowed that rejection along with genuine scope failures, so a body
  of `{"mode":"context"}` came back 200 with an empty block instead of the
  documented parameter error. Runtime failures still degrade into
  `stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
  server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
  than the 30s fuse, so a rewrite that finished inside its own budget was
  aborted client-side, discarding the whole response — including the
  uncompressed block the server returns when a rewrite fails — and falling
  back to `/recall`. The deadline is only extended when the body actually
  requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
  `plugin.recallContextTimeoutMs` pins it.

* chore(plugins): sync shared modules into the zcode snapshot

* fix(retrieval): align context quotas and plugin defaults

Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(retrieval): preserve recall compatibility

Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-05 12:31:21 +08:00
t0saki 29ba81b53e feat(plugins): identify memory plugin traffic with a User-Agent (#3532)
Every harness memory plugin sent OpenViking requests under Node's default
`user-agent: node`, so server-side and gateway logs could not attribute
traffic to a harness or a plugin version.

Send `openviking-memory-<harness>/<version>` from all six harnesses
(claude-code, codex, opencode, pi, cursor, trae) on both the data plane
and the MCP proxy. Versions come from each harness's own manifest, or
from OPENVIKING_INTEGRATION_VERSION for cursor/trae; an unreadable
manifest degrades to 0.0.0 rather than throwing inside a short-lived hook.

The header is purely informational — neither the open-source server nor
the commercial gateway consumes it, and no auth or identity header
changes.
2026-07-27 15:31:21 +08:00
t0saki 2a81edc707 Add workspace peer mode for memory plugins (#3099) 2026-07-09 17:53:36 +08:00
agentandZaynJarvis 3757a143f0 [codex] fix memory plugin recall and auth handling (#2676)
* fix codex memory plugin backlog handling

* [codex] fix install.sh validators: optional peer + parenthesized plugin list state

Two false-positive validations surfaced when re-running the installer
on a no-peer config after #2598 merged:

1. cached .mcp.json validator unconditionally required the
   X-OpenViking-Actor-Peer mapping. Since syncMcpConfig now omits that
   header when no peer is configured (#2598 commit 216dbb5e), the
   validator failed every no-peer install. Make it conditional, and
   symmetrically flag a stale header when no peer is configured (the
   same way bearer_token_env_var handles HAS_API_KEY both ways).

2. 'codex plugin list' validator grep expected '<id>  installed,
   enabled' but codex CLI actually prints '<id> (installed, enabled)'
   with parens. Loosen the regex to accept both forms.

- ov-credentials.mjs gains a 'has-peer-id' CLI command symmetric to
  'has-api-key', so install.sh can detect peer presence via the same
  resolver used at render time. Underlying peerId resolution is
  already covered by ov-credentials.test.mjs cases 1, 2, and 3.
- install.sh detects HAS_PEER_ID, mirrors the bearer pattern at the
  validator, and makes the codex-plugin-list regex tolerate the
  parenthesized state form.

Verified by re-running setup-helper/install.sh non-interactively
against the no-peer ovcli.conf: 'Plugin install looks valid.' now
prints with no warnings. Hand-traced both HAS_PEER_ID branches of
the new validator (peer-configured / peer-absent, header-present
/ header-absent) to confirm the four cases behave correctly.

* fix codex backfill capture parsing

* refactor codex background transcript capture

* fix: state

* fix:bug

* fix: auto capture

* fix

* fix

* Revert experimental Codex capture changes

---------

Co-authored-by: ZaynJarvis <zaynjarvis@gmail.com>
2026-06-18 12:45:06 +08:00
Zayn Jarvis c6990e4cd6 [codex] tighten memory recall, capture, and resume (#2598)
* Fix Codex memory hook recall noise and stop timeouts

* Tighten Codex recall compression output

* Add Codex archive resume and capture filtering

* Wrap Codex memory injection for capture filtering

* Detect Codex recall compressor profile

* Refresh Codex compressor profile on startup

* fix codex ov credential resolution

* [codex] resolve compressor profile via models_cache.json, not codex exec probe

SessionStart used to spawn 'codex exec' sequentially against each
candidate model to detect which one would respond — up to 3 probes ×
~15s timeout each, on every session start, even on resume. The
configured-on-startup default made this a guaranteed first-page-load
tax of several seconds.

Replace the probe with a lookup against codex's own model catalogue
(~/.codex/models_cache.json, refreshed by codex CLI's etag-backed
fetch). The first candidate whose slug is present wins. SessionStart
now goes cache-first: load the persisted profile if any, only resolve
on cache miss. The runtime compress path (auto-recall) deletes the
cached profile on any compress failure so the next SessionStart
re-resolves against the current catalogue.

- recall-compressor-profile.mjs:
  * loadCodexModelsCache(env) reads ~/.codex/models_cache.json; missing
    cache yields {present:false,slugs:Set()}.
  * resolveRecallCompressorProfile picks the first available candidate
    by slug; falls back optimistically to the first candidate when
    the catalogue is missing.
  * invalidateRecallCompressorProfileCache() rms the persisted file.
  * detectRecallCompressorProfile is now cache-first and never spawns.
- auto-recall.mjs: runCodexCompressor invalidates the cache on spawn
  error, timeout, non-zero exit, and read failure (best-effort,
  no error surface to user).
- recall-compressor-profile.test.mjs: 11 unit tests covering catalogue
  read, candidate selection (with/without configured first), missing
  catalogue fallback, configured_off path, invalidate, cache-first
  detect, and re-resolve after invalidate.

Notes:
- buildCodexExecArgs is still exported so auto-recall can spawn the
  actual compress run; the change only removes the *probe* spawn, not
  the compress spawn.
- recallCompressDetectTtlMs and recallCompressDetectTimeoutMs are
  preserved in config for back-compat; the timeout no longer matters
  but the TTL still bounds how stale a cached profile may be.

* [codex] omit X-OpenViking-Actor-Peer env_http_headers when no peer configured

syncMcpConfig used to unconditionally write all three OV header→env
mappings. The wrapper strips empty OPENVIKING_PEER_ID before exec'ing
codex, so an unset env var would silently flip the header to "" — the
OV side then has to disambiguate that from "no peer scope". Match the
bearer_token_env_var pattern: present only when there's something to
send. Also drops a stale X-OpenViking-Actor-Peer entry when the peer
is unset (e.g. after switching ovcli configs).

- Existing test 4 became two cases: with-peer keeps the mapping,
  without-peer drops it (symmetric to bearer).
- New test asserts an in-place drop when the cached .mcp.json had a
  stale peer mapping but the active config no longer has a peer.

* [codex] runtime_failed compressor marker stops same-session retry storms

Previous fix invalidated the profile cache on compress failure. Within a
single codex session that still bled `recallCompressTimeoutMs` of wall
time per UserPromptSubmit because the next hook reread cache (miss),
fell back to fallbackRecallCompressorProfile, and tried the same model.

Replace plain invalidate with a runtime_failed sentinel cached in the
profile slot. UserPromptSubmit's compressMemoryContext already short-
circuits on `profile.enabled === false`, so the marker stops further
spawns for the rest of the codex process. The next SessionStart cache-
first detect treats `source === 'runtime_failed'` as cache miss and
re-resolves against the current models_cache.json, so a transient
failure self-recovers across codex restarts without operator action.
detect_on_startup=false respects the marker (no auto-recover, matches
the "manual control" intent of that flag).

- recall-compressor-profile.mjs:
  * markRecallCompressorRuntimeFailed(cfg, {failedModel}) writes the
    disabled sentinel.
  * detectRecallCompressorProfile branches on cached.source ===
    'runtime_failed': cache hit otherwise, recover-via-resolve when
    startup-detect on, respect marker when off.
- auto-recall.mjs::runCodexCompressor: swap invalidate-on-error with
  markRecallCompressorRuntimeFailed(cfg, {failedModel: profile.model}).
- recall-compressor-profile.test.mjs: 4 new tests covering marker
  write, cross-restart recovery picking a different slug, and the
  detect_on_startup=false honor path. 20/20 pass.

invalidateRecallCompressorProfileCache is kept as a public API for
explicit operator use (e.g. a future `ov codex reset-compressor`
command), but is no longer called from the runtime path.
2026-06-16 02:05:18 +08:00
Qin Haojie ff258768c2 feat(memory): 引入 User/Peer 记忆隔离模型 (#2236)
* feat(memory): introduce user and peer memory isolation

Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.

* feat(memory): align session identity around peer IDs

* feat(search): pass peer id through retrieval

* refactor(memory): remove agent identity from integrations

* fix(memory): isolate peer identity from self extraction

* fix(tau2): provision benchmark user configs

* fix(auth): allow admin keys to access data APIs

* fix(openclaw): enable peer memory policy for peer roles

* fix(openclaw): resolve sender for peer recall

* refactor(session): simplify memory extraction routing

* refactor(ov-cli): reduce formatting-only diff

* refactor(message): remove unused message helpers

* refactor(retrieval): simplify peer target resolution

* refactor(namespace): remove deprecated agent namespace policy

* fix(agent): propagate peer id through integrations

* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
Zayn Jarvis a7d27920cd fix(plugin/codex): simplify commit hook messages (#2036) 2026-05-14 13:04:01 +08:00
e92180a7e1 feat(plugin/codex): add lifecycle hooks (recall, capture, pre-compact) to codex-memory-plugin (#1957)
* feat(plugin/codex): add lifecycle hooks (recall, capture, pre-compact)

Brings the codex-memory-plugin to feature parity with the claude-code-memory-plugin
by wiring the four Codex lifecycle hooks via `hooks.json`:

- SessionStart  -> bootstrap-runtime.mjs (npm ci into ${CODEX_PLUGIN_DATA}/runtime)
- UserPromptSubmit -> auto-recall.mjs (search OV, inject via hookSpecificOutput.additionalContext)
- Stop -> auto-capture.mjs (incremental transcript capture + last_assistant_message commit)
- PreCompact -> pre-compact-capture.mjs (full transcript -> single OV session -> commit)

Differences from the Claude Code plugin baked into the scripts:

- Codex output schema does not allow `decision: "approve"`; no-op is `{}`
- Stop/PreCompact only support `systemMessage`, not `additionalContext`
- Plugin envs are CODEX_PLUGIN_ROOT / CODEX_PLUGIN_DATA
- Config section is `codex` (was `claude_code`); config file defaults to
  `~/.openviking/ovcli.conf`, falling back to legacy `~/.openviking/ov.conf`

Other changes:

- src/memory-server.ts now reads ovcli.conf-style configs (top-level `url`,
  `api_key`, `account`, `user`, `agent_id`) so the plugin works against
  hosted OpenViking deployments out of the box. Env-var-only operation
  (OPENVIKING_URL set, no config file) is also supported.
- .mcp.json points at scripts/start-memory-server.mjs, which boots the same
  runtime the hooks use, so the MCP path benefits from npm-ci bootstrap.
- README rewritten with architecture diagram, validation SOP, configuration
  reference, and a Codex-vs-Claude-Code differences table.

Validated end-to-end against an OpenViking deployment:

- Auto-recall returns ranked memories with full content and emits
  hookSpecificOutput.additionalContext.
- Auto-capture (last_assistant_message path) creates a session, commits, and
  the OV pipeline extracts events + preferences within ~60s.
- Pre-compact-capture posts a full 4-turn transcript to one OV session,
  commits with archived=true, and produces structured leaf memories
  (preferences, events, entities) under viking://user/<user>/memories/.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(plugin/codex): drop SessionStart, split Stop=add_message vs PreCompact=commit

Codex's `Stop` hook fires per turn, not at session end, so committing per-Stop
over-fragments memory extraction. And codex re-fires `SessionStart` on short
reconnects, so registering an `npm ci` bootstrap there reinstalls the runtime
unnecessarily.

This change keeps one long-lived OpenViking session per codex `session_id`
across all `Stop` invocations, and only triggers the OV memory extractor on
`PreCompact` (or via an idle-sweep best-effort commit when codex exits without
compacting).

- hooks.json: drop SessionStart entry; keep UserPromptSubmit/Stop/PreCompact
- scripts/session-state.mjs (new): per-codex-session state under
  ~/.openviking/codex-plugin-state/, tracks ovSessionId + capturedTurnCount
- scripts/auto-capture.mjs (Stop): incremental add_message only, idle-sweep at
  the tail to commit stale codex sessions (default IDLE_TTL=30 min, override
  with OPENVIKING_CODEX_IDLE_TTL_MS)
- scripts/pre-compact-capture.mjs (PreCompact): catch-up append + commit the
  long-lived OV session, then null out ovSessionId so the next Stop opens a
  fresh OV session for the post-compact half
- MCP runtime install stays lazy in start-memory-server.mjs (already there);
  no SessionStart hook means short reconnects don't re-trigger npm ci
- VERIFICATION.md: end-to-end SOP against a live OV server (~3 min)
- bump plugin to 0.3.0

Verified end-to-end against ov.zaynjarvis.com:
  Stop adds turns idempotently and incrementally; PreCompact commits to
  history/archive_001/ with extractor producing memories under
  viking://user/<user>/memories/profile.md after ~30 s; post-compact Stop
  opens a fresh OV session; idle-sweep commits stale state files.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(plugin/codex): replace idle-sweep with SessionStart(source=clear) commit

Per Zayn's followup ("非必要不要加 idle commit"): drop the idle-sweep added
in the previous commit and use codex's actual context-disappearing signal —
SessionStart with source=clear — to commit orphaned sessions.

Codex hook signal map:
- /compact         → PreCompact      ✅ commit (already)
- /clear           → SessionStart(source=clear) for the NEW session_id;
                     the prior transcript is orphaned. Now committed.
- /new             → SessionStart(source=startup); ambiguous with fresh
                     codex startup, so we don't act on it.
- /resume / short reconnect → SessionStart(source=resume|startup); no-op
                     to avoid corrupting still-active sessions.
- SIGTERM/Ctrl+C/exit → no hook fires. Documented as a known gap; users
                     should /compact before /exit if they want commit.

Changes:
- new scripts/session-start-commit.mjs: gates internally on source=clear,
  iterates listStates(), and commits any state file whose codexSessionId
  != the new SessionStart session_id, then clears that state file
- hooks/hooks.json: re-register SessionStart pointing at the new script
  (timeout 30s)
- scripts/auto-capture.mjs: remove sweepIdleSessions() and
  IDLE_TTL_MS env handling; Stop is now strictly add_message
- README/VERIFICATION.md: update arch diagram, replace idle-sweep step
  with SessionStart(source=clear) verify (positive + negative paths),
  add "Known gap: SIGTERM/exit are silent" section
- bump to 0.3.1

Verified end-to-end against ov.zaynjarvis.com:
  Stop add+idempotent ✓
  SessionStart source=startup → {} ✓
  SessionStart source=resume → {} ✓
  SessionStart source=clear → committed prior OV session, history/archive_001/
  appeared, profile.md gained "Favorite snack: dark chocolate" within 30 s.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(plugin/codex): SessionStart matcher = "clear" (native dispatcher gate)

Codex's hooks dispatcher matches the SessionStart hook's `matcher` field
against the SessionStart `source` value. Setting matcher to "clear" means
codex won't even spawn our script on `source=startup` or `source=resume`
(short reconnects); we previously gated this in-script. The internal
source check in session-start-commit.mjs is kept as defense-in-depth.

Source: codex-rs/hooks/src/events/session_start.rs `select_handlers(...,
matcher_input: Some(request.source.as_str()))` and
codex-rs/hooks/src/events/common.rs `is_exact_matcher` — "clear" is
all-alphanumeric so it's matched as exact equality, not regex.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(plugin/codex): active-window heuristic + idle-TTL sweep at SessionStart (v0.4.0)

Source of truth: examples/codex-memory-plugin/DESIGN.md (added in this commit).

Behavioral changes:

- SessionStart matcher widens from `clear` to `clear|startup`. Both sources
  run the same active-window heuristic; `resume` is a hard no-op (still fires
  on short reconnects).
- Heuristic (DESIGN.md §3): count state files (excluding new session_id) within
  ACTIVE_WINDOW_MS (default 2 min). 0 → noop, 1 → commit it (just-ended
  session), ≥2 → skip and rely on idle TTL. Tunable via
  OPENVIKING_CODEX_ACTIVE_WINDOW_MS.
- Idle-TTL sweep returns at the tail of session-start-commit.mjs only (not
  every Stop). Default IDLE_TTL_MS = 30 min via OPENVIKING_CODEX_IDLE_TTL_MS.
  Catches SIGTERM/Ctrl+C/`/exit` orphans and the ≥2-active skip path.
- Stop hook deliberately does NOT sweep — state-write-on-every-turn already
  gives us the freshness signal. Marker comment added.
- Stop hook adds post-compact transcript-shrink defense: if
  allTurns.length < state.capturedTurnCount, reset capturedTurnCount = 0.
- Commit-on-failure preserves state everywhere (PreCompact, heuristic,
  idle sweep). A non-2xx /commit no longer clears ovSessionId; the next
  sweep retries.
- session-state.mjs saveState now uses atomic write (tmpfile + rename) for
  crash safety. listStates ignores the brief `<id>.json.tmp` window.

Bump: package.json + .codex-plugin/plugin.json → 0.4.0.

Docs: README "How It Works" gained a DESIGN.md pointer and rewrites the
SessionStart section to reflect heuristic + idle TTL. VERIFICATION.md step 6
now exercises all four heuristic branches (0/1/≥2 active, idle TTL, resume).

Phase-2 resume context inject documented in DESIGN.md but explicitly out of
scope here.

Verified locally with synthetic stdin tests against a fake OV server:
1-active commit, ≥2-active skip, idle TTL sweep, resume noop,
unreachable-server keeps state.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(plugin/codex): align config loading with claude-code plugin

Addresses three review points on PR #1957:

1. Honor OPENVIKING_CLI_CONFIG_FILE for the ovcli.conf override path
   (matches the convention used by `ov` CLI and claude-code-memory-plugin).
   OPENVIKING_CONFIG_FILE stays as the ov.conf override; for backward
   compat it still works when pointed at an ovcli-shaped file.

2. Strict env-first priority for every connection / identity field
   (baseUrl, apiKey, account, user, agentId). Env vars now win over
   ovcli.conf, which wins over ov.conf's codex.* block / server.*,
   which wins over built-in defaults.

3. Unify hook and MCP-server config loading: src/memory-server.ts now
   imports loadConfig from scripts/config.mjs (relative path stays
   valid post-compile because servers/ and scripts/ are siblings),
   eliminating the divergent account/user/agentId fallback chains
   the PR-Agent reviewer flagged.

Auth header: emit Authorization: Bearer (primary, required by OpenViking
Cloud) plus the legacy X-API-Key during the transition window. All six
fetch sites updated (4 hook scripts + memory-server.ts + compiled
servers/memory-server.js).

README: document the new resolution chain, OPENVIKING_CLI_CONFIG_FILE,
OPENVIKING_BEARER_TOKEN alias, and the Authorization: Bearer migration.

* docs(plugin/codex): put installation first

* fix(plugin/codex): harden runtime and capture paths

* docs(plugin/codex): align local marketplace name

* docs(plugin/codex): add one-line installer

* fix(plugin/codex): support branch installer testing

* fix(plugin/codex): keep installer env surface stable

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: zhengxiao.wu <zhengxiao.wu@bytedance.com>
2026-05-13 19:25:30 +08:00