Files
OpenViking/examples/opencode-plugin/INSTALL-ZH.md
T
2cc96e393e feat(retrieval): assemble auto-recall context server-side via /search mode="context" (#3534)
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"

Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.

This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.

New assembly kernel under openviking/retrieve/context_assembler/:

- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
  readable floor, then overview, then full for high-scoring entries. An
  oversized tier falls back to the previous one instead of being truncated,
  bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
  Summary section, code files reuse code_outline signatures, long documents use
  a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
  directories carry no stored abstract; their full tier stays capped at
  overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
  presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
  every harness inherits cross-turn dedup; exclude_uris remains as the
  stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
  element per entry. Every tier carries its URI, so the model can always drill
  down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
  timeout fuses, and a failed rewrite still returns the unrewritten block.
  Retrieval failures are counted into stats rather than silently yielding an
  empty block.

Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.

* refactor(retrieval): give context tiers a per-category default

The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.

Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.

Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".

Assembly fixes found alongside:

- Removing the abstract cap exemption would turn an oversized abstract into
  a bare URI, so it now falls back to overview first — for memory that is a
  cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
  `asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
  such attribute; usage was structurally always null. It now reads the model
  instance's tracker and reports only when the call count moved by exactly
  one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
  fail, and the file was never rewritten, so it could not heal. Records are
  now coerced on read and dropped on the next write, along with records left
  ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
  to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
  forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
  `viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
  bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
  started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
  resolved a different profile than `POST /recall`; an unknown `detail` value
  raised `KeyError` through the whole call instead of degrading.

* feat(codex): inject profile context on session start

Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(retrieval): raise rewrite timeout default to 30s

* docs(agents): document low-latency recall settings

* fix(codex): prefer luna as recall compressor fallback

* refactor(plugins): unify recall compression setting

* feat(plugins): enable recall compression by default

* docs(agents): use absolute links in image docs

* fix(retrieval): address context assembly review feedback

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* test: trim redundant context assembly coverage

* fix(retrieval): address second-round context assembly review

- Drop the backticked `/search` from the deprecated-recall row in both API
  overviews. The reference checker scans the whole row after the method cell
  for backticked paths, so it read the description as a route named
  `POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
  belongs to the Rust CLI, which writes `root_api_key`, `output`,
  `echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
  two Python readers had drifted into stricter subsets, so the shipped example
  already failed to load in both. Adding the new `plugin` section to a working
  ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
  Retrieval validates query and image_url before searching, and the gather
  fuse swallowed that rejection along with genuine scope failures, so a body
  of `{"mode":"context"}` came back 200 with an empty block instead of the
  documented parameter error. Runtime failures still degrade into
  `stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
  server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
  than the 30s fuse, so a rewrite that finished inside its own budget was
  aborted client-side, discarding the whole response — including the
  uncompressed block the server returns when a rewrite fails — and falling
  back to `/recall`. The deadline is only extended when the body actually
  requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
  `plugin.recallContextTimeoutMs` pins it.

* chore(plugins): sync shared modules into the zcode snapshot

* fix(retrieval): align context quotas and plugin defaults

Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(retrieval): preserve recall compatibility

Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-05 12:31:21 +08:00

7.9 KiB
Raw Permalink Blame History

安装 OpenViking OpenCode 统一插件

这个插件新增了一个面向 OpenCode 的统一 OpenViking 插件:

  • 面向 memory、resources 和 code context 的 OpenViking MCP 工具
  • 长期记忆、session 同步、生命周期边界 commit、自动 recall

这是仓库中唯一继续维护的 OpenCode 插件示例。这个插件不再安装 skills/openviking/SKILL.md,也不要求 agent 使用 ov 命令。模型工具由 Claude Code 和 Codex 记忆插件同款的 stdio MCP proxy 提供。

前置条件

需要先准备:

  • OpenCode
  • OpenViking HTTP Server
  • Node.js 18+
  • 如果服务端启用了认证,需要可用的 OpenViking API Key

建议先启动 OpenViking:

openviking-server --config ~/.openviking/ov.conf

检查服务:

curl http://localhost:1933/health

安装方式一:发布包安装

普通用户推荐通过 OpenCode 的 package plugin 机制启用:

{
  "plugin": ["@openviking/opencode-plugin"]
}

安装方式二:源码安装

用于开发调试或 PR 测试。OpenCode 推荐插件目录:

~/.config/opencode/plugins

在仓库根目录执行:

mkdir -p ~/.config/opencode/plugins/openviking
cp examples/opencode-plugin/wrappers/openviking.js ~/.config/opencode/plugins/openviking.js
cp examples/opencode-plugin/index.mjs examples/opencode-plugin/package.json ~/.config/opencode/plugins/openviking/
cp -r examples/opencode-plugin/lib ~/.config/opencode/plugins/openviking/
cp -r examples/opencode-plugin/servers ~/.config/opencode/plugins/openviking/

安装后结构应类似:

~/.config/opencode/plugins/
├── openviking.js
└── openviking/
    ├── index.mjs
    ├── package.json
    ├── lib/
    └── servers/

顶层 openviking.js 只负责把 OpenCode 能发现的一级 .js 入口转发到插件目录:

export { OpenVikingPlugin, default } from "./openviking/index.mjs"

这个 wrapper 只用于上面这种源码安装目录结构。npm 包安装会通过 package.json 直接加载 index.mjs。 源码安装请使用 .js wrapper;OpenCode 的本地插件扫描器会发现 JavaScript/TypeScript 插件文件。

如果你使用 npm 包方式安装,也可以将 examples/opencode-plugin 作为一个普通 OpenCode 插件包使用。

配置

创建用户级配置文件:

~/.config/opencode/openviking-config.json

示例配置:

{
  "enabled": true,
  "timeoutMs": 30000,
  "repoContext": { "enabled": true, "cacheTtlMs": 60000 },
  "autoRecall": {
    "enabled": true,
    "limit": 6,
    "scoreThreshold": 0.35,
    "maxContentChars": 500,
    "preferAbstract": true,
    "tokenBudget": 2000,
    "minQueryLength": 3
  },
  "commitTokenThreshold": 20000,
  "commitKeepRecentCount": 10,
  "profileTokenBudget": 10000,
  "resumeContextBudget": 32000
}

autoRecall.limit 是遗留的配额缩放输入,不是最终结果上限。显式设置为 1 到 5 时,有效总配额仍为 6,因为六个 coding 分类会各保留一个检索槽位。

推荐通过环境变量提供 API Key,而不是写入配置文件:

export OPENVIKING_API_KEY="your-api-key-here"

API key 会从环境变量或 ~/.openviking/ovcli.conf 读取,并由 hooks 和 MCP proxy 作为 Authorization: Bearer ... 发送。account 和 user 是 trusted mode 身份头,会作为 X-OpenViking-Account、X-OpenViking-User 发送;使用 user/admin API key 的 API_KEY mode 时应留空。 peerId 会作为 X-OpenViking-Actor-Peer 用于数据面的 memory/resource 请求;捕获 session message 时仍写入 body peer_id。需要 peer 维度路由时请显式配置。

OPENVIKING_API_KEY、OPENVIKING_ACCOUNT、OPENVIKING_USER、 OPENVIKING_PEER_ID 优先级高于 openviking-config.json 里的同名配置。

高级场景可以用 OPENVIKING_PLUGIN_CONFIG 指向其他配置文件路径。

验证

修改插件或 OpenViking 配置后,需要重启 OpenCode。

进入新的 OpenCode session 后,可以让 agent 浏览 OpenViking memory,或搜索一个已索引的资源。插件应暴露 OpenViking MCP server,OpenCode 中的工具名会带 openviking_ 前缀:

  • openviking_recall、openviking_search、openviking_find
  • openviking_read、openviking_list、openviking_grep、openviking_glob
  • openviking_remember、openviking_add_resource、openviking_forget、openviking_health
  • openviking_list_watches、openviking_cancel_watch

如果行为异常,先查看运行时文件:

ls ~/.config/opencode/openviking/
tail -n 100 ~/.config/opencode/openviking/openviking-memory.log

如果使用本地 server,也确认 OpenViking 可访问:

curl http://localhost:1933/health

可用 MCP 工具

插件会通过 OpenCode config 注册 OpenViking stdio MCP proxy。服务端实际返回的 tools/list 是最终工具清单;当前 OpenViking server 暴露:

  • openviking_recall:面向当前任务的平衡召回
  • openviking_search:跨 memories/resources/skills 的深度语义检索
  • openviking_find:快速语义检索
  • openviking_remember:存储重要事实或决策,供记忆提取
  • openviking_read:读取一个或多个 viking:// 文件
  • openviking_list:列出 viking:// 目录
  • openviking_grep:精确文本或正则搜索
  • openviking_glob:glob 文件匹配
  • openviking_add_resource:添加 URL、本地文件、sitemap 或 feed
  • openviking_forget:在用户明确确认后删除 viking:// URI
  • openviking_list_watches / openviking_cancel_watch:查看或取消资源 watch
  • openviking_health:检查 OpenViking server 健康状态

使用建议:

  • 概念性问题用 openviking_search
  • 精确符号、函数名、类名、报错字符串用 openviking_grep
  • 枚举文件用 openviking_glob
  • 读取内容用 openviking_read
  • 探索目录结构用 openviking_list
  • 删除前必须先获得用户明确确认,再调用 openviking_forget
  • 如果 agent 误用 OpenCode 本地 read、glob、grep 工具访问 viking:// URI,插件会阻止这次本地文件系统调用,并提示改用 MCP 工具。

openviking_add_resource 本地文件

openviking_add_resource 支持三类输入:

  • 远端 http(s) URL:直接调用 /api/v1/resources
  • 本地文件路径:先调用 /api/v1/resources/temp_upload,再用返回的 temp_file_id 添加资源
  • file:// URL:按本地文件处理

相对路径会按 OpenCode 当前项目目录解析。示例:

openviking_add_resource(path="https://example.com/spec.md", to="viking://resources/spec")
openviking_add_resource(path="./docs/notes.md", to="viking://resources/notes.md")
openviking_add_resource(path="file:///home/alice/project/notes.md", description="project notes")

当前仍不支持本地目录自动打 zip 上传;传入目录时会返回明确错误。

运行时文件

插件默认会把运行时文件写入:

~/.config/opencode/openviking/

可能包含:

  • openviking-memory.log
  • openviking-session-state.json

可以通过配置里的 runtime.dataDir 修改这个目录。

这些是本地运行时文件,不建议提交到版本库。

故障排查

问题 排查方向
插件没有加载 package 安装检查 ~/.config/opencode/opencode.json 是否包含 @openviking/opencode-plugin;源码安装检查 ~/.config/opencode/plugins/openviking.js 是否存在
MCP tools 连到了错误的 server 检查 ~/.openviking/ovcli.conf,或用 OPENVIKING_* 环境变量 / OPENVIKING_PLUGIN_CONFIG 指向正确配置
OpenViking 返回 401 / 403 检查 OPENVIKING_API_KEY;trusted-mode 部署还要检查 OPENVIKING_ACCOUNT 和 OPENVIKING_USER
recall 为空 确认 OpenViking 中已有 memories/resources,并且 autoRecall.enabled 为 true
本地 openviking_add_resource 失败 传入文件路径而不是目录;目前还不支持自动上传本地目录