* feat(retrieval): assemble auto-recall context server-side via /search mode="context"
Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.
This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.
New assembly kernel under openviking/retrieve/context_assembler/:
- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
readable floor, then overview, then full for high-scoring entries. An
oversized tier falls back to the previous one instead of being truncated,
bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
Summary section, code files reuse code_outline signatures, long documents use
a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
directories carry no stored abstract; their full tier stays capped at
overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
every harness inherits cross-turn dedup; exclude_uris remains as the
stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
element per entry. Every tier carries its URI, so the model can always drill
down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
timeout fuses, and a failed rewrite still returns the unrewritten block.
Retrieval failures are counted into stats rather than silently yielding an
empty block.
Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.
* refactor(retrieval): give context tiers a per-category default
The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.
Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.
Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".
Assembly fixes found alongside:
- Removing the abstract cap exemption would turn an oversized abstract into
a bare URI, so it now falls back to overview first — for memory that is a
cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
`asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
such attribute; usage was structurally always null. It now reads the model
instance's tracker and reports only when the call count moved by exactly
one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
fail, and the file was never rewritten, so it could not heal. Records are
now coerced on read and dropped on the next write, along with records left
ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
`viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
resolved a different profile than `POST /recall`; an unknown `detail` value
raised `KeyError` through the whole call instead of degrading.
* feat(codex): inject profile context on session start
Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): raise rewrite timeout default to 30s
* docs(agents): document low-latency recall settings
* fix(codex): prefer luna as recall compressor fallback
* refactor(plugins): unify recall compression setting
* feat(plugins): enable recall compression by default
* docs(agents): use absolute links in image docs
* fix(retrieval): address context assembly review feedback
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test: trim redundant context assembly coverage
* fix(retrieval): address second-round context assembly review
- Drop the backticked `/search` from the deprecated-recall row in both API
overviews. The reference checker scans the whole row after the method cell
for backticked paths, so it read the description as a route named
`POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
belongs to the Rust CLI, which writes `root_api_key`, `output`,
`echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
two Python readers had drifted into stricter subsets, so the shipped example
already failed to load in both. Adding the new `plugin` section to a working
ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
Retrieval validates query and image_url before searching, and the gather
fuse swallowed that rejection along with genuine scope failures, so a body
of `{"mode":"context"}` came back 200 with an empty block instead of the
documented parameter error. Runtime failures still degrade into
`stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
than the 30s fuse, so a rewrite that finished inside its own budget was
aborted client-side, discarding the whole response — including the
uncompressed block the server returns when a rewrite fails — and falling
back to `/recall`. The deadline is only extended when the body actually
requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
`plugin.recallContextTimeoutMs` pins it.
* chore(plugins): sync shared modules into the zcode snapshot
* fix(retrieval): align context quotas and plugin defaults
Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): preserve recall compatibility
Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
8.4 KiB
Install the Unified OpenViking OpenCode Plugin
This plugin adds one unified OpenViking plugin for OpenCode:
- OpenViking MCP tools for memory, resources, and code context
- Long-term memory, session synchronization, lifecycle commit, and automatic recall
This is the only OpenCode plugin example maintained in this repository. It does not install skills/openviking/SKILL.md, and it does not require the agent to use the ov command. Model tools are provided by the same stdio MCP proxy used by the Claude Code and Codex memory plugins.
Prerequisites
Prepare the following first:
- OpenCode
- OpenViking HTTP Server
- Node.js 18+
- A valid OpenViking API key if authentication is enabled on the server
Start OpenViking first:
openviking-server --config ~/.openviking/ov.conf
Check the service:
curl http://localhost:1933/health
Installation Method 1: Published Package
Normal users are recommended to enable it through OpenCode's package plugin mechanism:
{
"plugin": ["@openviking/opencode-plugin"]
}
Installation Method 2: Source Install
Use this method for development, debugging, or PR testing. OpenCode's recommended plugin directory is:
~/.config/opencode/plugins
Run the following commands from the repository root:
mkdir -p ~/.config/opencode/plugins/openviking
cp examples/opencode-plugin/wrappers/openviking.js ~/.config/opencode/plugins/openviking.js
cp examples/opencode-plugin/index.mjs examples/opencode-plugin/package.json ~/.config/opencode/plugins/openviking/
cp -r examples/opencode-plugin/lib ~/.config/opencode/plugins/openviking/
cp -r examples/opencode-plugin/servers ~/.config/opencode/plugins/openviking/
After installation, the layout should look like this:
~/.config/opencode/plugins/
├── openviking.js
└── openviking/
├── index.mjs
├── package.json
├── lib/
└── servers/
The top-level openviking.js forwards the first-level .js entry that OpenCode can discover to the actual plugin directory:
export { OpenVikingPlugin, default } from "./openviking/index.mjs"
This wrapper is only for source installs with the directory layout shown above. npm package installs load index.mjs directly through package.json.
Use the .js wrapper for source installs; OpenCode's local plugin scanner discovers JavaScript/TypeScript plugin files.
If you install through an npm package, you can also use examples/opencode-plugin as a normal OpenCode plugin package.
Configuration
Create the user-level configuration file:
~/.config/opencode/openviking-config.json
Example configuration:
{
"enabled": true,
"timeoutMs": 30000,
"repoContext": { "enabled": true, "cacheTtlMs": 60000 },
"autoRecall": {
"enabled": true,
"limit": 6,
"scoreThreshold": 0.35,
"maxContentChars": 500,
"preferAbstract": true,
"tokenBudget": 2000,
"minQueryLength": 3
},
"commitTokenThreshold": 20000,
"commitKeepRecentCount": 10,
"profileTokenBudget": 10000,
"resumeContextBudget": 32000
}
autoRecall.limit is a legacy quota-scaling input, not a final result cap.
Explicit values from 1 through 5 produce an effective total quota of 6 because
each coding category keeps one retrieval slot.
It is recommended to provide the API key through an environment variable instead of writing it into the configuration file:
export OPENVIKING_API_KEY="your-api-key-here"
API keys are resolved from environment variables or ~/.openviking/ovcli.conf and sent as Authorization: Bearer ... by both hooks and the MCP proxy. account and user are trusted-mode identity headers sent as X-OpenViking-Account and X-OpenViking-User; leave them empty when using API-key mode with user/admin API keys. peerId is sent as X-OpenViking-Actor-Peer on data-plane memory/resource requests; captured session messages store it as body peer_id.
OPENVIKING_API_KEY, OPENVIKING_ACCOUNT, OPENVIKING_USER, and OPENVIKING_PEER_ID take precedence over the corresponding values in openviking-config.json.
For advanced setups, use OPENVIKING_PLUGIN_CONFIG to point to another configuration file path.
Verify
Restart OpenCode after changing plugin or OpenViking configuration.
In a new OpenCode session, ask the agent to browse OpenViking memory or search for a known indexed resource. The plugin should expose the OpenViking MCP server, with tools namespaced by OpenCode as openviking_*:
openviking_recall,openviking_search,openviking_findopenviking_read,openviking_list,openviking_grep,openviking_globopenviking_remember,openviking_add_resource,openviking_forget,openviking_healthopenviking_list_watches,openviking_cancel_watch
If anything looks wrong, check the runtime files:
ls ~/.config/opencode/openviking/
tail -n 100 ~/.config/opencode/openviking/openviking-memory.log
For a local server, also confirm OpenViking is reachable:
curl http://localhost:1933/health
Available MCP Tools
The plugin registers OpenViking's stdio MCP proxy through OpenCode config. The server's real tools/list response is the source of truth; current OpenViking servers expose:
openviking_recall: balanced current-task recall.openviking_search: deep semantic retrieval across memories, resources, and skills.openviking_find: fast semantic retrieval.openviking_remember: store important facts or decisions for memory extraction.openviking_read: read one or moreviking://files.openviking_list: list aviking://directory.openviking_grep: exact text or regex search.openviking_glob: glob file matching.openviking_add_resource: add a URL, local file, sitemap, or feed.openviking_forget: delete aviking://URI after explicit user confirmation.openviking_list_watches/openviking_cancel_watch: inspect or cancel resource watches.openviking_health: check OpenViking server health.
Usage guidance:
- Use
openviking_searchfor conceptual questions. - Use
openviking_grepfor exact symbols, function names, class names, or error strings. - Use
openviking_globto enumerate files. - Use
openviking_readto read content. - Use
openviking_listto explore directory structure. - Before deleting anything, obtain explicit user confirmation first; then call
openviking_forget. - If an agent tries to use OpenCode's local
read,glob, orgreptools on aviking://URI, the plugin blocks that call and points it to the MCP tools.
Local Files with openviking_add_resource
openviking_add_resource supports three input types:
- Remote
http(s)URL: directly calls/api/v1/resources - Local file path: first calls
/api/v1/resources/temp_upload, then adds the resource using the returnedtemp_file_id file://URL: handled as a local file
Relative paths are resolved against the current OpenCode project directory. Examples:
openviking_add_resource(path="https://example.com/spec.md", to="viking://resources/spec")
openviking_add_resource(path="./docs/notes.md", to="viking://resources/notes.md")
openviking_add_resource(path="file:///home/alice/project/notes.md", description="project notes")
Automatic zip upload for local directories is not supported yet. Passing a directory will return a clear error.
Runtime Files
By default, the plugin writes runtime files to:
~/.config/opencode/openviking/
Possible files include:
openviking-memory.logopenviking-session-state.json
You can change this directory with runtime.dataDir in the configuration.
These are local runtime files and should not be committed to the repository.
Troubleshooting
| Issue | What to check |
|---|---|
| Plugin does not load | For package installs, confirm ~/.config/opencode/opencode.json contains @openviking/opencode-plugin; for source installs, confirm ~/.config/opencode/plugins/openviking.js exists |
| MCP tools call the wrong server | Check ~/.openviking/ovcli.conf, or set OPENVIKING_* env vars / OPENVIKING_PLUGIN_CONFIG to the intended config path |
| 401 / 403 from OpenViking | Verify OPENVIKING_API_KEY; for trusted-mode deployments, also verify OPENVIKING_ACCOUNT and OPENVIKING_USER |
| Recall is empty | Confirm OpenViking has indexed memories/resources and autoRecall.enabled is true |
Local openviking_add_resource fails |
Pass a file path, not a directory; local directories are not uploaded automatically yet |