* feat(resources): support source headers for HTTP imports
Allow HTTP(S) resource imports to pass request headers for authenticated object storage sources. Fetch authenticated sources into a request-scoped snapshot and keep credentials out of parser inputs and queued jobs.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* fix(resources): reserve source headers from args
Treat source_headers as a top-level add_resource field so it is rejected from args and filtered consistently from queued processor inputs.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat(resources): accept TOS auth through args
Replace generic source header passthrough with explicit tos_signature and tos_access args for HTTP TOS object imports. Keep credentials request-scoped and out of parser and queue payloads.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* refactor(resources): name transient TOS args
Centralize TOS credentials that are accepted from args but excluded from durable resource jobs.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Allow openai-codex VLM configuration to carry reasoning_effort through credential normalization and into Responses API requests so deployments can tune model effort without out-of-tree patches.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(sdk): sync go/ts/python SDKs with server find/search, recall, and admin changes
Server-side changes recently landed that the language SDKs had drifted from:
- find/search results now return `tags` and no longer return
`category`/`match_reason`/`relations`/`overview` (#3730). Go's strict
struct was the only one broken; update MatchedContext accordingly.
- new admin endpoints for agent-evolution and per-account settings (#3695).
- public `search/recall` endpoint was missing from all SDKs.
Changes:
- python: add `level`/`since`/`until`/`time_field` to find/search; add an
`extra` escape hatch to find/search/add_resource/write/batch_write so new
server fields can be passed without an SDK bump (only forwarded when set,
preserving `level=0`); add `recall` and the four admin methods.
- go: fix MatchedContext (add Tags, drop removed fields), add Recall and the
four admin methods.
- typescript: type MatchedContext/FindResult, add RecallOptions, add `recall`
and the four admin methods.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(sdk): unify options APIs and sync latest server interfaces
- migrate complex Python SDK calls to typed options dictionaries
- add dedicated context search and consistent extra-field handling
- align Go and TypeScript options with omission-aware serialization
- support session config, event tags, Agent Evolution date filters,
OpenViking Assets, batch write, downloads, and create_parent
- refresh SDK tests and examples across all three languages
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): address options API review findings
- fix Go session extra merging and Python message precedence
- adapt LangChain calls to the Python options API
- migrate repository examples, tests, and documentation
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): complete options migration and message parity
- migrate remaining Python SDK benchmarks to options dictionaries
- normalize empty parts consistently for single and batch messages
- add regression guards for repository SDK call sites
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): align reindex options after main rebase
- preserve reindex tags in Python typed options
- add reindex extra support for Go and TypeScript
- reject official fields passed through extra across SDKs
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(sdk): support legacy keyword options
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
docs(sdk): use explicit Python SDK arguments
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): support set tags extra options
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): expose Go add resource options
Expose AddType and ProcessingMode through Go AddResourceOptions and serialize them to the resources API. Add a regression test covering the resulting request payload.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(sdk): flatten core Python client options
Co-authored-by: TRAE CLI <traecli@bytedance.com>
docs(sdk): align Python examples with flattened options
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): preserve core API compatibility
Co-authored-by: TRAE CLI <traecli@bytedance.com>
refactor(python-sdk): move resource hints to options
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): align resource option callers
Co-authored-by: TRAE CLI <traecli@bytedance.com>
test(sdk): cover recursive reindex forwarding
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): preserve Go options compatibility
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(python-sdk): expose message peer id
Co-authored-by: TRAE CLI <traecli@bytedance.com>
test(python-sdk): consolidate options coverage
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(python-sdk): add parts and flatten image search
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* docs(sdk): align Python call examples
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
* feat(uri)!: reject uid-less current-user shorthand in favor of viking://~
viking://user/<segment> (memories/resources/skills/peers/privacy/sessions
without a user id) was ambiguous with a user literally named after the
segment, and a user actually named e.g. "memories" was unreachable for
USER/ADMIN callers. Now that the viking://~ home alias (#4167) covers the
same need unambiguously, the shorthand fails closed at the request
boundary instead of expanding:
- resolve_current_user_uri raises NamespaceShapeError with a corrective
hint naming both viking://~/<rest> and the explicit-uid form. Silently
parsing the reserved segment as a peer user id would misdirect reads
and writes, so rejection is the only safe removal.
- Bare viking://user falls through to the canonical parser and keeps
container semantics (a user key listing it sees only its own space).
- The self-id escape stays: a caller whose user_id equals a reserved
name keeps viking://user/<own-id> as their canonical root. ROOT-role
literal parsing and the legacy viking://session alias are unchanged.
- AddTargetsConfig normalizes stored legacy config spellings
(viking://user/resources|skills) to the viking://~ form at validation
so existing ov.conf/user_config deployments keep working; the accepted
per-user spelling is now viking://~/resources and viking://~/skills.
- usage_reporter keeps canonicalizing the historical shorthand found in
old transcripts and additionally recognizes viking://~/memories/.
BREAKING CHANGE: requests using the uid-less viking://user/<segment>
spelling now fail with 400; use viking://~/<segment> or an explicit
viking://user/{user_id}/<segment> URI.
* refactor(clients): migrate first-party emitters to the viking://~ home alias
Every in-repo client that emitted the removed uid-less current-user
shorthand now sends viking://~/... instead: vikingbot fallbacks and
default sentinels, the LangChain store/tools defaults, the shared
recall-core.mjs (all synced plugin copies), the codex/claude-code/
openclaw/openwebui/dsh/zcode/pi plugin emitters, quick-app examples,
Go SDK example, tau2 benchmark targets, and the eval golden dataset.
Compat kept where legacy strings live in stored user configs: bot and
ov_dream sentinels accept both spellings while emitting only ~, and
recall-core still rewrites legacy viking://user/<reserved> config values
client-side. langchain_openviking._uri now classifies viking://~ with
the explicit-user shape so canonicalized server responses keep matching
a ~ root. Plugin READMEs note the server requirement for the alias.
* docs: replace current-user shorthand guidance with the viking://~ home alias
Rewrite every EN/ZH doc and model-facing prompt that advertised the
uid-less viking://user/<segment> spelling: URI concept catalogue,
context-types/storage/extraction/retrieval/session/privacy concepts,
configuration guide (with the legacy add_targets auto-normalization
note), resources/skills/sessions/retrieval/admin API references, FAQ,
capability reference, and the openviking-memory / ov-experience-memory /
openclaw / ov-resources skills. The stale MCP viking://user/<path>
dialect passage in the MCP guide is replaced by ~ guidance, and bare
viking://user is documented as the container of user spaces.
* test(api): migrate live API session-used tests off the removed shorthand
tests/api_test/sessions sent uid-less viking://user/skills/... URIs to
record_used, which the request boundary now rejects with 400 (caught by
the API & CLI Integration Tests CI job; these tests need a live server
and are not part of the local suites). The api_test client authenticates
as an admin-role user key, so the viking://~ home alias expands for it.
tests/api_test/common/test_edge_cases.py is left as is: it asserts a 400
for a non-resource add target, which still holds.
* feat: add freshness-aware parent aggregation
Defer wide-directory abstract/overview regeneration until the configured freshness threshold is reached while continuing changed-file semantic and vector processing.
Persist freshness metadata atomically, make parent bubbling L0-aware, preserve separate semantic/vector statuses, and keep explicit waits synchronous.
Rebuild every sampled summary on threshold refresh and always retry directory vectorization so stale sidecars or transient vector failures cannot be silently accepted.
Add focused coverage for freshness policy, pending-state consumption, sampled-summary refresh, vector retries, and parent bubbling.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive, and applied to memory
* feat: ov reindex support --recursive, and applied to memory
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat(uri): add viking://~ home alias for the caller's user root
Accept `viking://~` (and `viking://~/<suffix>`) as a server-side alias
for the authenticated caller's user namespace root. The alias is
expanded at the request boundary by resolve_current_user_uri for
USER/ADMIN identities, so every control plane that already funnels
through validate_request_viking_uri (REST, MCP, and therefore CLI/SDK
clients) gets it with no client changes.
Design points:
- `~` is a reserved token that can never collide with a real user id
(validate_user_id's charset excludes it), so segment-0 aliasing does
not weaken canonical-first parsing.
- The canonical parser (resolve_uri) rejects the alias outright,
mirroring the legacy-session handling: root-role requests, internal
callers, and storage paths fail closed instead of materializing a
literal '~' directory.
- The alias is accepted but never advertised: scope error copy
("Must be one of: ...") filters it at both the parser and the public
validator, and VikingURI.build refuses to mint it, so responses and
persisted data stay canonical.
- MCP search now resolves exclude_uris with the same strictness as the
REST search router (closes the one entry point that skipped it).
* refactor(mcp): drop tilde mention from tool docstrings
Agents pass viking://~ through verbatim and responses echo the
canonical form, so the alias is self-explanatory on contact; carrying
the sentence in 12 tool descriptions costs every MCP session tokens.
Docs keep the alias documented for humans.
* feat(dsh): serve tools over the shared stdio MCP proxy
Replace the dsh bundle's seven hand-registered `viking_*` tools with the
OpenViking MCP surface, reached through the same stdio proxy every other
memory integration starts, and collapse the four duplicated proxy
entrypoints onto a shared config builder.
The bundle now mounts `@deepseek-ai/dsh-mcp-client` (which ships with dsh
itself) on `servers/mcp-proxy.mjs`. Pointing an MCP SDK client straight at
the server's `/mcp` endpoint does not work: with `stateless_http=True` the
server still answers `GET /mcp` with an idle 200 SSE stream, and once the
SDK client opens that standalone stream it stops resolving POST responses,
so `tools/list` never returns. The stdio proxy owns the transport itself
and is unaffected.
`trimSlash`, `normalizePath`, `uniq`, the watched-credential-path list and
the cfg -> proxyConfig mapping existed in four near-identical copies
(claude-code, codex, opencode, agent-plugins; the last one carried a
"keep in sync with claude-code" comment). They move to
`memory-plugin-shared/lib/mcp-proxy-config.mjs` and all five entrypoints —
including the new dsh one — now shape their config through
`buildMcpProxyConfig`. Behavior is preserved per field, including codex's
explicit `mcpUrl` override, claude-code's `ovcli.conf` credential-source
probe, and opencode's extra watched config file.
The bridge is mounted last in `apply()` so a proxy that fails to start
cannot hold up profile injection, recall, capture, commit, or the URI
guard registrations above it.
* feat(dsh): add to the unified installer and ship the shared skill
The bundle now registers its own isolated `ctx.skills` provider serving the
shared `openviking-memory` skill, so DSH gets the same guidance the Claude
Code, Codex, and Cursor integrations ship. `sync.mjs` distributes the skill
to the bundle, and the provider uses `includeDefaultRoots: false` so it
never shadows DSH's own project/user skill catalog.
`install.sh` grows a `dsh` harness id, auto-detected like the others, plus a
profile prompt that defaults to `web` (`--dsh-profile` / `OPENVIKING_DSH_PROFILE`
answer it up front). The installer always installs the published package:
`dsh plugin` forwards to pnpm, and a linked source tree cannot resolve the
dsh peers the bundle imports because Node resolves them from the checkout's
realpath rather than from the profile.
Documentation is restructured around installing rather than internals. The
integration page now leads with the one-line installer and keeps behavior at
the level the other harness pages use, with configuration in a details block;
design rationale moves to the bundle README, which itself leads with Install
and groups the rationale under "Design notes". Capability-reference claims
that dsh is outside the unified installer are corrected.
* chore(dsh): release 0.2.0
The MCP tool surface, the stdio proxy transport, and the bundled skill all
change what the bundle does for an existing user, so this is a minor bump
rather than a patch. 0.1.0 remains the native-`viking_*` tool surface.
* docs(dsh): note pnpm's 24h minimum release age
pnpm 11 refuses releases younger than minimumReleaseAge (24 hours by
default), and surfaces it as a registry 404, so installing a freshly
published version reads as "the package does not exist".
* fix(dsh): honour dev source mode in the installer
install_dsh ignored SOURCE_MODE and always fetched the published package,
so selecting "current checkout" installed npm's build instead of the
working tree and validation still reported success.
npm is the bundle's only distribution channel, so the github/tos choice
does not apply to it: every mode except dev now installs the published
package, and dev packs the checkout with npm pack first. It has to arrive
as a real package rather than a link, because a linked source tree
resolves its dsh peers from its own realpath and misses the profile's
hoisted node_modules. The install line reports which source was used.
* fix(dsh): make repeated installs actually overwrite
Two ways a re-run silently kept stale code:
pnpm treats an already-satisfied version as a no-op regardless of which
tarball the file: dependency points at, so a dev re-install after editing
the checkout left the previous build in place. Local installs now drop the
package before adding it back; that is confined to local sources, since
doing it for the registry path would leave nothing installed when add
fails.
A bare package name has the same effect in reverse: a profile holding a
dev build satisfies it, so switching back to the published package was a
no-op. The registry path now asks for @latest.
The packed tarball is named after a fingerprint of the checkout's shipped
files, so an unchanged checkout skips the pack and keeps a stable path in
the profile lockfile.
* fix(feishu): support legacy doc imports
Keep legacy Feishu doc resources on the doc API path instead of routing doccn tokens through docx blocks.
* test(feishu): trim legacy doc coverage
Keep legacy doc import coverage focused on public Feishu accessor behavior and remove extra mock-level assertions.
* docs: fix integration docs and comments that contradict the code
- codex: credential resolution in the default `auto` mode is env-first — `credentials.mjs`
only falls back to `ovcli.conf` when no credential env var is set, while the docs and the
`config.mjs` header comment claimed `ovcli.conf` wins by default. Also document
`OPENVIKING_CREDENTIAL_SOURCE=cli`, which was undocumented.
- codex: the four hook scripts send the key as `X-API-Key` in addition to
`Authorization: Bearer`; the README documented Bearer only.
- claude-code: the OV session id is `cc-<cc_session_id>` verbatim (`deriveHarnessSessionId`
does no hashing), not `cc-<sha256(cc_session_id)>`.
- claude-code: `hooks.json` registers 9 hooks, not 7 — the responsibilities table was
missing the `PreToolUse` `viking://` guard and the `PostToolUse` skill-experience hook.
- claude-code: archival is triggered client-side (the `Stop` hook commits once
server-reported pending tokens cross `commitTokenThreshold`, default 20000, plus
unconditional commits from `PreCompact` / `SessionEnd` / `SubagentStop`). The README
attributed it to a server-side `auto_commit_threshold`, but
`memory.session_auto_commit.default_enabled` is false and no plugin sends a policy.
- trae / opencode: the MCP proxy transparently exposes the full server tool set (16 tools);
the docs listed a 4-item sample or 11-13 tools and omitted `tree` / `write` / `edit`.
- trae-cli: the installer registers the MCP server as `openviking-memory`, but the verify
step told users to look for `openviking`.
- pi: the manual install block omitted the `pi install <dest>` registration step that the
one-click installer runs, so a hand-copied extension is never registered.
- install.sh: `--uninstall` handles cursor, trae, trae-cn, trae-cli and zcode; the `--help`
text still said Cursor/TRAE only.
- mcp_endpoint.py: the module docstring enumerated 13 tools and omitted `recall`,
`list_watches` and `cancel_watch`; replaced the stale enumeration with a pointer to the
`@mcp.tool` registrations.
* docs: add a cross-integration capability reference page
The agent-integrations section had per-integration install guides but no place
to compare integrations against each other. This adds one bilingual page that
does that, and wires it into the existing pages in both directions.
- New page `docs/{en,zh}/agent-integrations/16-capability-reference.md`: a
dimension-first comparison of every OpenViking integration — active tool
surface, automatic hook surface, install/credential/config layering, recall
and injection, session and commit lifecycle (including a shutdown-path x
harness end-state matrix), compaction takeover, write/delete boundaries,
degradation, and a per-harness profile card for each integration.
- Sidebar: `StructuredSidebarCopy` gains an optional `topItems` field so a
section can list flat entries next to its overview; agent-integrations uses
it to place the new page beside the overview. Other sections are unaffected.
- Links both ways: the overview and all 14 per-integration pages link to the
reference, and the reference links back to each integration page from its
profile card, from the non-coding integration table, and from the custom
agent integration paths. Section cross-references (§x.x) are real in-page
anchor links, generated from the built heading ids.
- trae-cli is documented as TraeCode CLI 2.0 only, installed through a codex
plugin alias; 1.0 and its standalone plugin are called out as unsupported.
- The MCP tool surface is described as 15 tools throughout, matching the
removal of the `recall` tool in favour of `search` with `mode="context"`.
Pages outside this change that still mention an MCP `recall` tool
(04-codex, 12-cursor, 15-agent-plugins, guides/06-mcp-integration) need a
follow-up sweep once that removal lands.
* docs: 更新服务端 MCP 工具面描述,简化信息并明确更新方式
* docs(hermes): recommend memory setup openviking
Point Hermes docs at `hermes memory setup openviking` and describe the
real wizard. Shorten Volcengine console agent guides to TOS + cloud API
key, and mark docs/images as console-only.
* docs(hermes): drop setup filler
Keep the command and the two connection paths. Remove picker
explanations and wizard narration.
* docs: keep images AGENTS.md.local local-only
Ignore AGENTS.md.local like AGENTS.md. Drop the unused
docs/images/README.md.
* docs(console): keep harness setup to command plus API key
Drop installer narration, idempotency notes, and copied site
guides. Console pages only need the TOS command and cloud key.
* docs(console): restore Install / Verify / Troubleshoot
Keep the short cloud setup, put it back under the three section
headings the console pages use.
* docs(console): add Reference links
Point each harness page at docs.openviking.net, the coding-agent
blog where it exists, and the example source.
* docs(console): label Reference as manual settings and blog
Use Docs on Manual Settings for the full site page. Use Blog
about how it works where a how-it-works writeup exists.
* feat: add OpenViking memory integration for TRAE CLI
Add TRAE CLI lifecycle hooks and MCP proxy support, wire the integration into the shared installer, and cover idempotent install and uninstall behavior.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(trae-cli): cover archive installs and hook payload aliases
* fix: keep TRAE CLI installation explicit
Leave TRAE Desktop detection unchanged and avoid auto-selecting TRAE CLI. TRAE CLI remains available through an explicit harness selection or --harness trae-cli.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(trae-cli): auto-select installed CLI commands
Detect traecli and traex only when they are available in PATH, then mark and select the TRAE CLI harness automatically.
---------
Co-authored-by: “bianhaonan” <“bianhaonan@bytedance.com”>
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(pdf): refactor MinerU parsing to the official file_parse API
* feat(pdf): remove mineru_api_key from configuration and examples
* feat(pdf): preflight MinerU /health during service initialization
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* feat(plugins): add Agent Plugins 1.0 portable package
Add agent-plugins/, an Agent Plugins 1.0 conformant package
(https://agent-plugins.org/specification) that any conforming client can
load: plugin.json manifest, an openviking-memory skill teaching the
hook-less recall + persist loop, and an mcp.json stdio entry running a
stdio -> streamable-HTTP proxy that resolves credentials from
OPENVIKING_* env -> ~/.openviking/ovcli.conf -> ~/.openviking/ov.conf,
same as the ov CLI.
servers/shared/* are generated copies of memory-plugin-shared/lib, wired
into sync.mjs / sync.test.mjs TARGETS so they cannot drift silently.
config.mjs / debug-log.mjs / mcp-proxy.mjs are adapted from
claude-code-memory-plugin with the hook-tuning knobs dropped.
plugin.test.mjs validates spec conformance (schema URLs and matching
spec versions, name rules, closed manifest root, semver, skill
frontmatter, referenced files staying inside the plugin root, node
--check on all .mjs) and runs in CI via pr.yml.
The skill treats tree/write/edit as optional, since they only exist on
servers that carry #3936.
Docs: docs/{en,zh}/agent-integrations/15-agent-plugins.md, registered in
the VitePress sidebar and the integration overview tables, plus a link
from the three root READMEs. The docs recommend the per-client plugin
whenever the harness has hooks, with the shared installer one-liner.
Based on #3994 by @ZaynJarvis.
Co-Authored-By: Zayn Jarvis <zaynjarvis@gmail.com>
* docs(agent-plugins): pluralize README title
---------
Co-authored-by: Zayn Jarvis <zaynjarvis@gmail.com>
* feat(reindex): support tag updates
Add replace and append tag modes to reindex vector rebuilds, preserve omission-aware behavior, and propagate options through background tasks and namespace rebuilds. Align Python, TypeScript, Go, and CLI interfaces with tests and documentation.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(reindex): lock file targets exactly
Use an exact path lock for existing file targets while retaining tree locks for directories and prune-orphans scopes. Add a regression test for single-file reindex.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(reindex): handle prune file targets
Use exact locks for existing file targets in prune-orphans mode while retaining tree scope for missing targets. Document the existing Go ReindexOptions wait semantics and add lock regression coverage.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(storage): keep non-memory appends free of memory trailers
ContentWriteCoordinator._write_in_place routed every append through
MemoryFileUtils, which strips the existing trailing newline and appends
a reserved MEMORY_FIELDS metadata trailer, even for resource/skill files
where MEMORY_FIELDS is not a reserved format (see content_visibility).
Append to non-memory files now concatenates raw content instead, matching
POSIX append semantics and the documented visibility rules.
* feat(mcp): add write tool with exact-string edit support
Agents could not use viking:// as a working directory through MCP: no
tool could create or update file content. Add a write tool covering full
writes (mode=replace as create-or-overwrite, append, strict create) and
targeted edits (a list of {old_string, new_string, replace_all}
exact-string replacements applied in order, all-or-nothing), following
the Write/Edit conventions of common agent harnesses.
Edits read via read_visible and write back through the content-write
coordinator, so memory metadata trailers are preserved and semantic /
vector re-indexing triggers as with any other write. Parent directories
are created automatically by the storage layer. Descriptions spell out
writable scopes (resources, user memories/resources, agent) and the
wait=true knob for read-after-write search consistency.
Also update the stale tool-count comment in app.py and the MCP tool
tables in the en/zh guides (13 -> 14 tools).
* feat(mcp): add tree tool, split targeted edits into edit tool
tree renders the recursive directory tree under a viking:// URI,
indented by depth with file sizes, for whole-layout orientation;
level_limit/node_limit bound the output and include_abstract adds
per-file summaries. Missing directories report "(nothing under ...)"
instead of an error, matching the read tool's convention.
edit(uri, old_string, new_string, replace_all) takes over the targeted
exact-string replacement that previously lived in write's edits array,
matching the classic Edit tool signature harnesses already train on.
write now only does full-content writes (content + mode), removing the
mutually-exclusive content/edits schema ambiguity. Edits still read via
read_visible and write back through the content-write coordinator, so
memory metadata trailers are preserved and re-indexing triggers as with
any other write.
* test(plugin): update canonical MCP tool list for tree/write/edit
The marketplace test pins the server-registered MCP tool list; add the
new tree, write, and edit tools to fix plugin-tests CI.
* feat(storage): support plain files at the user scope root
Agents treating viking:// as a working directory naturally drop files
like viking://user/zeus-persona.md at the user root, but the write
coordinator only accepted the memories/ and resources/ subtrees.
Two changes make that work:
- Namespace shorthand: a dotted first segment under viking://user/ is a
file name, not a user id (canonical user ids are dot-free by
convention), so viking://user/zeus-persona.md now canonicalizes to
viking://user/<current-user>/zeus-persona.md, matching how the
reserved memories/resources/skills segments already shorthand.
Dot-free segments still address an explicit user, and an exact match
with the current user id still wins.
- Coordinator: plain files directly under the user root (or in
non-managed subdirectories) anchor their semantic refresh at the
parent directory. The managed subtrees skills/, peers/, privacy/ and
sessions/ remain read-only with an actionable error message.
* fix(namespace): narrow user-root shorthand to text-file extensions
Review on #3936 (codex /review-pr) flagged that treating any dotted
segment as a user-root file shorthand would silently re-route canonical
URIs for valid dotted user ids (e.g. alice.smith) into the current
user space. Shorthand now triggers only when the first segment ends
in a common text-file extension; dotted or email-style user ids keep
resolving as canonical user ids. Adds regression tests pinning both
behaviors.
* fix(mcp): resolve user URIs against current user
* test(mcp): pin plain-file writes directly at the user root
The user-root shorthand exists so an agent can drop viking://user/persona.md
into its workspace, but every new test went through an intermediate directory
(viking://user/project/zeus-persona.md), leaving the no-directory shape — the
one that anchors the write coordinator's refresh at the user root itself —
uncovered. Add the missing case.
Also correct the write tool docstring: the create-extension allowlist applies
to any newly created file, including one created by mode="replace" falling
back to create, not only to an explicit mode="create".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Coding-agent plugins capped a tool part's `tool_output` at 2000 chars before
POSTing it to `/api/v1/sessions/{id}/messages`. That cap sits below the server's
own externalization threshold (`tool_output_externalization.threshold_chars`,
default 20000), so output in the 2k-20k band was destroyed for no reason and
anything larger never reached `ToolResultStore` - leaving `tool_output_ref`
permanently empty and the `/tool-results` read-back path unusable.
Raise the `captureToolMaxChars` default to 1000000 (a guard against pathological
payloads, not a truncation policy) and lift the opencode/pi clamps that would
otherwise pin it back to 20000. claude-code had no knob at all - two hardcoded
`TOOL_OUTPUT_PART_MAX_CHARS = 2000` constants - so it gains the same config
entry and both capture scripts now read it.
Also stop pi from sending tool output twice: for a tool-only payload the
rawText-derived text part re-rendered the same output the tool part carries.
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix: remove heima partner, clean up auth docs, add web-studio unsupported auth banner
- Remove heima from partner list in README (en/zh/ja)
- Remove unsupported env var references (OPENVIKING_AUTH_MODE, OPENVIKING_USERNAME,
OPENVIKING_PASSWORD) from LDAP auth docs
- Remove temporary switch bash snippets from auth docs
- Fix ldap_password description
- Add web-studio unsupported-auth-mode banner for oidc/ldap servers
* fix: address OIDC/LDAP review comments on auth plugin design
Key changes driven by PR review:
- **Role mapping**: OIDC and LDAP external identities always resolve to
USER role. Removed map_role() calls and group_membership-based role
mapping. Admin access is gated by the root API key mechanism only.
- **LDAP credential extraction**: Removed query-parameter-based username/
password extraction (security concern — passwords in URLs can leak via
shell history, proxy logs, and monitoring). Clients must use Basic Auth
header or form data.
- **OIDC identifier sanitization**: Auth0 and other providers may include
characters like "|" in the `sub` claim. These are now replaced with "_"
to produce valid OpenViking user identifiers.
- **Dead code removal**: Removed _extract_groups, memberof_attribute,
require_root_api_key_for_admin, _initialize_api_key_manager, and
get_request_context_checks from both plugins since they are no longer
needed.
- **Docs**: Removed query-parameter curl example, memberof_attribute and
require_root_api_key_for_admin config references.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix(auth): bind lazy OIDC imports at module scope
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(retrieval): honor tier ceilings and stop cooling unserved recalls
Follow-up to #3534, from its post-merge review round.
- The abstract-to-overview substitute now applies only to categories whose
stored abstract is the whole file body. A resource or skill whose abstract is
missing (`processing_mode=vectors_only`) or over the per-entry cap read its
body and returned an overview instead, which for a short file is the body
almost verbatim — crossing the opt-in deepening boundary those categories are
documented to have, and doing it even under an explicit `detail="abstract"`.
They now degrade to a bare URI and their body is never read.
- A digest reporting `no_relevant` blanks `rendered`, so the client injects
nothing, yet those URIs still entered the dedup ledger and were cooled for
`dedup_turns` turns. That contradicted the ledger's own bare-URI grace rule
and held memories back from the later turn they were relevant to.
- Flat retrieval reaches built-in memory types outside the four named ones
(`cases`, `patterns`, `tools`, `trajectories`, skill-usage memories) and
reported them as an undeclared `memories` category that no tier or penalty
table covered, so other-peer hits skipped the score penalty and callers could
not pin their tier. The catch-all is now a declared category with both; it
stays out of `quotas`, whose buckets it would overlap. Skill-usage memories
also stop being misread as the `skills` category.
- ZCode, OpenCode and pi own an OV session id but did not forward it, so their
recalls silently ran without query expansion or cross-turn dedup.
- The context-request deadline covered only the server's 30s rewrite fuse, but
the pipeline is serial: expansion, retrieval and budgeting all precede it.
45s covers both fuses and the work between them.
- `plugin` config scope and the `/recall` successor example now match what the
code actually does.
* fix(retrieval): make the context deadline and expansion opt-out reachable
Forwarding a session id turns on server-side query expansion, an LLM call with
its own 5s fuse, but neither the deadline that was supposed to cover it nor the
switch that turns it off reached the two harnesses this PR newly enabled it for.
- `contextRequestTimeoutMs()` now derives the deadline from the request body
rather than from `cfg` plus a rewrite flag. The body is what states which
server stages will run: a session takes the expansion fuse, `rewrite` takes
the digest fuse, and a bare retrieval takes neither and keeps the caller's own
budget. Reading `cfg` alone could not tell those apart.
- OpenCode pinned `timeoutMs: 5000` after spreading the helper's options and pi
ignored them entirely, so the helper's deadline was dead code in both. Their
own budgets are now defaults rather than ceilings. OpenCode's 5s in particular
was shorter than the expansion fuse it had just enabled, so a legal request
would have been aborted client-side and dropped back to the path with neither
dedup nor expansion.
- OpenCode and pi read `OPENVIKING_RECALL_QUERY_EXPANSION` (and
`recallQueryExpansion` in their own config files) and set the `configured`
flag the shared body builder requires, so the documented opt-out exists where
the cost was introduced.
- The integration overview no longer implies every harness reads the same
environment knobs, and describes the deadline as per-stage rather than
rewrite-only.
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"
Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.
This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.
New assembly kernel under openviking/retrieve/context_assembler/:
- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
readable floor, then overview, then full for high-scoring entries. An
oversized tier falls back to the previous one instead of being truncated,
bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
Summary section, code files reuse code_outline signatures, long documents use
a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
directories carry no stored abstract; their full tier stays capped at
overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
every harness inherits cross-turn dedup; exclude_uris remains as the
stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
element per entry. Every tier carries its URI, so the model can always drill
down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
timeout fuses, and a failed rewrite still returns the unrewritten block.
Retrieval failures are counted into stats rather than silently yielding an
empty block.
Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.
* refactor(retrieval): give context tiers a per-category default
The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.
Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.
Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".
Assembly fixes found alongside:
- Removing the abstract cap exemption would turn an oversized abstract into
a bare URI, so it now falls back to overview first — for memory that is a
cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
`asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
such attribute; usage was structurally always null. It now reads the model
instance's tracker and reports only when the call count moved by exactly
one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
fail, and the file was never rewritten, so it could not heal. Records are
now coerced on read and dropped on the next write, along with records left
ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
`viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
resolved a different profile than `POST /recall`; an unknown `detail` value
raised `KeyError` through the whole call instead of degrading.
* feat(codex): inject profile context on session start
Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): raise rewrite timeout default to 30s
* docs(agents): document low-latency recall settings
* fix(codex): prefer luna as recall compressor fallback
* refactor(plugins): unify recall compression setting
* feat(plugins): enable recall compression by default
* docs(agents): use absolute links in image docs
* fix(retrieval): address context assembly review feedback
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test: trim redundant context assembly coverage
* fix(retrieval): address second-round context assembly review
- Drop the backticked `/search` from the deprecated-recall row in both API
overviews. The reference checker scans the whole row after the method cell
for backticked paths, so it read the description as a route named
`POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
belongs to the Rust CLI, which writes `root_api_key`, `output`,
`echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
two Python readers had drifted into stricter subsets, so the shipped example
already failed to load in both. Adding the new `plugin` section to a working
ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
Retrieval validates query and image_url before searching, and the gather
fuse swallowed that rejection along with genuine scope failures, so a body
of `{"mode":"context"}` came back 200 with an empty block instead of the
documented parameter error. Runtime failures still degrade into
`stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
than the 30s fuse, so a rewrite that finished inside its own budget was
aborted client-side, discarding the whole response — including the
uncompressed block the server returns when a rewrite fails — and falling
back to `/recall`. The deadline is only extended when the body actually
requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
`plugin.recallContextTimeoutMs` pins it.
* chore(plugins): sync shared modules into the zcode snapshot
* fix(retrieval): align context quotas and plugin defaults
Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): preserve recall compatibility
Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>