* fix(watch): avoid global lock during resource moves
Allow unrelated watch operations to continue while viking_fs.mv is running, while serializing overlapping resource paths with a move fence. Preserve transaction integrity across persistence failures, rollback, and caller cancellation.
* feat(resource): add URI mutation coordinator
* refactor(watch): separate target rewrites from resource moves
* refactor(fs): own resource move watch transaction
* feat(watch): coordinate refreshes with URI mutations
* refactor(core): share URI mutation coordinator
* test(watch): focus URI move coverage
* test(watch): reuse existing URI move coverage
---------
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
Keep expired terminal tasks cached when persistent deletion fails so cleanup can retry without resurrecting stale records after restart.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Materialize session-aware context requests through SessionService before
loading the recall ledger, while retaining the messages.jsonl guard for
partially initialized sessions.
Add coverage for first-turn ledger writes, same-session message capture,
materialization failure recovery, and stateless requests with session
features disabled.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(memory): support event tag filtering
Add session-level default event tags, commit-time overrides, durable queue propagation, and first-write vector index tagging. Include config update APIs and coverage for serialization, concurrency, extraction, and HTTP behavior.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(memory): expose event tags in SDKs and CLI
Add session default tag configuration, config updates, and commit-time event tag overrides across embedded Python, standalone Python, TypeScript, Go, and the Rust CLI. Preserve explicit empty-tag semantics and document each public interface.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(sdk): align legacy session tag APIs
Forward commit-time event tags through the legacy Python HTTP shims and align BaseClient session signatures without adding a new abstract-method requirement for existing subclasses.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(session): allow updating auto-commit policy
Extend PATCH session config to atomically update event tags and auto-commit settings. Merge policy objects by field, use explicit null to disable automatic commits, preserve omitted fields, and expose the contract across SDKs and CLI.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(session): align session config interfaces
Replace the generic session create config JSON flag with explicit event-tag and auto-commit options. Preserve omitted, object, and null auto-commit semantics across HTTP, embedded clients, SDKs, and CLI, reject ambiguous null policy fields, and handle nullable event configuration consistently.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test(session): trim redundant event tag tests
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
- keep handed-off leases in the registry with pending_handoff so auto-refresh continues
- include lease_ref in PathLockHandoffRef for same-process adopt fast path
- rotate both lease_ref and ownership_ref on adopt to invalidate stale producer capabilities
- reject stale capability use in release, release_selected and refresh
- validate owner_id, lock_paths and covered_paths before local adopt
- reject replayed fallback adopt when the same owner/path token is already held
- add tests for pending handoff refresh, retryable adopt race, forged coverage and replay rejection
Read replicas load the API key store once at startup and never rewrite
it, so a user registered/rotated/removed on the writer stays invisible
(new key -> "Invalid API Key"; removed key -> still accepted).
Add an optional background watcher that polls the shared key store and
reloads the in-memory index only when it actually changes:
- APIKeyManager.reload(): strictly read-only refresh that rebuilds state
and swaps it in atomically, never writing or migrating plaintext keys.
- compute_store_signature(): cheap (path, size, modTime) signature over
accounts.json + every users.json so the watcher skips unchanged polls.
- ApiKeyAuthPlugin starts/stops the watcher behind api_key_watch_enabled
(default off) with api_key_watch_interval_seconds; AuthPlugin.shutdown()
is wired into app shutdown to cancel it cleanly.
Add coverage for reload convergence, read-only/no-migrate guarantees,
uninitialized-store tolerance, signature change detection, and watcher
reload/skip/shutdown behavior.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
_smart_split_content documents that it enforces both a token limit
(max_size) and a hard character limit. But when a single paragraph was
oversized by tokens yet under the character limit, the force-split loop
stepped through it by max_chars only, so it emitted a chunk that still
exceeded max_size tokens.
This is reachable with ordinary long-form CJK text: _estimate_token_count
weights CJK at ~0.7 token/char, so a ~5000-char Chinese paragraph is
~3500 tokens (over the 2048 default) while staying under the char limit,
and was returned as a single over-budget chunk.
Bound the force-split step by min(max_chars, max_size / MAX_TOKENS_PER_CHAR),
where MAX_TOKENS_PER_CHAR is the worst-case (CJK) density already used by
_estimate_token_count, now extracted into a shared constant so the two stay
in sync.
Add regression tests for the token budget and content preservation.
fix_fields_data backfills schema fields absent from a row's data with their
defaults, but it first short-circuited on
`len(field_data_dict) >= len(field_meta_dict)`, using field count as a proxy
for "all schema fields are present".
That proxy is wrong: a row can have as many keys as the schema (or more) while
still missing a specific field — e.g. a row written before a new field was
added that also carries an extra internal/non-schema key. In that case the
fill loop was skipped and the missing field was silently omitted rather than
defaulted, so downstream reads see an incomplete record.
Remove the count guard. The loop already skips fields that are present, so
complete inputs are returned unchanged; only genuinely missing fields are now
filled.
Add regression tests: a missing field with matching key count is backfilled
(default value and type default), and complete data is returned unchanged.
_gh_slug claims to produce "GitHub-style" heading slugs but collapsed runs
of whitespace (`re.sub(r"\s+", "-", s)`). GitHub's reference slugger
(github-slugger) maps each space to its own hyphen (`.replace(/ /g, '-')`)
and does not collapse.
Because punctuation is stripped before spaces are converted, a heading like
"Foo & Bar" leaves two adjacent spaces where "&" was. GitHub renders this as
"foo--bar", but _gh_slug produced "foo-bar". The intra-document link rewriter
(_rewrite_link) compares _gh_slug(heading) against the link fragment, so an
author-written link such as `guide.md#foo--bar` failed to match its heading
and was left unrewritten after the target doc was split into sections.
Replace `\s+` with `\s` so each whitespace character maps to one hyphen,
matching GitHub. Simple single-space headings are unaffected.
Add regression tests covering punctuation headings and ordinary headings.
The range checks used strict inequalities (< instead of <=), rejecting
the valid boundary values defined by the WGS-84 geographic standard:
- latitude ±90 (the North and South Poles)
- longitude ±180 (the antimeridian / International Date Line)
Any resource whose geo_point is exactly on a pole or the date line cannot
be indexed, and a geo_range query centred on those boundaries raises
ValueError instead of executing the search.
Change both checks to <=. Values strictly outside the valid interval
(e.g. 181, -91) continue to raise ValueError as before.
Add a regression test covering all four boundary endpoints and all four
out-of-range coordinates.
The except block in VikingFS.mkdir() had no re-raise, so any backend
failure that was not an already-exists error (permission denied, quota
exceeded, I/O errors, lock-lease violations) — and even already-exists
errors with exist_ok=False — was silently discarded and mkdir() returned
as if the directory had been created. Callers on the write hot path
(ovpack import, parsers, session, privacy) then write into a directory
that may not exist, and the original actionable error is lost.
Re-raise the original exception unless it is an already-exists error
tolerated by exist_ok=True.
Also update tests/misc/test_mkdir.py, which still mocked fs.agfs.mkdir
even though mkdir() now goes through the AsyncAGFSClient wrapper
(self._async_agfs) — the swallowed-attribute-error made the stale tests
pass/fail for the wrong reasons. Add regression tests covering error
propagation for both exist_ok values.
_copy_dir_through_vikingfs() drives the copy phase of mv() for non-temp
directories, but enumerated the source with the agent-facing ls() default
node_limit=1000. Any directory level with more than 1000 visible entries
was copied only partially, and mv() then unconditionally deleted the
source recursively — permanently losing every entry past the cap, while
the vector index (remapped via the uncapped _collect_uris) kept pointing
at URIs that no longer exist anywhere.
Pass the module's LS_ALL_NODES sentinel, which exists precisely for
internal callers that must enumerate an entire directory.
Add a regression test that fails without the fix.
The row serializer writes each string's byte length as a UINT16 prefix. The
scalar `string` field guards this contract — a >65535-byte value raises a
clean, field-attributed ValueError. The `list<string>` element path uses the
identical UINT16 prefix but had no such guard, so an oversized element instead
raised a raw `struct.error: 'H' format requires 0 <= number <= 65535` from
deep inside struct.pack_into — a different exception type, naming no field.
Callers that catch ValueError (matching the documented scalar contract) do not
catch this, and the error gives no clue which field/element overflowed. Apply
the same bounds check to list<string> elements so inclusion in a list does not
silently downgrade the type-checked contract the scalar path upholds.
Add a regression test asserting the clean ValueError for an oversized element
and that an in-bounds (incl. multibyte) list still round-trips.
Several routers returned Response(status='error', error=ErrorInfo(...))
directly, so FastAPI shipped them with HTTP 200 instead of the mapped
status from ERROR_CODE_TO_HTTP_STATUS. Web-studio's sessions API only
survived because of a code-based fallback; other clients (e.g. a generic
HTTP retry layer) would treat these as success and never surface the
error.
Switch the seven return sites to error_response() so the canonical
mapping drives the HTTP status, and add the three previously-unmapped
codes (INTERNAL_ERROR, NO_VECTOR_DB, INVALID_FILTER) to the map.
Residual follow-up from PR #1764 (ac9f679a).
Co-authored-by: ming <silverchris@foxmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Route internal viking:// Markdown targets through the existing resource navigator instead of the attachment download endpoint. Preserve download fallback behavior for previews without a navigation callback and add a click regression test.
Test: NODE_OPTIONS='--localstorage-file=/tmp/openviking-vitest-localstorage' pnpm test\nTest: pnpm build
FastMCP's Streamable HTTP session manager stores sessions in an
in-process dictionary. In multi-instance deployments, an initialize
request handled by one OpenViking instance creates a session that is
not available to subsequent requests routed to another instance,
causing intermittent 404 "Session not found" errors.
This resulted in MCP tool discovery failures (for example, OpenCode
reporting "Failed to get tools") and repeated 404 errors in server logs.
OpenViking MCP tools do not rely on FastMCP session state. The
remember() tool flow creates and manages its own OpenViking session,
so enabling stateless Streamable HTTP mode allows the MCP endpoint to
work correctly behind load balancers and other stateless deployment
topologies.
Co-authored-by: wangyu134 <wangyu134@58.com>
Add push trigger (paths: examples/openclaw-plugin/**) to the release
workflow, keep publish jobs enabled on push, serialize runs with a
concurrency group, and retry the ClawHub verify step which races the
registry's read-after-write / rate-limit window.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
parse() resolves SecretRef values to plain strings before returning, but the
parsed type still inherited string | OpenVikingSecretRef from the input type,
breaking tsc -p tsconfig.build.json (release prepack) since #3618.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Derive the statusline recall count from unique viking:// references in the injected context and bump the plugin to 0.4.4.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Keep retention options when retryable Claude session commits are queued and replayed. Add coverage for retryable storage conflicts.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Treat messages.jsonl as the materialization boundary for session-aware recall,
repair partial session roots during the existing authoritative append path,
and preserve Claude capture cursors when writes never reach the server.
Also replay explicitly retryable storage conflicts across memory plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix: remove heima partner, clean up auth docs, add web-studio unsupported auth banner
- Remove heima from partner list in README (en/zh/ja)
- Remove unsupported env var references (OPENVIKING_AUTH_MODE, OPENVIKING_USERNAME,
OPENVIKING_PASSWORD) from LDAP auth docs
- Remove temporary switch bash snippets from auth docs
- Fix ldap_password description
- Add web-studio unsupported-auth-mode banner for oidc/ldap servers
* fix: address OIDC/LDAP review comments on auth plugin design
Key changes driven by PR review:
- **Role mapping**: OIDC and LDAP external identities always resolve to
USER role. Removed map_role() calls and group_membership-based role
mapping. Admin access is gated by the root API key mechanism only.
- **LDAP credential extraction**: Removed query-parameter-based username/
password extraction (security concern — passwords in URLs can leak via
shell history, proxy logs, and monitoring). Clients must use Basic Auth
header or form data.
- **OIDC identifier sanitization**: Auth0 and other providers may include
characters like "|" in the `sub` claim. These are now replaced with "_"
to produce valid OpenViking user identifiers.
- **Dead code removal**: Removed _extract_groups, memberof_attribute,
require_root_api_key_for_admin, _initialize_api_key_manager, and
get_request_context_checks from both plugins since they are no longer
needed.
- **Docs**: Removed query-parameter curl example, memberof_attribute and
require_root_api_key_for_admin config references.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix(auth): bind lazy OIDC imports at module scope
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(retrieval): honor tier ceilings and stop cooling unserved recalls
Follow-up to #3534, from its post-merge review round.
- The abstract-to-overview substitute now applies only to categories whose
stored abstract is the whole file body. A resource or skill whose abstract is
missing (`processing_mode=vectors_only`) or over the per-entry cap read its
body and returned an overview instead, which for a short file is the body
almost verbatim — crossing the opt-in deepening boundary those categories are
documented to have, and doing it even under an explicit `detail="abstract"`.
They now degrade to a bare URI and their body is never read.
- A digest reporting `no_relevant` blanks `rendered`, so the client injects
nothing, yet those URIs still entered the dedup ledger and were cooled for
`dedup_turns` turns. That contradicted the ledger's own bare-URI grace rule
and held memories back from the later turn they were relevant to.
- Flat retrieval reaches built-in memory types outside the four named ones
(`cases`, `patterns`, `tools`, `trajectories`, skill-usage memories) and
reported them as an undeclared `memories` category that no tier or penalty
table covered, so other-peer hits skipped the score penalty and callers could
not pin their tier. The catch-all is now a declared category with both; it
stays out of `quotas`, whose buckets it would overlap. Skill-usage memories
also stop being misread as the `skills` category.
- ZCode, OpenCode and pi own an OV session id but did not forward it, so their
recalls silently ran without query expansion or cross-turn dedup.
- The context-request deadline covered only the server's 30s rewrite fuse, but
the pipeline is serial: expansion, retrieval and budgeting all precede it.
45s covers both fuses and the work between them.
- `plugin` config scope and the `/recall` successor example now match what the
code actually does.
* fix(retrieval): make the context deadline and expansion opt-out reachable
Forwarding a session id turns on server-side query expansion, an LLM call with
its own 5s fuse, but neither the deadline that was supposed to cover it nor the
switch that turns it off reached the two harnesses this PR newly enabled it for.
- `contextRequestTimeoutMs()` now derives the deadline from the request body
rather than from `cfg` plus a rewrite flag. The body is what states which
server stages will run: a session takes the expansion fuse, `rewrite` takes
the digest fuse, and a bare retrieval takes neither and keeps the caller's own
budget. Reading `cfg` alone could not tell those apart.
- OpenCode pinned `timeoutMs: 5000` after spreading the helper's options and pi
ignored them entirely, so the helper's deadline was dead code in both. Their
own budgets are now defaults rather than ceilings. OpenCode's 5s in particular
was shorter than the expansion fuse it had just enabled, so a legal request
would have been aborted client-side and dropped back to the path with neither
dedup nor expansion.
- OpenCode and pi read `OPENVIKING_RECALL_QUERY_EXPANSION` (and
`recallQueryExpansion` in their own config files) and set the `configured`
flag the shared body builder requires, so the documented opt-out exists where
the cost was introduced.
- The integration overview no longer implies every harness reads the same
environment knobs, and describes the deadline as per-stage rather than
rewrite-only.