Directory L1 overviews previously fed the model bare `[N]` file indices and
`[N] filename: summary` source lines, then post-processed by replacing `[N]`
with the entry filename. Because the model naturally copied the filename next
to the index (e.g. `### [1] filename` or `→ [1] filename`), the substitution
produced duplicated headings like `### filename filename` in both Quick
Navigation and Detailed Description sections.
Replace the index-reference scheme with collision-free link placeholders:
- Feed each entry a compact placeholder `(link: viking://input_sample_fN)` for
files and `viking://input_sample_cN` for subdirectories, and instruct the
model to emit standard Markdown links `[display title](placeholder)`.
- Resolve placeholders back to real `viking://` URIs (built from
`dir_uri + "/" + name`) in post-processing via `_replace_link_references`,
replacing the old `_replace_index_references`.
- Apply consistently across the single, truncated, and batched generation
paths; batched partials resolve placeholders before merge so the merge step
needs no further substitution.
This eliminates the duplicated-heading class of bugs entirely (bare `[N]`
collisions are far less likely than reused digits) and yields clickable,
accurate resource links in the generated overview, matching the Markdown
link convention already used elsewhere (session memory extraction).
Co-authored-by: Maojia Sheng <shengmaojia@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat(uri)!: reject uid-less current-user shorthand in favor of viking://~
viking://user/<segment> (memories/resources/skills/peers/privacy/sessions
without a user id) was ambiguous with a user literally named after the
segment, and a user actually named e.g. "memories" was unreachable for
USER/ADMIN callers. Now that the viking://~ home alias (#4167) covers the
same need unambiguously, the shorthand fails closed at the request
boundary instead of expanding:
- resolve_current_user_uri raises NamespaceShapeError with a corrective
hint naming both viking://~/<rest> and the explicit-uid form. Silently
parsing the reserved segment as a peer user id would misdirect reads
and writes, so rejection is the only safe removal.
- Bare viking://user falls through to the canonical parser and keeps
container semantics (a user key listing it sees only its own space).
- The self-id escape stays: a caller whose user_id equals a reserved
name keeps viking://user/<own-id> as their canonical root. ROOT-role
literal parsing and the legacy viking://session alias are unchanged.
- AddTargetsConfig normalizes stored legacy config spellings
(viking://user/resources|skills) to the viking://~ form at validation
so existing ov.conf/user_config deployments keep working; the accepted
per-user spelling is now viking://~/resources and viking://~/skills.
- usage_reporter keeps canonicalizing the historical shorthand found in
old transcripts and additionally recognizes viking://~/memories/.
BREAKING CHANGE: requests using the uid-less viking://user/<segment>
spelling now fail with 400; use viking://~/<segment> or an explicit
viking://user/{user_id}/<segment> URI.
* refactor(clients): migrate first-party emitters to the viking://~ home alias
Every in-repo client that emitted the removed uid-less current-user
shorthand now sends viking://~/... instead: vikingbot fallbacks and
default sentinels, the LangChain store/tools defaults, the shared
recall-core.mjs (all synced plugin copies), the codex/claude-code/
openclaw/openwebui/dsh/zcode/pi plugin emitters, quick-app examples,
Go SDK example, tau2 benchmark targets, and the eval golden dataset.
Compat kept where legacy strings live in stored user configs: bot and
ov_dream sentinels accept both spellings while emitting only ~, and
recall-core still rewrites legacy viking://user/<reserved> config values
client-side. langchain_openviking._uri now classifies viking://~ with
the explicit-user shape so canonicalized server responses keep matching
a ~ root. Plugin READMEs note the server requirement for the alias.
* docs: replace current-user shorthand guidance with the viking://~ home alias
Rewrite every EN/ZH doc and model-facing prompt that advertised the
uid-less viking://user/<segment> spelling: URI concept catalogue,
context-types/storage/extraction/retrieval/session/privacy concepts,
configuration guide (with the legacy add_targets auto-normalization
note), resources/skills/sessions/retrieval/admin API references, FAQ,
capability reference, and the openviking-memory / ov-experience-memory /
openclaw / ov-resources skills. The stale MCP viking://user/<path>
dialect passage in the MCP guide is replaced by ~ guidance, and bare
viking://user is documented as the container of user spaces.
* test(api): migrate live API session-used tests off the removed shorthand
tests/api_test/sessions sent uid-less viking://user/skills/... URIs to
record_used, which the request boundary now rejects with 400 (caught by
the API & CLI Integration Tests CI job; these tests need a live server
and are not part of the local suites). The api_test client authenticates
as an admin-role user key, so the viking://~ home alias expands for it.
tests/api_test/common/test_edge_cases.py is left as is: it asserts a 400
for a non-resource add target, which still holds.
* feat: add freshness-aware parent aggregation
Defer wide-directory abstract/overview regeneration until the configured freshness threshold is reached while continuing changed-file semantic and vector processing.
Persist freshness metadata atomically, make parent bubbling L0-aware, preserve separate semantic/vector statuses, and keep explicit waits synchronous.
Rebuild every sampled summary on threshold refresh and always retry directory vectorization so stale sidecars or transient vector failures cannot be silently accepted.
Add focused coverage for freshness policy, pending-state consumption, sampled-summary refresh, vector retries, and parent bubbling.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive, and applied to memory
* feat: ov reindex support --recursive, and applied to memory
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* perf(resource): make wait=false ingestion durable
Move source preparation and parser work behind the durable AddResource task
boundary, while keeping task-owned credentials private and short-lived.
* fix(feishu): preflight add-resource roots
* test(fs): align tree rel_path fakes with binding
* fix(vikingdb): normalize all date_time range filters in API key client
OpenViking compiles TimeRange down to the internal `range` DSL, but the
commercial VikingDB data plane (Bearer API-key auth) expects `time_range`
for date_time fields and `range` only for numeric fields. The API-key
client does not run the local engine's filter conversion, so `range`
nodes on date_time fields were sent verbatim and mis-handled.
Normalize `range` -> `time_range` for every schema date_time field by
reusing the canonical VALID_TIME_FIELDS constant, covering both
`created_at` and `updated_at` instead of hardcoding a single field name.
Numeric `range` nodes and nested boolean filter structure are preserved,
and filters already emitted as `time_range` pass through unchanged. Only
the request body `filter` is rewritten; upsert/update data is untouched.
Add regression tests covering the converted created_at/updated_at date
filters, an unchanged numeric filter, and time_range idempotency.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(vikingdb): normalize date_time filters in AK/SK client
The API-key client already rewrites `range` filter nodes on date_time
fields to VikingDB's `time_range` operator, but the AK/SK-signed
`VolcengineCollection` shares the same commercial data-plane endpoints
and had the identical latent bug: `TimeRange` expressions compile down
to the internal `range` DSL, which the commercial API only accepts for
numeric fields.
Mirror the API-key fix in `VolcengineCollection._data_post` so both
auth modes normalize `range` -> `time_range` for `created_at`/`updated_at`
while leaving numeric `range` nodes untouched. Add AK/SK coverage for
both date_time fields and for idempotency of already-`time_range` input.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(storage): keep non-memory appends free of memory trailers
ContentWriteCoordinator._write_in_place routed every append through
MemoryFileUtils, which strips the existing trailing newline and appends
a reserved MEMORY_FIELDS metadata trailer, even for resource/skill files
where MEMORY_FIELDS is not a reserved format (see content_visibility).
Append to non-memory files now concatenates raw content instead, matching
POSIX append semantics and the documented visibility rules.
* feat(mcp): add write tool with exact-string edit support
Agents could not use viking:// as a working directory through MCP: no
tool could create or update file content. Add a write tool covering full
writes (mode=replace as create-or-overwrite, append, strict create) and
targeted edits (a list of {old_string, new_string, replace_all}
exact-string replacements applied in order, all-or-nothing), following
the Write/Edit conventions of common agent harnesses.
Edits read via read_visible and write back through the content-write
coordinator, so memory metadata trailers are preserved and semantic /
vector re-indexing triggers as with any other write. Parent directories
are created automatically by the storage layer. Descriptions spell out
writable scopes (resources, user memories/resources, agent) and the
wait=true knob for read-after-write search consistency.
Also update the stale tool-count comment in app.py and the MCP tool
tables in the en/zh guides (13 -> 14 tools).
* feat(mcp): add tree tool, split targeted edits into edit tool
tree renders the recursive directory tree under a viking:// URI,
indented by depth with file sizes, for whole-layout orientation;
level_limit/node_limit bound the output and include_abstract adds
per-file summaries. Missing directories report "(nothing under ...)"
instead of an error, matching the read tool's convention.
edit(uri, old_string, new_string, replace_all) takes over the targeted
exact-string replacement that previously lived in write's edits array,
matching the classic Edit tool signature harnesses already train on.
write now only does full-content writes (content + mode), removing the
mutually-exclusive content/edits schema ambiguity. Edits still read via
read_visible and write back through the content-write coordinator, so
memory metadata trailers are preserved and re-indexing triggers as with
any other write.
* test(plugin): update canonical MCP tool list for tree/write/edit
The marketplace test pins the server-registered MCP tool list; add the
new tree, write, and edit tools to fix plugin-tests CI.
* feat(storage): support plain files at the user scope root
Agents treating viking:// as a working directory naturally drop files
like viking://user/zeus-persona.md at the user root, but the write
coordinator only accepted the memories/ and resources/ subtrees.
Two changes make that work:
- Namespace shorthand: a dotted first segment under viking://user/ is a
file name, not a user id (canonical user ids are dot-free by
convention), so viking://user/zeus-persona.md now canonicalizes to
viking://user/<current-user>/zeus-persona.md, matching how the
reserved memories/resources/skills segments already shorthand.
Dot-free segments still address an explicit user, and an exact match
with the current user id still wins.
- Coordinator: plain files directly under the user root (or in
non-managed subdirectories) anchor their semantic refresh at the
parent directory. The managed subtrees skills/, peers/, privacy/ and
sessions/ remain read-only with an actionable error message.
* fix(namespace): narrow user-root shorthand to text-file extensions
Review on #3936 (codex /review-pr) flagged that treating any dotted
segment as a user-root file shorthand would silently re-route canonical
URIs for valid dotted user ids (e.g. alice.smith) into the current
user space. Shorthand now triggers only when the first segment ends
in a common text-file extension; dotted or email-style user ids keep
resolving as canonical user ids. Adds regression tests pinning both
behaviors.
* fix(mcp): resolve user URIs against current user
* test(mcp): pin plain-file writes directly at the user root
The user-root shorthand exists so an agent can drop viking://user/persona.md
into its workspace, but every new test went through an intermediate directory
(viking://user/project/zeus-persona.md), leaving the no-directory shape — the
one that anchors the write coordinator's refresh at the user root itself —
uncovered. Add the missing case.
Also correct the write tool docstring: the create-extension allowlist applies
to any newly created file, including one created by mode="replace" falling
back to create, not only to an explicit mode="create".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(memory): support event tag filtering
Add session-level default event tags, commit-time overrides, durable queue propagation, and first-write vector index tagging. Include config update APIs and coverage for serialization, concurrency, extraction, and HTTP behavior.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(memory): expose event tags in SDKs and CLI
Add session default tag configuration, config updates, and commit-time event tag overrides across embedded Python, standalone Python, TypeScript, Go, and the Rust CLI. Preserve explicit empty-tag semantics and document each public interface.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(sdk): align legacy session tag APIs
Forward commit-time event tags through the legacy Python HTTP shims and align BaseClient session signatures without adding a new abstract-method requirement for existing subclasses.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(session): allow updating auto-commit policy
Extend PATCH session config to atomically update event tags and auto-commit settings. Merge policy objects by field, use explicit null to disable automatic commits, preserve omitted fields, and expose the contract across SDKs and CLI.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(session): align session config interfaces
Replace the generic session create config JSON flag with explicit event-tag and auto-commit options. Preserve omitted, object, and null auto-commit semantics across HTTP, embedded clients, SDKs, and CLI, reject ambiguous null policy fields, and handle nullable event configuration consistently.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test(session): trim redundant event tag tests
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
fix_fields_data backfills schema fields absent from a row's data with their
defaults, but it first short-circuited on
`len(field_data_dict) >= len(field_meta_dict)`, using field count as a proxy
for "all schema fields are present".
That proxy is wrong: a row can have as many keys as the schema (or more) while
still missing a specific field — e.g. a row written before a new field was
added that also carries an extra internal/non-schema key. In that case the
fill loop was skipped and the missing field was silently omitted rather than
defaulted, so downstream reads see an incomplete record.
Remove the count guard. The loop already skips fields that are present, so
complete inputs are returned unchanged; only genuinely missing fields are now
filled.
Add regression tests: a missing field with matching key count is backfilled
(default value and type default), and complete data is returned unchanged.
The range checks used strict inequalities (< instead of <=), rejecting
the valid boundary values defined by the WGS-84 geographic standard:
- latitude ±90 (the North and South Poles)
- longitude ±180 (the antimeridian / International Date Line)
Any resource whose geo_point is exactly on a pole or the date line cannot
be indexed, and a geo_range query centred on those boundaries raises
ValueError instead of executing the search.
Change both checks to <=. Values strictly outside the valid interval
(e.g. 181, -91) continue to raise ValueError as before.
Add a regression test covering all four boundary endpoints and all four
out-of-range coordinates.
The except block in VikingFS.mkdir() had no re-raise, so any backend
failure that was not an already-exists error (permission denied, quota
exceeded, I/O errors, lock-lease violations) — and even already-exists
errors with exist_ok=False — was silently discarded and mkdir() returned
as if the directory had been created. Callers on the write hot path
(ovpack import, parsers, session, privacy) then write into a directory
that may not exist, and the original actionable error is lost.
Re-raise the original exception unless it is an already-exists error
tolerated by exist_ok=True.
Also update tests/misc/test_mkdir.py, which still mocked fs.agfs.mkdir
even though mkdir() now goes through the AsyncAGFSClient wrapper
(self._async_agfs) — the swallowed-attribute-error made the stale tests
pass/fail for the wrong reasons. Add regression tests covering error
propagation for both exist_ok values.
_copy_dir_through_vikingfs() drives the copy phase of mv() for non-temp
directories, but enumerated the source with the agent-facing ls() default
node_limit=1000. Any directory level with more than 1000 visible entries
was copied only partially, and mv() then unconditionally deleted the
source recursively — permanently losing every entry past the cap, while
the vector index (remapped via the uncapped _collect_uris) kept pointing
at URIs that no longer exist anywhere.
Pass the module's LS_ALL_NODES sentinel, which exists precisely for
internal callers that must enumerate an entire directory.
Add a regression test that fails without the fix.
The row serializer writes each string's byte length as a UINT16 prefix. The
scalar `string` field guards this contract — a >65535-byte value raises a
clean, field-attributed ValueError. The `list<string>` element path uses the
identical UINT16 prefix but had no such guard, so an oversized element instead
raised a raw `struct.error: 'H' format requires 0 <= number <= 65535` from
deep inside struct.pack_into — a different exception type, naming no field.
Callers that catch ValueError (matching the documented scalar contract) do not
catch this, and the error gives no clue which field/element overflowed. Apply
the same bounds check to list<string> elements so inclusion in a list does not
silently downgrade the type-checked contract the scalar path upholds.
Add a regression test asserting the clean ValueError for an oversized element
and that an in-bounds (incl. multibyte) list still round-trips.