* fix(vikingdb): normalize all date_time range filters in API key client
OpenViking compiles TimeRange down to the internal `range` DSL, but the
commercial VikingDB data plane (Bearer API-key auth) expects `time_range`
for date_time fields and `range` only for numeric fields. The API-key
client does not run the local engine's filter conversion, so `range`
nodes on date_time fields were sent verbatim and mis-handled.
Normalize `range` -> `time_range` for every schema date_time field by
reusing the canonical VALID_TIME_FIELDS constant, covering both
`created_at` and `updated_at` instead of hardcoding a single field name.
Numeric `range` nodes and nested boolean filter structure are preserved,
and filters already emitted as `time_range` pass through unchanged. Only
the request body `filter` is rewritten; upsert/update data is untouched.
Add regression tests covering the converted created_at/updated_at date
filters, an unchanged numeric filter, and time_range idempotency.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(vikingdb): normalize date_time filters in AK/SK client
The API-key client already rewrites `range` filter nodes on date_time
fields to VikingDB's `time_range` operator, but the AK/SK-signed
`VolcengineCollection` shares the same commercial data-plane endpoints
and had the identical latent bug: `TimeRange` expressions compile down
to the internal `range` DSL, which the commercial API only accepts for
numeric fields.
Mirror the API-key fix in `VolcengineCollection._data_post` so both
auth modes normalize `range` -> `time_range` for `created_at`/`updated_at`
while leaving numeric `range` nodes untouched. Add AK/SK coverage for
both date_time fields and for idempotency of already-`time_range` input.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(storage): keep non-memory appends free of memory trailers
ContentWriteCoordinator._write_in_place routed every append through
MemoryFileUtils, which strips the existing trailing newline and appends
a reserved MEMORY_FIELDS metadata trailer, even for resource/skill files
where MEMORY_FIELDS is not a reserved format (see content_visibility).
Append to non-memory files now concatenates raw content instead, matching
POSIX append semantics and the documented visibility rules.
* feat(mcp): add write tool with exact-string edit support
Agents could not use viking:// as a working directory through MCP: no
tool could create or update file content. Add a write tool covering full
writes (mode=replace as create-or-overwrite, append, strict create) and
targeted edits (a list of {old_string, new_string, replace_all}
exact-string replacements applied in order, all-or-nothing), following
the Write/Edit conventions of common agent harnesses.
Edits read via read_visible and write back through the content-write
coordinator, so memory metadata trailers are preserved and semantic /
vector re-indexing triggers as with any other write. Parent directories
are created automatically by the storage layer. Descriptions spell out
writable scopes (resources, user memories/resources, agent) and the
wait=true knob for read-after-write search consistency.
Also update the stale tool-count comment in app.py and the MCP tool
tables in the en/zh guides (13 -> 14 tools).
* feat(mcp): add tree tool, split targeted edits into edit tool
tree renders the recursive directory tree under a viking:// URI,
indented by depth with file sizes, for whole-layout orientation;
level_limit/node_limit bound the output and include_abstract adds
per-file summaries. Missing directories report "(nothing under ...)"
instead of an error, matching the read tool's convention.
edit(uri, old_string, new_string, replace_all) takes over the targeted
exact-string replacement that previously lived in write's edits array,
matching the classic Edit tool signature harnesses already train on.
write now only does full-content writes (content + mode), removing the
mutually-exclusive content/edits schema ambiguity. Edits still read via
read_visible and write back through the content-write coordinator, so
memory metadata trailers are preserved and re-indexing triggers as with
any other write.
* test(plugin): update canonical MCP tool list for tree/write/edit
The marketplace test pins the server-registered MCP tool list; add the
new tree, write, and edit tools to fix plugin-tests CI.
* feat(storage): support plain files at the user scope root
Agents treating viking:// as a working directory naturally drop files
like viking://user/zeus-persona.md at the user root, but the write
coordinator only accepted the memories/ and resources/ subtrees.
Two changes make that work:
- Namespace shorthand: a dotted first segment under viking://user/ is a
file name, not a user id (canonical user ids are dot-free by
convention), so viking://user/zeus-persona.md now canonicalizes to
viking://user/<current-user>/zeus-persona.md, matching how the
reserved memories/resources/skills segments already shorthand.
Dot-free segments still address an explicit user, and an exact match
with the current user id still wins.
- Coordinator: plain files directly under the user root (or in
non-managed subdirectories) anchor their semantic refresh at the
parent directory. The managed subtrees skills/, peers/, privacy/ and
sessions/ remain read-only with an actionable error message.
* fix(namespace): narrow user-root shorthand to text-file extensions
Review on #3936 (codex /review-pr) flagged that treating any dotted
segment as a user-root file shorthand would silently re-route canonical
URIs for valid dotted user ids (e.g. alice.smith) into the current
user space. Shorthand now triggers only when the first segment ends
in a common text-file extension; dotted or email-style user ids keep
resolving as canonical user ids. Adds regression tests pinning both
behaviors.
* fix(mcp): resolve user URIs against current user
* test(mcp): pin plain-file writes directly at the user root
The user-root shorthand exists so an agent can drop viking://user/persona.md
into its workspace, but every new test went through an intermediate directory
(viking://user/project/zeus-persona.md), leaving the no-directory shape — the
one that anchors the write coordinator's refresh at the user root itself —
uncovered. Add the missing case.
Also correct the write tool docstring: the create-extension allowlist applies
to any newly created file, including one created by mode="replace" falling
back to create, not only to an explicit mode="create".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(memory): support event tag filtering
Add session-level default event tags, commit-time overrides, durable queue propagation, and first-write vector index tagging. Include config update APIs and coverage for serialization, concurrency, extraction, and HTTP behavior.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(memory): expose event tags in SDKs and CLI
Add session default tag configuration, config updates, and commit-time event tag overrides across embedded Python, standalone Python, TypeScript, Go, and the Rust CLI. Preserve explicit empty-tag semantics and document each public interface.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(sdk): align legacy session tag APIs
Forward commit-time event tags through the legacy Python HTTP shims and align BaseClient session signatures without adding a new abstract-method requirement for existing subclasses.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(session): allow updating auto-commit policy
Extend PATCH session config to atomically update event tags and auto-commit settings. Merge policy objects by field, use explicit null to disable automatic commits, preserve omitted fields, and expose the contract across SDKs and CLI.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(session): align session config interfaces
Replace the generic session create config JSON flag with explicit event-tag and auto-commit options. Preserve omitted, object, and null auto-commit semantics across HTTP, embedded clients, SDKs, and CLI, reject ambiguous null policy fields, and handle nullable event configuration consistently.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test(session): trim redundant event tag tests
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
fix_fields_data backfills schema fields absent from a row's data with their
defaults, but it first short-circuited on
`len(field_data_dict) >= len(field_meta_dict)`, using field count as a proxy
for "all schema fields are present".
That proxy is wrong: a row can have as many keys as the schema (or more) while
still missing a specific field — e.g. a row written before a new field was
added that also carries an extra internal/non-schema key. In that case the
fill loop was skipped and the missing field was silently omitted rather than
defaulted, so downstream reads see an incomplete record.
Remove the count guard. The loop already skips fields that are present, so
complete inputs are returned unchanged; only genuinely missing fields are now
filled.
Add regression tests: a missing field with matching key count is backfilled
(default value and type default), and complete data is returned unchanged.
The range checks used strict inequalities (< instead of <=), rejecting
the valid boundary values defined by the WGS-84 geographic standard:
- latitude ±90 (the North and South Poles)
- longitude ±180 (the antimeridian / International Date Line)
Any resource whose geo_point is exactly on a pole or the date line cannot
be indexed, and a geo_range query centred on those boundaries raises
ValueError instead of executing the search.
Change both checks to <=. Values strictly outside the valid interval
(e.g. 181, -91) continue to raise ValueError as before.
Add a regression test covering all four boundary endpoints and all four
out-of-range coordinates.
The except block in VikingFS.mkdir() had no re-raise, so any backend
failure that was not an already-exists error (permission denied, quota
exceeded, I/O errors, lock-lease violations) — and even already-exists
errors with exist_ok=False — was silently discarded and mkdir() returned
as if the directory had been created. Callers on the write hot path
(ovpack import, parsers, session, privacy) then write into a directory
that may not exist, and the original actionable error is lost.
Re-raise the original exception unless it is an already-exists error
tolerated by exist_ok=True.
Also update tests/misc/test_mkdir.py, which still mocked fs.agfs.mkdir
even though mkdir() now goes through the AsyncAGFSClient wrapper
(self._async_agfs) — the swallowed-attribute-error made the stale tests
pass/fail for the wrong reasons. Add regression tests covering error
propagation for both exist_ok values.
_copy_dir_through_vikingfs() drives the copy phase of mv() for non-temp
directories, but enumerated the source with the agent-facing ls() default
node_limit=1000. Any directory level with more than 1000 visible entries
was copied only partially, and mv() then unconditionally deleted the
source recursively — permanently losing every entry past the cap, while
the vector index (remapped via the uncapped _collect_uris) kept pointing
at URIs that no longer exist anywhere.
Pass the module's LS_ALL_NODES sentinel, which exists precisely for
internal callers that must enumerate an entire directory.
Add a regression test that fails without the fix.
The row serializer writes each string's byte length as a UINT16 prefix. The
scalar `string` field guards this contract — a >65535-byte value raises a
clean, field-attributed ValueError. The `list<string>` element path uses the
identical UINT16 prefix but had no such guard, so an oversized element instead
raised a raw `struct.error: 'H' format requires 0 <= number <= 65535` from
deep inside struct.pack_into — a different exception type, naming no field.
Callers that catch ValueError (matching the documented scalar contract) do not
catch this, and the error gives no clue which field/element overflowed. Apply
the same bounds check to list<string> elements so inclusion in a list does not
silently downgrade the type-checked contract the scalar path upholds.
Add a regression test asserting the clean ValueError for an oversized element
and that an in-bounds (incl. multibyte) list still round-trips.
The bytes_row STRING type uses a uint16 length prefix, capping a single
field at 65535 bytes. Add a new TEXT field type (enum value 9) that mirrors
STRING semantics (utf-8 str round-trip) but uses a uint32 length prefix,
lifting the per-field limit to ~4GB.
TEXT is added only at the physical bytes_row layer, across all serializers
that must stay byte-identical: the C++ engine (bytes_row.h/.cpp), the abi3
boundary (abi3_engine_backend.cpp, decoding to str not bytes), the pure
Python fallback (store/bytes_row.py), and the engine API (_python_api.py).
Existing types and the CandidateData.fields field are untouched, so old
on-disk data stays readable without reindex.
Fields opt into the new type via metadata={"field_type": FieldType.text}.
Add TestTextFieldType covering >65535-byte round-trips, py<->cpp cross
read/write, binary consistency, and metadata-based declaration.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Route the query() random-sampling branch through search_by_vector with a
client-generated random vector (config.embedding.dimension) so every
backend behaves consistently, instead of each backend's server-side
search_by_random. This also makes Qdrant/OpenGauss truly random rather
than a deterministic scroll/scan.
Drop the raise_on_error path entirely per request: query(), the
Collection wrapper, and HttpCollection.search_by_random no longer take
raise_on_error, and delete() no longer requests it. As a result, HTTP
filter-based deletion id lookups now return empty on non-200 instead of
raising.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: add audio and video understanding via VLM
* docs: design media resource guards
* fix: bound media staging concurrency
* fix: cap unknown-size media staging
* test: stage media in routing fake
* test: exercise media staging callbacks
* test: trim media understanding coverage
* chore: 清理实现计划文档
* fix: 修复多凭证切换问题
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
* fix(server): stop exporting raw query strings and buffering zip responses in observability
Sweep findings: B-03, B-13. Prevent query secrets from reaching traces and keep ZIP responses streaming.
(cherry picked from commit d8ac3dc33b)
* fix(session): tolerate missing/corrupt archive in Phase-2 replay
(NotFoundError / _ArchiveMessagesCorruptError) on a missing or corrupt
archive messages.jsonl instead of returning []. That PR added skip-on-
missing tolerance to the read path (_get_uncovered_archive_messages) and to
resume_queued_commit, but not to the Phase-2 commit replay path
(_prepare_phase2_archive_messages), which calls _read_archive_messages
unguarded while rolling earlier failed archives into the current commit.
Consequence: a terminally-failed earlier archive whose messages.jsonl is
missing/corrupt (legacy "no messages" terminal data, or produced by #3417's
own archive_read terminal path) makes every subsequent commit's Phase-2
extraction raise -> caught by _run_memory_extraction's except -> the current
archive is terminal-failed too. Because the poisoned archive is only removed
from replay once "covered" (which requires a later archive to complete), and
no later archive can ever complete, the session's memory extraction is
permanently poisoned. Raw messages are safe, but extraction is stuck.
Fix: wrap the replay-loop _read_archive_messages call in the same tolerance
_get_uncovered_archive_messages already uses -- skip + warn on not-found
(_is_storage_not_found) and on _ArchiveMessagesCorruptError, re-raise real
storage failures. The skipped archive stays in covered_failed so the current
archive's .done marks it covered, clearing the poison permanently.
Adds a regression test asserting the replay skips a failed archive with a
missing messages.jsonl (and marks it covered) instead of raising, and that a
real storage failure still propagates.
Follow-up to #3417.
(cherry picked from commit 5b8ec9e68a)
* fix(client): align client surfaces without leaking memory metadata
Reconstructs the client-parity work from upstream PR #3439 on current main and strips reserved memory metadata before line slicing in both embedded and HTTP reads.
Based-on: 48b411d58c
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(index): propagate semantic vectorization failures safely
Reconstructs upstream PR #3437 on current main, carries enqueue failures through SemanticDagExecutor, and drains the attempt's embedding tracker before retry-visible failure propagation.
Based-on: 02387deb09
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(core): close privacy and embedding failure gaps
* fix(memory): strip repeated metadata trailers
* fix(core): close public memory visibility gaps
* ci: skip embedding-dependent resource test without secrets
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
- add storage.agfs.pathlock.lock_timeout_secs
- use pathlock default timeout instead of hardcoded zero in wrapper
- map legacy storage.transaction.lock_timeout when new config is unset
- remote redolog by using persistent `session_commit` queue.