Treat messages.jsonl as the materialization boundary for session-aware recall,
repair partial session roots during the existing authoritative append path,
and preserve Claude capture cursors when writes never reach the server.
Also replay explicitly retryable storage conflicts across memory plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): honor tier ceilings and stop cooling unserved recalls
Follow-up to #3534, from its post-merge review round.
- The abstract-to-overview substitute now applies only to categories whose
stored abstract is the whole file body. A resource or skill whose abstract is
missing (`processing_mode=vectors_only`) or over the per-entry cap read its
body and returned an overview instead, which for a short file is the body
almost verbatim — crossing the opt-in deepening boundary those categories are
documented to have, and doing it even under an explicit `detail="abstract"`.
They now degrade to a bare URI and their body is never read.
- A digest reporting `no_relevant` blanks `rendered`, so the client injects
nothing, yet those URIs still entered the dedup ledger and were cooled for
`dedup_turns` turns. That contradicted the ledger's own bare-URI grace rule
and held memories back from the later turn they were relevant to.
- Flat retrieval reaches built-in memory types outside the four named ones
(`cases`, `patterns`, `tools`, `trajectories`, skill-usage memories) and
reported them as an undeclared `memories` category that no tier or penalty
table covered, so other-peer hits skipped the score penalty and callers could
not pin their tier. The catch-all is now a declared category with both; it
stays out of `quotas`, whose buckets it would overlap. Skill-usage memories
also stop being misread as the `skills` category.
- ZCode, OpenCode and pi own an OV session id but did not forward it, so their
recalls silently ran without query expansion or cross-turn dedup.
- The context-request deadline covered only the server's 30s rewrite fuse, but
the pipeline is serial: expansion, retrieval and budgeting all precede it.
45s covers both fuses and the work between them.
- `plugin` config scope and the `/recall` successor example now match what the
code actually does.
* fix(retrieval): make the context deadline and expansion opt-out reachable
Forwarding a session id turns on server-side query expansion, an LLM call with
its own 5s fuse, but neither the deadline that was supposed to cover it nor the
switch that turns it off reached the two harnesses this PR newly enabled it for.
- `contextRequestTimeoutMs()` now derives the deadline from the request body
rather than from `cfg` plus a rewrite flag. The body is what states which
server stages will run: a session takes the expansion fuse, `rewrite` takes
the digest fuse, and a bare retrieval takes neither and keeps the caller's own
budget. Reading `cfg` alone could not tell those apart.
- OpenCode pinned `timeoutMs: 5000` after spreading the helper's options and pi
ignored them entirely, so the helper's deadline was dead code in both. Their
own budgets are now defaults rather than ceilings. OpenCode's 5s in particular
was shorter than the expansion fuse it had just enabled, so a legal request
would have been aborted client-side and dropped back to the path with neither
dedup nor expansion.
- OpenCode and pi read `OPENVIKING_RECALL_QUERY_EXPANSION` (and
`recallQueryExpansion` in their own config files) and set the `configured`
flag the shared body builder requires, so the documented opt-out exists where
the cost was introduced.
- The integration overview no longer implies every harness reads the same
environment knobs, and describes the deadline as per-stage rather than
rewrite-only.
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"
Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.
This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.
New assembly kernel under openviking/retrieve/context_assembler/:
- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
readable floor, then overview, then full for high-scoring entries. An
oversized tier falls back to the previous one instead of being truncated,
bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
Summary section, code files reuse code_outline signatures, long documents use
a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
directories carry no stored abstract; their full tier stays capped at
overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
every harness inherits cross-turn dedup; exclude_uris remains as the
stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
element per entry. Every tier carries its URI, so the model can always drill
down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
timeout fuses, and a failed rewrite still returns the unrewritten block.
Retrieval failures are counted into stats rather than silently yielding an
empty block.
Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.
* refactor(retrieval): give context tiers a per-category default
The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.
Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.
Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".
Assembly fixes found alongside:
- Removing the abstract cap exemption would turn an oversized abstract into
a bare URI, so it now falls back to overview first — for memory that is a
cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
`asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
such attribute; usage was structurally always null. It now reads the model
instance's tracker and reports only when the call count moved by exactly
one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
fail, and the file was never rewritten, so it could not heal. Records are
now coerced on read and dropped on the next write, along with records left
ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
`viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
resolved a different profile than `POST /recall`; an unknown `detail` value
raised `KeyError` through the whole call instead of degrading.
* feat(codex): inject profile context on session start
Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): raise rewrite timeout default to 30s
* docs(agents): document low-latency recall settings
* fix(codex): prefer luna as recall compressor fallback
* refactor(plugins): unify recall compression setting
* feat(plugins): enable recall compression by default
* docs(agents): use absolute links in image docs
* fix(retrieval): address context assembly review feedback
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test: trim redundant context assembly coverage
* fix(retrieval): address second-round context assembly review
- Drop the backticked `/search` from the deprecated-recall row in both API
overviews. The reference checker scans the whole row after the method cell
for backticked paths, so it read the description as a route named
`POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
belongs to the Rust CLI, which writes `root_api_key`, `output`,
`echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
two Python readers had drifted into stricter subsets, so the shipped example
already failed to load in both. Adding the new `plugin` section to a working
ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
Retrieval validates query and image_url before searching, and the gather
fuse swallowed that rejection along with genuine scope failures, so a body
of `{"mode":"context"}` came back 200 with an empty block instead of the
documented parameter error. Runtime failures still degrade into
`stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
than the 30s fuse, so a rewrite that finished inside its own budget was
aborted client-side, discarding the whole response — including the
uncompressed block the server returns when a rewrite fails — and falling
back to `/recall`. The deadline is only extended when the body actually
requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
`plugin.recallContextTimeoutMs` pins it.
* chore(plugins): sync shared modules into the zcode snapshot
* fix(retrieval): align context quotas and plugin defaults
Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): preserve recall compatibility
Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
Surface the vector store's search_tags on each matched context (returned
under the "tags" key to match the tags filter param) and remove the
result fields the retrieval pipeline never populates (category,
match_reason, relations, overview).
Co-authored-by: TRAE CLI <noreply@bytedance.com>
The bytes_row STRING type uses a uint16 length prefix, capping a single
field at 65535 bytes. Add a new TEXT field type (enum value 9) that mirrors
STRING semantics (utf-8 str round-trip) but uses a uint32 length prefix,
lifting the per-field limit to ~4GB.
TEXT is added only at the physical bytes_row layer, across all serializers
that must stay byte-identical: the C++ engine (bytes_row.h/.cpp), the abi3
boundary (abi3_engine_backend.cpp, decoding to str not bytes), the pure
Python fallback (store/bytes_row.py), and the engine API (_python_api.py).
Existing types and the CandidateData.fields field are untouched, so old
on-disk data stays readable without reindex.
Fields opt into the new type via metadata={"field_type": FieldType.text}.
Add TestTextFieldType covering >65535-byte round-trips, py<->cpp cross
read/write, binary consistency, and metadata-based declaration.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: add audio and video understanding via VLM
* docs: design media resource guards
* fix: bound media staging concurrency
* fix: cap unknown-size media staging
* test: stage media in routing fake
* test: exercise media staging callbacks
* test: trim media understanding coverage
* chore: 清理实现计划文档
* fix: 修复多凭证切换问题
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
* fix(server): stop exporting raw query strings and buffering zip responses in observability
Sweep findings: B-03, B-13. Prevent query secrets from reaching traces and keep ZIP responses streaming.
(cherry picked from commit d8ac3dc33b)
* fix(session): tolerate missing/corrupt archive in Phase-2 replay
(NotFoundError / _ArchiveMessagesCorruptError) on a missing or corrupt
archive messages.jsonl instead of returning []. That PR added skip-on-
missing tolerance to the read path (_get_uncovered_archive_messages) and to
resume_queued_commit, but not to the Phase-2 commit replay path
(_prepare_phase2_archive_messages), which calls _read_archive_messages
unguarded while rolling earlier failed archives into the current commit.
Consequence: a terminally-failed earlier archive whose messages.jsonl is
missing/corrupt (legacy "no messages" terminal data, or produced by #3417's
own archive_read terminal path) makes every subsequent commit's Phase-2
extraction raise -> caught by _run_memory_extraction's except -> the current
archive is terminal-failed too. Because the poisoned archive is only removed
from replay once "covered" (which requires a later archive to complete), and
no later archive can ever complete, the session's memory extraction is
permanently poisoned. Raw messages are safe, but extraction is stuck.
Fix: wrap the replay-loop _read_archive_messages call in the same tolerance
_get_uncovered_archive_messages already uses -- skip + warn on not-found
(_is_storage_not_found) and on _ArchiveMessagesCorruptError, re-raise real
storage failures. The skipped archive stays in covered_failed so the current
archive's .done marks it covered, clearing the poison permanently.
Adds a regression test asserting the replay skips a failed archive with a
missing messages.jsonl (and marks it covered) instead of raising, and that a
real storage failure still propagates.
Follow-up to #3417.
(cherry picked from commit 5b8ec9e68a)
* fix(client): align client surfaces without leaking memory metadata
Reconstructs the client-parity work from upstream PR #3439 on current main and strips reserved memory metadata before line slicing in both embedded and HTTP reads.
Based-on: 48b411d58c
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(index): propagate semantic vectorization failures safely
Reconstructs upstream PR #3437 on current main, carries enqueue failures through SemanticDagExecutor, and drains the attempt's embedding tracker before retry-visible failure propagation.
Based-on: 02387deb09
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(core): close privacy and embedding failure gaps
* fix(memory): strip repeated metadata trailers
* fix(core): close public memory visibility gaps
* ci: skip embedding-dependent resource test without secrets
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(ragfs): preserve cache visibility on partial S3 deletes
Surface exact and per-object S3 deletion failures, while always invalidating the affected directory and stat cache scope after a recursive delete attempt.
Source-PR: #3407
Original-Commit: 8d6addf28e
* fix(session): preserve legacy policy and peer identity compatibility
Parse string false and other legacy boolean-like memory policy values without silently enabling extraction or breaking persisted configs. Encode mixed-script peers losslessly, while retaining their former lossy IDs as read-only retrieval and extraction aliases.
Source-PR: #3422
Original-Commit: 0dfd5a9ed9
* fix(memory): drain timer flush tasks during shutdown
Retain the shielded timer flush task and await it when close cancels the timer loop, so batch failures are observed and submitters are resolved without unhandled task exceptions.
Source-PR: #3438
Original-Commit: ca1d74e164
* fix(storage): preserve peer isolation and cache correctness
* fix(ingest): reserve encoded peer namespace
* ci: skip embedding-dependent resource test without secrets
* fix(ragfs): invalidate caches after partial remove
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
Print-to-PDF producers routinely emit several image XObjects drawn at the
exact same position on a page (a background layer plus a content layer).
Because `_extract_image_from_page` rasterises the page *region* rather than
decoding the XObject itself, every one of them renders to identical bytes —
so a document with two stacked full-page layers wrote two byte-identical
PNGs per page and referenced both from the generated markdown.
Dedup within each page, in two steps:
- bbox first, so a repeat is skipped before paying for the render;
- a content hash as a backstop, for bboxes that differ slightly but still
rasterise to the same bytes.
Both sets are per-page, so a header logo repeated across pages is still
kept once on every page. `meta["images_deduplicated"]` reports how many
were skipped.
Measured on an 8-page article exported from a web page: 16 saved PNGs -> 8,
16 markdown image references -> 8, local conversion 3.8s -> 2.4s.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
- add storage.agfs.pathlock.lock_timeout_secs
- use pathlock default timeout instead of hardcoded zero in wrapper
- map legacy storage.transaction.lock_timeout when new config is unset
- remote redolog by using persistent `session_commit` queue.
* feat: implement server-resolved OpenViking Assets manifests
Add the openviking-assets/1 declaration flow with server-owned configuration parsing and native Rust CLI execution.
- Resolve one flat Manifest against one Catalog through an authenticated server endpoint with strict schema and Git semantic validation.
- Reject recursive includes and unsafe clone URLs; return a resolved plan without submitting resources or running server-side batches.
- Keep local credential aliases, manifest state, dry-run, failure isolation, and per-asset create/sync orchestration in the CLI.
- Generate normalized stable asset identities on the server and remove the CLI direct SHA-1 dependency.
- Update flat examples and add server resolver/API plus Rust CLI coverage.
* feat: implement server-resolved OpenViking Assets manifests
* feat: implement server-resolved OpenViking Assets manifests
* fix(pathlock): tolerate missing lock token after recursive delete
* feat: implement server-resolved OpenViking Assets manifests
* feat: implement server-resolved OpenViking Assets manifests
Give with_openviking_context a deterministic lifecycle owner, reuse adapter clients without sharing invocation state, and preserve loop-scoped async behavior. Make component copies lifecycle-safe and reject post-close use before history access.
Copy connection and retriever configuration without cloning caller-owned live clients or carrying owned client caches into copied adapters. Guard optional Pydantic private state before resetting compatibility caches.
* feat(agent-evolution): reload global switch at commit time
* feat(agent-evolution): expose configured account in status
* test(agent-evolution): cover account in status response
* fix(agent-evolution): align live config reload semantics
* fix(agent-evolution): tolerate non-object live config
* fix(usage-reporter): use snake case count fields
* feat(usage-reporter): add file log sink
* fix
* fix: address live reload and usage sink review findings
* fix(usage-reporter): complete file sink compatibility
* fix: make experience snapshot source unambiguous
* docs(usage-reporter): align count record implementation plan
* fix: address agent evolution review blockers
* fix(usage-reporter): preserve Windows rollover deadline
* fix(usage-reporter): encode file records as JSON envelopes
* fix(usage-reporter): use snake case unique id
* fix(parse): distinguish mpegts from TypeScript ts
* fix(parse): tighten mpegts ts routing semantics
* fix(semantic): use file name for media summary type
---------
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
* fix(langchain): make async recording concurrency-safe
* fix(langchain): scope async state to each invocation
* fix(langchain): preserve cancellation progress on Python 3.10