Update the Codex OAuth setup default and examples after GPT-5.4 retirement.
Document migration for saved configurations and exercise the actual wizard
model selection in the existing setup test.
Fixes#5072
* feat(pathlock): pathlock support cache runtime
* perf(pathlock): shard Redis lock hashes by account
Route global, system, and account paths to separate HASH keys while retaining namespace slot affinity. Reject cross-scope batches and preserve lock lifecycle behavior within each scope.
* fix(pathlock): fix batch Redis stale token deletions too many results to unpack problem
* fix(test): update pathlock cache provider docs
* fix(pathlock): After the Redis lock times out, the incremental update reverts to an incorrect lock scope.
* feat(compile): let OV own external task lifecycle
Keep durable task state, recovery, query, and cancellation in OpenViking while external providers execute the workload.
* feat(compile): align external session protocol
Persist external session state in OV and adapt VikingBot to the documented create, status, and cancel contract.
* fix(compile): infer external API from base URL
Remove the redundant enable switch so generated Base Server configuration activates Compile directly from its configured endpoint.
* refactor(compile): simplify task lifecycle controls
Remove client-facing runtime and wait controls, and retire legacy OV routes in favor of the generic Task API. Bound transient status polling failures so unavailable providers fail the owned task.
* fix(compile): restore server runtime deadline
* fix(compile): enforce timeout in OV task polling
* feat(compile): expose args in CLI and SDKs
* fix(compile): correct provider retry boundaries
* fix(compile): bound cancellation convergence
* fix(compile): retry submit without runtime cap
* refactor(compile): keep provider credentials minimal
* fix(compile): poll tasks every 30 seconds by default
* refactor(compile): standardize runtime task protocol
* fix(compile): wait for remote cancellation to settle
* session: add restricted Python DSL extraction protocol and make it the default
Introduce a restricted Python memory SDK output protocol as an alternative to
the JSON extraction protocol, and switch the default to python. Both protocols
share the same ResolvedOperations post-processing, schema rules, and patch-repair
path via a new ExtractionOutputProtocol abstraction.
- Add extraction_output_protocol/{base,json,python}.py; python compiles a
restricted AST into the same operations model as json.
- Default memory.extraction_output_format flips json -> python.
- Surface the offending source line on Python syntax errors and add targeted
triple-quote retry guidance for string-literal breaks.
- Preserve every distinct fact on canonical merges; remove hardcoded memory
type names from prompts so custom memory_types render dynamically.
- Downgrade benign batch-delete link-inheritance read failures to WARNING.
- Add memory_organization A/B benchmark and message_format pretty-printer.
Tests: extraction protocol, config loader, memory react suites pass.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: drop redundant entity split hint and duplicate abstractmethod
- entities.yaml: remove the size-triggered split hint; when to split/compact is
decided at read time by memory_maintenance_notice, so the static schema
description only keeps the identity semantics and fact-preservation rule.
- vlm/base.py: remove a duplicated @abstractmethod on get_completion_async.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* vlm: drop redundant *.vlm.call span decorators for trace parity
volcengine already dropped its @tracer("volcengine.vlm.call") wrapper to avoid
duplicate spans now that the request is logged via tracer.info(llm_input_messages=...).
Remove the symmetric litellm/openai decorators so all three backends behave the same.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* benchmark/locomo: fix commit_session kwarg for CLI AsyncHTTPClient
ov.AsyncHTTPClient resolves to openviking_cli.client._http_compat.AsyncHTTPClient,
whose commit_session takes a flat telemetry= kwarg and has no options= parameter.
Passing options={...} (the SDK-client shape) raised TypeError during import.
Use telemetry=True to match the CLI client, consistent with the other locomo
import scripts.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: don't parse error-target sentinels as URIs in extraction telemetry
The by-type extraction telemetry treated result.errors[].uri as a valid viking
URI and fell back to MemoryUpdater.memory_type_from_uri(), but that field is an
error *target* — it can be a sentinel like "unknown" or "events(page_id=100)".
VikingURI() then raised 'URI must start with viking://', turning a single
recorded extraction error into a crash of the whole long_term extraction step.
Count failed errors by the known uri->type map only, defaulting to "unknown".
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* bot: pass context_type via FindOptions after SDK find/search sync
The SDK find/search sync moved context_type from a top-level find() kwarg into
FindOptions. VikingBot still called client.find(context_type='memory') for peer
recall, so every per-turn type-quota recall raised 'unexpected keyword argument
context_type' and silently returned no memories. The answer agent then fell back
to manual multi-round search (iteration ~1.3 -> ~3.9) and accuracy dropped from
~83% to ~72-77%. Pass it via options={'context_type': 'memory'} instead.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* bot: adapt VikingClient.find to SDK FindOptions for context_type/filter
The SDK find/search sync moved context_type and filter out of top-level find()
kwargs into FindOptions. VikingClient.find still forwarded them as top-level
kwargs to the SDK client, so peer memory recall raised 'unexpected keyword
argument context_type' (and after the prior partial fix, 'options') and returned
no memories — the answer agent fell back to manual multi-round search, spiking
iteration ~1.3 -> ~4 and dropping accuracy ~83% -> ~76%.
Do the SDK adaptation once in VikingClient.find (pack context_type/filter into
options={...}); callers keep the stable VikingClient.find(context_type=...)
interface, so memory.py reverts to passing context_type= directly. Verified via
a single-question smoke: type_quota recall returns 13 memories, injection is
non-empty, iteration=1, answer correct.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: route event resolution repair through the output protocol
The event resolution-repair instruction was hardcoded to demand a JSON object,
but under the default Python protocol the repaired response is parsed by the
Python SDK compiler. When a first-pass event had out-of-bounds ranges, an
assistant-only span, or an ambiguous peer, the repair round returned JSON, the
compiler rejected it as an invalid program, retries were exhausted, and the
recoverable event memory was never written.
Add ExtractionOutputProtocol.render_resolution_repair(); JSON keeps the existing
JSON-object wording, Python asks for corrected sdk.create_events(...) calls.
_build_resolution_repair_instruction now delegates to the active protocol, like
patch-repair and the final instruction already do.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session,bot,benchmark: address PR review findings
- extract_loop: document 32768 as the extraction output floor (tuned for Doubao;
lower-max models override via vlm.max_tokens) and extract
_resolve_effective_max_output_tokens; ov.conf.example notes the override.
- python_protocol: alias non-identifier memory_type/field names on the Python DSL
surface only (real names kept in URIs/storage/JSON); map aliases back when
compiling, instead of hard-rejecting kebab-case custom schemas.
- run_full_eval.sh: move auto-commit + GIT_COMMIT_ID capture AFTER arg parsing so
--auto-commit is honored and run metadata records the committed HEAD.
- litellm_vlm: strip Gemini cache_control from the already-sanitized messages so
empty assistant turns are not reintroduced; sanitize_openai_messages passes
through non-dict entries.
- Tests for each fix.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: reject Python DSL alias collisions instead of silently overwriting
_identifier_alias() is not one-to-one: memory_type 'project-notes' and
'project_notes' (or fields 'note-body'/'note_body') fold to the same DSL alias.
The alias->real dict comprehensions would silently drop one, making a schema/
field unreachable and routing writes to the wrong target. Add
_validate_alias_uniqueness(), invoked in render_contract() and the compiler
__init__ (so parse() paths without render are also guarded), which fails loudly
with a rename hint. Distinct-identifier names never collide, so real configs are
unaffected. Tests cover type, field, and no-render parse collisions.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: bound str.replace result size before allocation in Python DSL
The _MAX_EXPRESSION_SIZE guard covered * (repeat) and + (concat) but not the
whitelisted string methods: only join() had a projected-size check, so
('x'*1000).replace('x','y'*10000) could still allocate a >1MB result and bypass
the limit. Add _check_replace_size() that bounds source + occurrences*(len(new)
-len(old)) BEFORE calling str.replace (which builds the whole result in C), so
the oversized string is never allocated. replace is the only whitelisted method
that can materially inflate output (join already guarded; upper/lower/strip/
split/startswith/endswith do not grow). Tests cover an oversized replace being
rejected and a normal replace passing.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: close f-string width and str %-format inflation in Python DSL
The replace guard alone was insufficient: f-string format specs (f"{'x':>1000001}")
and str %-formatting ("%1000001s" % "x") also turn a small integer literal into an
arbitrarily large string with no repeat operator, bypassing _MAX_EXPRESSION_SIZE.
Neither has a legitimate use in memory content, so disallow them outright rather
than bounding width inflation: reject any non-empty f-string format spec and reject
str/bytes %-formatting (numeric % still allowed). Combined with the existing
*/+/join/replace pre-allocation checks, all small-input->large-output amplifiers
are now closed. Tests cover f-string width rejection, plain f-string, and str %.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* fix(rerank): validate credentials for explicitly configured providers
RerankConfig checked required fields for openai and litellm only. An explicit
`provider: cohere` without api_key, or `provider: vikingdb` without ak/sk, was
accepted at load time. is_available() then returned False and
HierarchicalRetriever fell back to plain vector search, logging a single
info-level line saying rerank was not configured.
The new checks run against the effective provider, matching the existing
openai and litellm branches. Auto-detection is unaffected, since detecting
cohere already requires api_key and detecting vikingdb already requires ak and
sk. An empty RerankConfig() still resolves to no provider and stays valid.
Drops test_default_provider_is_vikingdb, which asserted a default that
auto-detection replaced and had been failing on main. Rewrites
test_unknown_provider_raises_value_error to actually cover an unknown provider
and adds coverage for the two providers that were missing validation.
* docs(configuration): state required credentials per rerank provider
The rerank section described credential inference but not the fields each
provider requires when provider is set explicitly.
---------
Co-authored-by: Terminator666666 <Terminator666666@users.noreply.github.com>
* fix: directory understanding_api & fix temp empty file
* fix: understanding_api error处理
* fix: api err msg
* fix: html test
* fix: support directory and HTML imports via UnderstandingAPI
* fix: support directory and HTML imports via UnderstandingAPI
* fix: director max files and zip msg
* fix: director max files and zip msg
* feat: support vector record IDs across filesystem APIs
- centralize deterministic vector record ID generation and migration handling
- allow stat and read to resolve file IDs with actionable missing-index diagnostics
- add selectable filesystem fields and script-friendly ls, tree, and glob output
- preserve complete IDs in simple output and honor tree --simple without fields
- restrict Python SDK ID passthrough to supported read-only endpoints
- preserve explicit empty glob extra_fields for metadata responses
- retain the dedicated Codex OAuth doctor diagnostic path
- document the public API behavior and add CLI, SDK, storage, and server regressions
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* fix(cli): preserve full record IDs in field tables
Render filesystem record IDs without abbreviation in both table and simple field modes so the values can be passed directly to stat and read. Add a regression test for normal table rendering.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: Maojia Sheng <shengmaojia@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* docs: add anydoc office converter design
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(parse): add AnydocConfig and anydoc adapter skeleton
Register anydoc parser config with firecrawl-anydoc dependency and a small
attribute adapter for binding name normalization ahead of converter wiring.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(parse): add anydoc conversion core
Serialize anydoc documents to GFM while preserving embedded images through the existing storage media pipeline.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(parse): address Task 2 review findings
Move inline-code backslash escaping out of the f-string expression for
Python 3.10/3.11 compatibility, and stop tracking the SDD task report
in the product tree.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(parse): wire Word and legacy Doc to anydoc
Route Word and real OLE documents through the shared converter while preserving configurable legacy fallbacks and OOXML disguise handling.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(parse): honor anydoc config and safe fallbacks
Wire application parser settings into the default registry and prevent unsupported ODT/RTF files from reaching python-docx.
* feat(parse): wire PowerPoint and EPUB to anydoc
Route supported presentation and EPUB formats through the shared converter while preserving safe format-specific legacy fallbacks.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(parse): wire Excel parser to anydoc
Route modern spreadsheet formats through anydoc with safe row truncation while preserving legacy process-pool conversion when anydoc is disabled.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(parse): preserve Excel ingestion options
Forward directory parse settings through the anydoc path and make skipped row truncation observable.
* docs(parse): document anydoc Office support
Reflect the expanded Office and EPUB format coverage while keeping PDF behavior explicitly unchanged.
* fix(parse): address final anydoc review findings
Resolve signatureless CSV conversion and preserve ingestion options across Office parsers while documenting and testing XLSB behavior.
* feat: unify office parsing with anydoc
* refactor: simplify anydoc renderer organization
* fix(parse): align anydoc parser config switch
* fix(anydoc): preserve legacy parser compatibility
* fix(anydoc): restore legacy safeguards
* refactor(anydoc): keep Office parsing on the unified path
* fix(anydoc): preserve config and benchmark compatibility
* fix(markdown): isolate link rewrite state per parse
---------
Co-authored-by: 张剑锋 <zhangjianfeng@ydjdev.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(cli): download directories as zip archives
* fix(download): cap directory archives
* fix(download): bound directory archives while they are built
The archive size cap was only enforced by `os.path.getsize()` after the
whole ZIP had been written, and `actual_total` counts file payload bytes
only. A tree made of empty directories or empty files therefore adds
per-entry ZIP headers that no check sees until the temp file is already
complete: with the limit set to 1 KiB, a 20k-entry tree writes 1.9 MB to
disk before being rejected.
Check the live write offset after every member so the temp archive stays
within the limit as it grows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CneCiyRjLaeDRYSWKngcJ
* docs(download): make the KG snippet re-runnable and sync the API catalog
The new knowledge-graph snippet extracts with a plain `unzip`, but the
note above it only tells the reader to delete the archive. Re-running it
leaves the previously extracted `./journal-kg/` in place, so `unzip`
stops at an overwrite prompt — and in a non-interactive shell it exits 1
without extracting anything. Use `unzip -o` and say what the note
actually has to cover.
Also update the endpoint catalog in api/01-overview.md, which still
described /content/download as file-bytes only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CneCiyRjLaeDRYSWKngcJ
* fix(download): build directory archives in memory, not in a temp file
The temp archive is handed to FileResponse with a BackgroundTask that
unlinks it, but starlette runs `background` only after a successful
send. Both Range-header error branches (starlette/responses.py:370,373)
`return await PlainTextResponse(...)` before reaching it, so a malformed
or unsatisfiable Range leaks the archive permanently — 22 such requests
leak 22 files in a local repro, up to 10 MiB each, with nothing to
reclaim them. asyncio.CancelledError misses the `except Exception`
cleanup for the same reason.
Since the archive is capped at 10 MiB anyway, build it in a BytesIO and
return it as a plain Response, exactly like the single-file branch. That
drops the temp file, the cleanup callback, and the tempfile/os/
FileResponse/BackgroundTask imports, and gives both branches the same
`Content-Disposition: attachment; filename*=UTF-8''...` form instead of
two different ones.
Directory downloads no longer honour Range. They never usefully did:
the archive is rebuilt per request and zipfile stamps time.localtime()
into every member, so resuming a range spliced two different archives.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CneCiyRjLaeDRYSWKngcJ
* fix(download): return 413 for oversized directory archives
RESOURCE_EXHAUSTED maps to 429, which tells clients the request is
rate-limited and worth retrying after a backoff. An archive over the
10 MiB cap fails because of the directory's own size, so every retry
re-walks the tree and re-zips it before failing again.
Add PAYLOAD_TOO_LARGE / 413 and raise it from the archive size check.
The code is plumbed through both status<->code maps (server app and
utils), the client's code->exception table, and the Rust CLI's status
mapping, so an over-cap `ov get` still surfaces a typed error rather
than falling through to INTERNAL.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CneCiyRjLaeDRYSWKngcJ
* feat(cli): write directory downloads as a named .zip in a target directory
`ov get viking://resources/myfolder ./myfolder` wrote the ZIP bytes to a
path named `myfolder` with no suffix: a regular file wearing a folder's
name, which `cd` rejects and `file` reports as ZIP data. The local path
was always used verbatim, so only the docs' hard-coded `./project.zip`
form produced a sane result.
Treat a target that is an existing directory — or omitted, meaning the
current directory — as the destination *directory*, and name the file
after the resource, appending `.zip` when the response came back as
`application/zip`. An explicit non-directory path is still used
verbatim, so `ov get <uri> ./explicit.zip` is unchanged. Nothing is
extracted; the archive is what lands.
get_bytes_with_type exposes the response Content-Type, which is how the
caller tells a raw file apart from a directory served as a ZIP;
get_bytes keeps its old signature for the TUI and its existing test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CneCiyRjLaeDRYSWKngcJ
* fix(download): bound archive entries and preflight targets
* fix(cli): preflight existing symlink targets
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Allow openai-codex VLM configuration to carry reasoning_effort through credential normalization and into Responses API requests so deployments can tune model effort without out-of-tree patches.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(uri)!: reject uid-less current-user shorthand in favor of viking://~
viking://user/<segment> (memories/resources/skills/peers/privacy/sessions
without a user id) was ambiguous with a user literally named after the
segment, and a user actually named e.g. "memories" was unreachable for
USER/ADMIN callers. Now that the viking://~ home alias (#4167) covers the
same need unambiguously, the shorthand fails closed at the request
boundary instead of expanding:
- resolve_current_user_uri raises NamespaceShapeError with a corrective
hint naming both viking://~/<rest> and the explicit-uid form. Silently
parsing the reserved segment as a peer user id would misdirect reads
and writes, so rejection is the only safe removal.
- Bare viking://user falls through to the canonical parser and keeps
container semantics (a user key listing it sees only its own space).
- The self-id escape stays: a caller whose user_id equals a reserved
name keeps viking://user/<own-id> as their canonical root. ROOT-role
literal parsing and the legacy viking://session alias are unchanged.
- AddTargetsConfig normalizes stored legacy config spellings
(viking://user/resources|skills) to the viking://~ form at validation
so existing ov.conf/user_config deployments keep working; the accepted
per-user spelling is now viking://~/resources and viking://~/skills.
- usage_reporter keeps canonicalizing the historical shorthand found in
old transcripts and additionally recognizes viking://~/memories/.
BREAKING CHANGE: requests using the uid-less viking://user/<segment>
spelling now fail with 400; use viking://~/<segment> or an explicit
viking://user/{user_id}/<segment> URI.
* refactor(clients): migrate first-party emitters to the viking://~ home alias
Every in-repo client that emitted the removed uid-less current-user
shorthand now sends viking://~/... instead: vikingbot fallbacks and
default sentinels, the LangChain store/tools defaults, the shared
recall-core.mjs (all synced plugin copies), the codex/claude-code/
openclaw/openwebui/dsh/zcode/pi plugin emitters, quick-app examples,
Go SDK example, tau2 benchmark targets, and the eval golden dataset.
Compat kept where legacy strings live in stored user configs: bot and
ov_dream sentinels accept both spellings while emitting only ~, and
recall-core still rewrites legacy viking://user/<reserved> config values
client-side. langchain_openviking._uri now classifies viking://~ with
the explicit-user shape so canonicalized server responses keep matching
a ~ root. Plugin READMEs note the server requirement for the alias.
* docs: replace current-user shorthand guidance with the viking://~ home alias
Rewrite every EN/ZH doc and model-facing prompt that advertised the
uid-less viking://user/<segment> spelling: URI concept catalogue,
context-types/storage/extraction/retrieval/session/privacy concepts,
configuration guide (with the legacy add_targets auto-normalization
note), resources/skills/sessions/retrieval/admin API references, FAQ,
capability reference, and the openviking-memory / ov-experience-memory /
openclaw / ov-resources skills. The stale MCP viking://user/<path>
dialect passage in the MCP guide is replaced by ~ guidance, and bare
viking://user is documented as the container of user spaces.
* test(api): migrate live API session-used tests off the removed shorthand
tests/api_test/sessions sent uid-less viking://user/skills/... URIs to
record_used, which the request boundary now rejects with 400 (caught by
the API & CLI Integration Tests CI job; these tests need a live server
and are not part of the local suites). The api_test client authenticates
as an admin-role user key, so the viking://~ home alias expands for it.
tests/api_test/common/test_edge_cases.py is left as is: it asserts a 400
for a non-resource add target, which still holds.
* feat: add freshness-aware parent aggregation
Defer wide-directory abstract/overview regeneration until the configured freshness threshold is reached while continuing changed-file semantic and vector processing.
Persist freshness metadata atomically, make parent bubbling L0-aware, preserve separate semantic/vector statuses, and keep explicit waits synchronous.
Rebuild every sampled summary on threshold refresh and always retry directory vectorization so stale sidecars or transient vector failures cannot be silently accepted.
Add focused coverage for freshness policy, pending-state consumption, sampled-summary refresh, vector retries, and parent bubbling.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive
* feat: ov reindex support --recursive, and applied to memory
* feat: ov reindex support --recursive, and applied to memory
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat(uri): add viking://~ home alias for the caller's user root
Accept `viking://~` (and `viking://~/<suffix>`) as a server-side alias
for the authenticated caller's user namespace root. The alias is
expanded at the request boundary by resolve_current_user_uri for
USER/ADMIN identities, so every control plane that already funnels
through validate_request_viking_uri (REST, MCP, and therefore CLI/SDK
clients) gets it with no client changes.
Design points:
- `~` is a reserved token that can never collide with a real user id
(validate_user_id's charset excludes it), so segment-0 aliasing does
not weaken canonical-first parsing.
- The canonical parser (resolve_uri) rejects the alias outright,
mirroring the legacy-session handling: root-role requests, internal
callers, and storage paths fail closed instead of materializing a
literal '~' directory.
- The alias is accepted but never advertised: scope error copy
("Must be one of: ...") filters it at both the parser and the public
validator, and VikingURI.build refuses to mint it, so responses and
persisted data stay canonical.
- MCP search now resolves exclude_uris with the same strictness as the
REST search router (closes the one entry point that skipped it).
* refactor(mcp): drop tilde mention from tool docstrings
Agents pass viking://~ through verbatim and responses echo the
canonical form, so the alias is self-explanatory on contact; carrying
the sentence in 12 tool descriptions costs every MCP session tokens.
Docs keep the alias documented for humans.
- VolcEngine selection now opens an access-tier submenu: Agent Plan
(api/plan/v3), Coding Plan (api/coding/v3), or pay-as-you-go API, with
doubao-seed-2.0-lite / doubao-embedding-vision defaults for the plans
- BytePlus gains the same submenu with ModelArk Coding Plan
(api/coding/v3, dola-seed-2.0-lite / skylark-embedding-vision) and
pay-as-you-go
- Embedding 'Other (manual)' becomes a Custom/manual submenu offering
interactive OpenAI-compatible prompts (URL, key, model, dimension)
alongside the editor hand-off
- Custom VLM/embedding prompts accept !back to return to the provider
menu; plan submenus support back navigation and seed defaults from the
existing config endpoint (unknown volces.com/bytepluses.com endpoints
are offered as a keep-current option)
- Two-step flow seeds the VLM plan menu from the embedding tier choice;
summaries now show the api_base so tiers are distinguishable
- Fix provider re-detection for BytePlus plan endpoints (domain-based,
since VolcEngine and BytePlus share the volcengine provider string)
- Document the plan endpoints in ov.conf.example
The official npm bin (ov.mjs) starts with #!/usr/bin/env node. The Python
entry-point fallback skipped every shebang file on PATH, so the Node wrapper
was never exec'd and users got 'ov binary not found'.
Only compare against and skip the current Python entry point itself; other
PATH executables (including the Node shebang wrapper) are exec'd normally, and
exec failures are no longer swallowed.
* feat(pdf): refactor MinerU parsing to the official file_parse API
* feat(pdf): remove mineru_api_key from configuration and examples
* feat(pdf): preflight MinerU /health during service initialization
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
Frontmatter was parsed into ParseResult.meta and removed from the body, but
that metadata is never persisted, so every ingested markdown file lost its
frontmatter with no way to read the fields back.
Parse frontmatter into meta unconditionally (it still drives doc_title) and
only remove it from the stored body when explicitly configured; that removal
is now off by default.
* feat(memory): support event tag filtering
Add session-level default event tags, commit-time overrides, durable queue propagation, and first-write vector index tagging. Include config update APIs and coverage for serialization, concurrency, extraction, and HTTP behavior.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(memory): expose event tags in SDKs and CLI
Add session default tag configuration, config updates, and commit-time event tag overrides across embedded Python, standalone Python, TypeScript, Go, and the Rust CLI. Preserve explicit empty-tag semantics and document each public interface.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(sdk): align legacy session tag APIs
Forward commit-time event tags through the legacy Python HTTP shims and align BaseClient session signatures without adding a new abstract-method requirement for existing subclasses.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat(session): allow updating auto-commit policy
Extend PATCH session config to atomically update event tags and auto-commit settings. Merge policy objects by field, use explicit null to disable automatic commits, preserve omitted fields, and expose the contract across SDKs and CLI.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(session): align session config interfaces
Replace the generic session create config JSON flag with explicit event-tag and auto-commit options. Preserve omitted, object, and null auto-commit semantics across HTTP, embedded clients, SDKs, and CLI, reject ambiguous null policy fields, and handle nullable event configuration consistently.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test(session): trim redundant event tag tests
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix: remove heima partner, clean up auth docs, add web-studio unsupported auth banner
- Remove heima from partner list in README (en/zh/ja)
- Remove unsupported env var references (OPENVIKING_AUTH_MODE, OPENVIKING_USERNAME,
OPENVIKING_PASSWORD) from LDAP auth docs
- Remove temporary switch bash snippets from auth docs
- Fix ldap_password description
- Add web-studio unsupported-auth-mode banner for oidc/ldap servers
* fix: address OIDC/LDAP review comments on auth plugin design
Key changes driven by PR review:
- **Role mapping**: OIDC and LDAP external identities always resolve to
USER role. Removed map_role() calls and group_membership-based role
mapping. Admin access is gated by the root API key mechanism only.
- **LDAP credential extraction**: Removed query-parameter-based username/
password extraction (security concern — passwords in URLs can leak via
shell history, proxy logs, and monitoring). Clients must use Basic Auth
header or form data.
- **OIDC identifier sanitization**: Auth0 and other providers may include
characters like "|" in the `sub` claim. These are now replaced with "_"
to produce valid OpenViking user identifiers.
- **Dead code removal**: Removed _extract_groups, memberof_attribute,
require_root_api_key_for_admin, _initialize_api_key_manager, and
get_request_context_checks from both plugins since they are no longer
needed.
- **Docs**: Removed query-parameter curl example, memberof_attribute and
require_root_api_key_for_admin config references.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix(auth): bind lazy OIDC imports at module scope
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"
Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.
This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.
New assembly kernel under openviking/retrieve/context_assembler/:
- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
readable floor, then overview, then full for high-scoring entries. An
oversized tier falls back to the previous one instead of being truncated,
bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
Summary section, code files reuse code_outline signatures, long documents use
a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
directories carry no stored abstract; their full tier stays capped at
overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
every harness inherits cross-turn dedup; exclude_uris remains as the
stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
element per entry. Every tier carries its URI, so the model can always drill
down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
timeout fuses, and a failed rewrite still returns the unrewritten block.
Retrieval failures are counted into stats rather than silently yielding an
empty block.
Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.
* refactor(retrieval): give context tiers a per-category default
The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.
Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.
Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".
Assembly fixes found alongside:
- Removing the abstract cap exemption would turn an oversized abstract into
a bare URI, so it now falls back to overview first — for memory that is a
cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
`asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
such attribute; usage was structurally always null. It now reads the model
instance's tracker and reports only when the call count moved by exactly
one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
fail, and the file was never rewritten, so it could not heal. Records are
now coerced on read and dropped on the next write, along with records left
ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
`viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
resolved a different profile than `POST /recall`; an unknown `detail` value
raised `KeyError` through the whole call instead of degrading.
* feat(codex): inject profile context on session start
Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): raise rewrite timeout default to 30s
* docs(agents): document low-latency recall settings
* fix(codex): prefer luna as recall compressor fallback
* refactor(plugins): unify recall compression setting
* feat(plugins): enable recall compression by default
* docs(agents): use absolute links in image docs
* fix(retrieval): address context assembly review feedback
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test: trim redundant context assembly coverage
* fix(retrieval): address second-round context assembly review
- Drop the backticked `/search` from the deprecated-recall row in both API
overviews. The reference checker scans the whole row after the method cell
for backticked paths, so it read the description as a route named
`POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
belongs to the Rust CLI, which writes `root_api_key`, `output`,
`echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
two Python readers had drifted into stricter subsets, so the shipped example
already failed to load in both. Adding the new `plugin` section to a working
ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
Retrieval validates query and image_url before searching, and the gather
fuse swallowed that rejection along with genuine scope failures, so a body
of `{"mode":"context"}` came back 200 with an empty block instead of the
documented parameter error. Runtime failures still degrade into
`stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
than the 30s fuse, so a rewrite that finished inside its own budget was
aborted client-side, discarding the whole response — including the
uncompressed block the server returns when a rewrite fails — and falling
back to `/recall`. The deadline is only extended when the body actually
requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
`plugin.recallContextTimeoutMs` pins it.
* chore(plugins): sync shared modules into the zcode snapshot
* fix(retrieval): align context quotas and plugin defaults
Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): preserve recall compatibility
Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
Surface the vector store's search_tags on each matched context (returned
under the "tags" key to match the tags filter param) and remove the
result fields the retrieval pipeline never populates (category,
match_reason, relations, overview).
Co-authored-by: TRAE CLI <noreply@bytedance.com>