* feat(rerank): add configurable HTTP timeout for OpenAI-compatible client
OpenAIRerankClient hardcoded a 30s HTTP timeout, which is insufficient for
local LLM servers (e.g. llama.cpp on ROCm) that incur model cold-start
latency on the first request after inactivity, causing ReadTimeout errors.
Add a `timeout` field to RerankConfig (default 30.0, backwards-compatible)
and thread it through OpenAIRerankClient.__init__, from_config, and the
requests.post call in rerank_batch. The timeout can now be set per-environment
in ov.conf, e.g. "timeout": 120.
Closes#2732
* docs: document rerank timeout config
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(web-studio): move playground actions into mobile drawer
* fix(web-studio): make mobile playground actions fullscreen
* docs: add web studio mobile screenshots
* chore: clear unused files
* fix(tests): fix unit test
* refactor(auth): introduce plugin-based authentication architecture
Replace the monolithic `openviking/server/auth.py` with an extensible
plugin-based auth system. This refactor extracts the three built-in modes
(`dev`, `api_key`, `trusted`) into separate `AuthPlugin` implementations,
adds a registry for third-party plugins, and preserves all existing behavior
while enabling custom authentication backends (e.g. LDAP, OIDC, mTLS).
Key changes:
- **New public API**: `AuthPlugin` (ABC) and `register_auth_plugin` decorator.
- **New registry**: `AuthPluginRegistry` supports runtime registration.
- **Built-in plugins**: `DevAuthPlugin`, `ApiKeyAuthPlugin`, `TrustedAuthPlugin`.
- **Config change**: `auth_mode` widened from `Literal` to `str` for custom modes.
- **Validation delegated**: `validate_server_config()` now delegates to the active
plugin's `validate_config()`, preserving existing validation semantics.
- **Router compatibility**: All existing `require_*` decorators and `resolve_identity`
/ `get_request_context` dependencies remain unchanged. Routers import the same
symbols from `openviking.server.auth`.
- **Tests**: `conftest.py` manually wires the DevAuthPlugin in ASGI tests (lifespan
not triggered). `test_auth.py` expanded with plugin registration and validation tests.
- **Docs**: `04-authentication.md` (en/zh) updated with plugin registration examples.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): fix trusted mode test
* fix(tests): fix unit test
---------
Co-authored-by: claude-sonnet-4-6 <noreply@anthropic.com>
* add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
* Handle poisoned known_keys mutexes in cache providers
* Add configurable cache support to ragfs python
* Add Redis cache provider support
* Prune native cache providers from default build
* Add RAGFS cache guides in Chinese and English
* delete .cargo/config.toml
* Fix stale cache invalidation across shared wrappers
* Document Yuanrong native concurrency limits
* add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
* Handle poisoned known_keys mutexes in cache providers
* Add configurable cache support to ragfs python
* Add Redis cache provider support
* Prune native cache providers from default build
* Add RAGFS cache guides in Chinese and English
* delete .cargo/config.toml
* Fix stale cache invalidation across shared wrappers
* Document Yuanrong native concurrency limits
* Limit directory cache entries and add regression test
* docs: add TOS s3fs backend test design
* Add runtime cache config override
* add build guides for mooncake and yuanrong
* fix _FakeConfig test with cache
* fix(ragfs): cache encrypted data below the encryption layer
- pass runtime cache config to the Rust binding
- cache ciphertext instead of decrypted content
- disable cache for encrypted multi-write mounts
- update tests and provider build documentation
* fix(ragfs): invalidate cache after same-mount raw copy
Ensure copy_within_mount invalidates the cached destination file and
parent directory after the raw backend fast path writes data directly.
This prevents stale reads and stale directory metadata when cache is
enabled and Python cp() uses the same-mount copy optimization.
Add a regression test covering overwrite copy followed by read/list.
* fix(ragfs): preserve multi-write discovery through cache layer
Allow MountableFS::as_multiwrite() to unwrap CachedFileSystem when
discovering the underlying MultiWriteWrappedFS. This preserves
multi-write admin paths and same-mount copy behavior for cache-enabled,
unencrypted multi-write mounts.
Add a regression test covering sync status, sync retry, same-mount copy,
and unmount behavior for cached unencrypted multi-write mounts.
* feat(ragfs): add cache-aware tree traversal mode
- add configurable tree traversal mode to cache policy
- keep default tree behavior delegated to backend
- allow cached traversal to reuse read_dir directory cache
- bypass cached traversal for multi-write backends
- add regression coverage for tree cache behavior and fallbacks
* docs: design cache-aware grep traversal
* feat(ragfs): add cache-aware grep traversal
Introduce a shared cache traversal mode for recursive APIs and use it to
optionally run grep through CachedFileSystem.
- add CacheTraversalMode with backend and cached_traversal modes
- keep CacheTreeMode as a compatibility alias
- route tree and grep through cached traversal only when explicitly enabled
- reuse cached read_dir entries and full-file reads during grep traversal
- keep multi-write traversal on the backend path
- expose storage.agfs.cache.traversal_mode in Python config
- raise max cached directory entries threshold to 4096
- add regression tests for grep cache traversal and traversal config
* Optimize cached grep generation validation
* Parallelize cached grep file scanning
---------
Co-authored-by: fang <fang@fangMacBook-Air.local>
* Fix Codex memory hook recall noise and stop timeouts
* Tighten Codex recall compression output
* Add Codex archive resume and capture filtering
* Wrap Codex memory injection for capture filtering
* Detect Codex recall compressor profile
* Refresh Codex compressor profile on startup
* fix codex ov credential resolution
* [codex] resolve compressor profile via models_cache.json, not codex exec probe
SessionStart used to spawn 'codex exec' sequentially against each
candidate model to detect which one would respond — up to 3 probes ×
~15s timeout each, on every session start, even on resume. The
configured-on-startup default made this a guaranteed first-page-load
tax of several seconds.
Replace the probe with a lookup against codex's own model catalogue
(~/.codex/models_cache.json, refreshed by codex CLI's etag-backed
fetch). The first candidate whose slug is present wins. SessionStart
now goes cache-first: load the persisted profile if any, only resolve
on cache miss. The runtime compress path (auto-recall) deletes the
cached profile on any compress failure so the next SessionStart
re-resolves against the current catalogue.
- recall-compressor-profile.mjs:
* loadCodexModelsCache(env) reads ~/.codex/models_cache.json; missing
cache yields {present:false,slugs:Set()}.
* resolveRecallCompressorProfile picks the first available candidate
by slug; falls back optimistically to the first candidate when
the catalogue is missing.
* invalidateRecallCompressorProfileCache() rms the persisted file.
* detectRecallCompressorProfile is now cache-first and never spawns.
- auto-recall.mjs: runCodexCompressor invalidates the cache on spawn
error, timeout, non-zero exit, and read failure (best-effort,
no error surface to user).
- recall-compressor-profile.test.mjs: 11 unit tests covering catalogue
read, candidate selection (with/without configured first), missing
catalogue fallback, configured_off path, invalidate, cache-first
detect, and re-resolve after invalidate.
Notes:
- buildCodexExecArgs is still exported so auto-recall can spawn the
actual compress run; the change only removes the *probe* spawn, not
the compress spawn.
- recallCompressDetectTtlMs and recallCompressDetectTimeoutMs are
preserved in config for back-compat; the timeout no longer matters
but the TTL still bounds how stale a cached profile may be.
* [codex] omit X-OpenViking-Actor-Peer env_http_headers when no peer configured
syncMcpConfig used to unconditionally write all three OV header→env
mappings. The wrapper strips empty OPENVIKING_PEER_ID before exec'ing
codex, so an unset env var would silently flip the header to "" — the
OV side then has to disambiguate that from "no peer scope". Match the
bearer_token_env_var pattern: present only when there's something to
send. Also drops a stale X-OpenViking-Actor-Peer entry when the peer
is unset (e.g. after switching ovcli configs).
- Existing test 4 became two cases: with-peer keeps the mapping,
without-peer drops it (symmetric to bearer).
- New test asserts an in-place drop when the cached .mcp.json had a
stale peer mapping but the active config no longer has a peer.
* [codex] runtime_failed compressor marker stops same-session retry storms
Previous fix invalidated the profile cache on compress failure. Within a
single codex session that still bled `recallCompressTimeoutMs` of wall
time per UserPromptSubmit because the next hook reread cache (miss),
fell back to fallbackRecallCompressorProfile, and tried the same model.
Replace plain invalidate with a runtime_failed sentinel cached in the
profile slot. UserPromptSubmit's compressMemoryContext already short-
circuits on `profile.enabled === false`, so the marker stops further
spawns for the rest of the codex process. The next SessionStart cache-
first detect treats `source === 'runtime_failed'` as cache miss and
re-resolves against the current models_cache.json, so a transient
failure self-recovers across codex restarts without operator action.
detect_on_startup=false respects the marker (no auto-recover, matches
the "manual control" intent of that flag).
- recall-compressor-profile.mjs:
* markRecallCompressorRuntimeFailed(cfg, {failedModel}) writes the
disabled sentinel.
* detectRecallCompressorProfile branches on cached.source ===
'runtime_failed': cache hit otherwise, recover-via-resolve when
startup-detect on, respect marker when off.
- auto-recall.mjs::runCodexCompressor: swap invalidate-on-error with
markRecallCompressorRuntimeFailed(cfg, {failedModel: profile.model}).
- recall-compressor-profile.test.mjs: 4 new tests covering marker
write, cross-restart recovery picking a different slug, and the
detect_on_startup=false honor path. 20/20 pass.
invalidateRecallCompressorProfileCache is kept as a public API for
explicit operator use (e.g. a future `ov codex reset-compressor`
command), but is no longer called from the runtime path.
Promote ov_intent_analysis_sft:v7_q8 to the recommended local Ollama
query-planner model. Add the bundled retrieval.ov_intent_analysis_sft_v7
prompt and map v7_q8 to it (v4_q8 mapping kept). Update the setup wizard
presets (v7 recommended, v4 retained, v1 dropped) and the configuration
docs (EN/ZH). Extend tests to cover the v7 mapping.
Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
v0.3.24 (2026-06-05) is the latest release but the changelog top was
still v0.3.23. Adds curated Highlights + Upgrade Notes for v0.3.24 in
both EN and ZH, matching the format of #2316 / #2088.
* ci: upload source zip and plugin installers to TOS on release
Add a standalone workflow (20. Release TOS Upload) that runs on release
publish (or manual dispatch with a tag for backfill) and uploads:
- the source archive to releases/<tag>/ and releases/latest/
- both memory-plugin install.sh scripts to versioned paths and to
stable root paths for a China-reachable one-liner URL
Reuses the existing TOS secrets (AK/SK/region/endpoint) with a new
TOS_RELEASE_BUCKET secret so release artifacts stay out of the docs
bucket. Missing secrets skip gracefully (fork-friendly); real upload
failures fail the workflow.
* ci: server-side copy for the latest source zip
* feat(plugins): GitHub-free TOS install path for memory plugins
Domestic users can't reach github.com / raw.githubusercontent.com, so the
existing one-liner installers stall at their step-3 `git clone`. Add a
GitHub-free path that sources everything from Volcengine TOS:
- Both install.sh learn OPENVIKING_REPO_ARCHIVE_URL: when set, fetch the
source from a zip (curl + unzip) instead of git clone. A
.openviking-archive-source marker makes re-runs idempotent and refuses
to clobber a git checkout or unrelated data at REPO_DIR.
- New setup-helper/tos-install.sh bootstrap per plugin: sets the TOS
archive URL, downloads the real install.sh from TOS to a temp file
(kept off the stdin pipe so prompts stay interactive), and delegates.
- release-tos.yml uploads both tos-install.sh alongside install.sh.
One-liner for users behind the GFW:
bash <(curl -fsSL https://ovrelease.tos-cn-beijing.volces.com/claude-code-memory-plugin/tos-install.sh)
The GitHub default path is unchanged; archive mode only activates when
OPENVIKING_REPO_ARCHIVE_URL is set.
* docs: document the TOS (GitHub-free) install path for memory plugins
Main agent-integration docs (zh/en, claude-code + codex) keep the GitHub
one-liner and add the TOS equivalent for regions where GitHub is hard to
reach. The CDN integration cards switch their install one-liner to the TOS
bootstrap only, since that gallery is served where GitHub raw is unreliable.
* docs: trim the TOS install note to one line
* Add search retrieval telemetry breakdown
* Fix pyagfs helper annotation imports
* docs: document search relation controls and telemetry fields
Document the new include_relations request parameter and the search telemetry summary fields so the API docs stay aligned with the latest retrieval changes.
* refactor(search): drop relation enrichment and trim telemetry
Remove relation fetching from the retrieval path and delete low-value search telemetry fields so retrieval stays simpler and the telemetry summary focuses on actionable diagnostics.
The push-OTP feature (mint an OTP in Studio to hand to an MCP client) was never
wired to a consumer: consume_otp had zero production callers and no endpoint or
grant ever redeemed an OTP. The 'full happy path' test actually exercised the
display_code flow, not OTP. So the sidebar footer's 'OAuth setup' entry minted a
code with nowhere to use it — dead, confusing UX.
Remove it end-to-end and repurpose the footer slot into an entry for the
cross-device verify page (enter the 6-char display_code), which previously had no
discoverable entry point in Studio.
Frontend:
- delete oauth-setup-dialog.tsx + /oauth/setup route (+ routeTree, i18n)
- extract CrossDeviceVerifyForm from verify.tsx; add CrossDeviceVerifyDialog
- footer 'OAuth verify' entry opens the verify dialog (desktop) / page (mobile)
Backend:
- drop issue_otp route + OTPRequest/OTPResponse, storage insert_otp/consume_otp,
oauth_config.otp_ttl_seconds, and the OTP-specific tests
- keep otp.py generate_otp (cross-device display_code) + hash_secret, the shared
_atomic_consume_code, and the oauth_codes.kind column
- convert the race/expiry/revoke/GC storage tests to auth-code rows
Docs: update 11-oauth, 06-mcp-integration, and the design doc to reflect removal.
* feat(cli): add query planner setup to init wizard
Let `openviking-server init` configure the optional lightweight
query_planner model. The wizard pulls the chosen Ollama model and writes
the query_planner config; the IntentAnalyzer selects the matching prompt
at retrieval time via a model->prompt-id mapping, so no prompt files are
copied and no prompts.templates_dir override is needed.
- intent_analyzer: QUERY_PLANNER_PROMPT_BY_MODEL maps the fine-tuned SFT
models to their bundled prompt id; unmapped models keep the default
retrieval.intent_analysis prompt.
- bundle retrieval/ov_intent_analysis_sft_v4.yaml (loaded by its own id).
- ollama detection + doctor now recognize query_planner Ollama usage.
- docs: describe the init flow and runtime prompt selection.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(cli): offer query planner on all paths, recommend it with an Ollama VLM
The init wizard offers the lightweight query planner after model setup. When the
chosen setup already uses an Ollama VLM (`ollama_running` is not None) the
planner rides on that running Ollama at near-zero extra cost, so the enable
prompt is tagged "(recommended)" and defaults to yes. For cloud / non-Ollama VLM
setups it is still offered, but defaults to no and drops the recommendation;
opting in there runs the Ollama install flow.
The Ollama state established during model setup is threaded through the wizard so
the planner reuses it instead of re-running the install dialog:
- `_wizard_ollama` / `_wizard_llamacpp` return `(config, ollama_running)`.
- `run_init` forwards that state to `_wizard_query_planner`.
Docs (zh/en) and tests updated accordingly.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Adding a shell-alias name (e.g. `cc` from `alias cc=claude`) to
OPENVIKING_CC_WRAP_EXTRA / OPENVIKING_CODEX_WRAP_EXTRA broke the wrapper:
bash expands the alias mid-eval and clobbers the base `claude`/`codex`
function (so `command cc` ends up running the C compiler), while zsh
aborts with a parse error on every shell start. Guard the wrapper-defining
loop to skip names that are already shell aliases — an alias already
routes through the base wrapper once it expands, so it needs no function.
Also reject heads starting with `-`, which `alias`/`command` would
otherwise misparse as an option.
Also document the custom-launch-command feature and the alias guidance:
- 8 agent-integration docs (en/zh main + CDN cards): brief install note,
plus two troubleshooting rows (wrapper-not-sourced, alias gap)
- claude/codex plugin READMEs (+ README_CN, which was missing the section
entirely): wrap the real target command, never the alias name
* docs: refresh Claude Code & Codex memory plugin integration docs
- Fix dead anchor #1-wrap-claude-to-inject-env-from-ovcliconf -> #configuring-mcp
in the Claude Code manual setup (zh/en agent-integrations + CDN cards)
- Add the post-install wrapper activation step (source .../wrapper.sh) to the
Claude Code docs, and make Codex's activation shell-agnostic
(source ~/.zshrc -> source .../codex-memory-plugin/setup-helper/wrapper.sh)
- Rewrite Claude Code manual step 1 to the guarded `source wrapper.sh` form
- Convert Claude Code "How it works" into a lifecycle bullet list
- Language polish pass across all eight Claude Code / Codex docs
- Sync docs/images/agents/zh/index.json summaries with the refreshed intros
* docs: foolproof the verify step against an inactive wrapper
Add a guard note to the Verify section of every Claude Code / Codex doc:
if `type claude` / `type codex` prints a path instead of "shell function",
the wrapper isn't active, so re-source it (or open a new terminal) before
launching — otherwise Claude Code silently connects to 127.0.0.1 with no
auth, and Codex starts without OPENVIKING_API_KEY and reports
"MCP server is not logged in". For Claude Code, also state explicitly to
launch `claude` from the terminal where the wrapper is active.
Applies to all eight surfaces (zh/en, agent-integrations guides + CDN cards).
Keep full background add-resource tasks limited to Git repositories so anti-crawler HTTP pages are parsed by the normal importer instead of failing during early source validation.