* fix(rerank): support DashScope nested request/response envelope
OpenAIRerankClient sent a flat request body ({"model", "query",
"documents"}) and parsed "results" at the top level of the response.
DashScope (qwen3-rerank) requires a nested envelope:
Request: {"model", "input": {"query", "documents"}, "parameters": ...}
Response: {"output": {"results": [...]}, "request_id", "usage"}
This caused DashScope rerank to silently fail — the response had no
top-level "results" key, so the client returned None.
Changes:
- Add _is_dashscope() to detect DashScope endpoints by host marker.
- Add _build_request_body() that produces the nested envelope for
DashScope and the flat body for standard OpenAI/Cohere services.
- Add _extract_results() that reads output.results for DashScope and
top-level results for standard services.
- Accept both "relevance_score" (singular, DashScope) and
"relevance_scores" (plural, some providers) in result items.
- Add 13 tests covering host detection, body construction, response
parsing, end-to-end mocked flows for both providers, plural key
handling, empty documents, and sparse results.
Fixes#3459
* fix(rerank): detect DashScope protocol by URL path, not hostname
Reviewer noted the previous hostname-based switch broke the documented
qwen3-rerank compatible-api endpoint (/compatible-api/v1/reranks), which
must use the flat OpenAI-style body and top-level results.
Switch to path-based detection: only /api/v1/services/rerank uses the
native nested input/output envelope; everything else (including the
DashScope compatible-api and generic OpenAI/Cohere gateways) keeps the
flat protocol. Rename _is_dashscope -> _uses_nested_envelope for clarity.
Add regression tests covering the compatible-api flat path and reconcile
the existing native-path fixtures to the nested envelope.
* docs(rerank): use qwen3-rerank for compatible-api example
The compatible-api/v1/reranks endpoint uses the flat OpenAI-compatible
protocol; qwen3-vl-rerank is a native-envelope model served at
/api/v1/services/rerank. Align the example model with the endpoint the
implementation selects by URL path.
---------
Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
* fix(storage): block deleting write-protected viking://agent root
The delete-namespace guard rejected viking:// and viking://user but not the
account-shared viking://agent root that the sibling write guard already
forbids, so a non-root user could rm(viking://agent, recursive=True) and
recursively wipe account-wide agent skills/endpoints/tools/payments. Mirror
the write guard's bare-agent-root rejection; concrete sub-paths stay deletable.
Parity with #2873.
* test(storage): cover delete rejection of bare viking://agent root
#2873 added a delete guard so rm("viking://") / rm("viking://user") raise
PermissionDenied before any side effect. mv() is the sister destructive op
(copy + recursive rm of the source) but only guards the source via the write
guard _ensure_mutable_access, which does not reject the bare "viking://"
account root. So a normal user calling mv from the bare root (HTTP filesystem,
WebDAV MOVE, CLI) reaches the recursive source delete the guard was meant to
forbid.
Apply _ensure_delete_access to the mv source (it is recursively rm'd), keeping
the write guard on the destination. The guard is additive, so it can only
reject more (the protected bare root), never loosen the existing
write-namespace checks; concrete subtree URIs used by internal callers
(watch_manager, semantic_processor) are unaffected.
Adds a parametrized regression test mirroring the existing rm protected-root
test: bare-root mv sources now raise before any AGFS stat/rm.
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
* Handle poisoned known_keys mutexes in cache providers
* Add configurable cache support to ragfs python
* Add Redis cache provider support
* Prune native cache providers from default build
* Add RAGFS cache guides in Chinese and English
* delete .cargo/config.toml
* Fix stale cache invalidation across shared wrappers
* Document Yuanrong native concurrency limits
* add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
* Handle poisoned known_keys mutexes in cache providers
* Add configurable cache support to ragfs python
* Add Redis cache provider support
* Prune native cache providers from default build
* Add RAGFS cache guides in Chinese and English
* delete .cargo/config.toml
* Fix stale cache invalidation across shared wrappers
* Document Yuanrong native concurrency limits
* Limit directory cache entries and add regression test
* docs: add TOS s3fs backend test design
* Add runtime cache config override
* add build guides for mooncake and yuanrong
* fix _FakeConfig test with cache
* fix(ragfs): cache encrypted data below the encryption layer
- pass runtime cache config to the Rust binding
- cache ciphertext instead of decrypted content
- disable cache for encrypted multi-write mounts
- update tests and provider build documentation
* fix(ragfs): invalidate cache after same-mount raw copy
Ensure copy_within_mount invalidates the cached destination file and
parent directory after the raw backend fast path writes data directly.
This prevents stale reads and stale directory metadata when cache is
enabled and Python cp() uses the same-mount copy optimization.
Add a regression test covering overwrite copy followed by read/list.
* fix(ragfs): preserve multi-write discovery through cache layer
Allow MountableFS::as_multiwrite() to unwrap CachedFileSystem when
discovering the underlying MultiWriteWrappedFS. This preserves
multi-write admin paths and same-mount copy behavior for cache-enabled,
unencrypted multi-write mounts.
Add a regression test covering sync status, sync retry, same-mount copy,
and unmount behavior for cached unencrypted multi-write mounts.
* feat(ragfs): add cache-aware tree traversal mode
- add configurable tree traversal mode to cache policy
- keep default tree behavior delegated to backend
- allow cached traversal to reuse read_dir directory cache
- bypass cached traversal for multi-write backends
- add regression coverage for tree cache behavior and fallbacks
* docs: design cache-aware grep traversal
* feat(ragfs): add cache-aware grep traversal
Introduce a shared cache traversal mode for recursive APIs and use it to
optionally run grep through CachedFileSystem.
- add CacheTraversalMode with backend and cached_traversal modes
- keep CacheTreeMode as a compatibility alias
- route tree and grep through cached traversal only when explicitly enabled
- reuse cached read_dir entries and full-file reads during grep traversal
- keep multi-write traversal on the backend path
- expose storage.agfs.cache.traversal_mode in Python config
- raise max cached directory entries threshold to 4096
- add regression tests for grep cache traversal and traversal config
* Optimize cached grep generation validation
* Parallelize cached grep file scanning
---------
Co-authored-by: fang <fang@fangMacBook-Air.local>
Cache async SDK clients per event loop to avoid cross-loop reuse in worker threads.
Move memory vectorization into semantic queue refresh and preserve target sync state for resource updates.
The OpenViking docker image still launched the legacy `openviking/console`
standalone service on port 8020. Now that web-studio is bundled into the OV
server itself at /studio (see #2156), that process is redundant and the
port is just a confusing artefact.
This change retires the old console (python package + 8020 + console-frontend
favicons) but **keeps the in-compose Caddy as a stable single-ingress on
port 1934**, just simplified to one upstream now that there's no 8020. The
server-side BFF at `openviking/server/routers/console.py` (under
`/api/v1/console/*`) is also kept — web-studio uses the same endpoints.
**The OAuth authorize page (`openviking/server/oauth/router.py`) is
deliberately untouched in this PR** — the console-link button and Quick
authorize same-origin panel will be re-pointed at web-studio in a focused
follow-up.
BREAKING CHANGES:
- Port 8020 is gone from the docker image and docker-compose.yml; Caddy at
1934 now forwards everything to 1933 (web-studio lives at /studio there).
Anything bookmarked at `http://host:8020/...` must migrate to
`http://host:1933/studio/`.
- `python -m openviking.console.bootstrap` no longer exists; the python
package `openviking.console` has been removed.
Pip packaging:
- web-studio dist is now shipped inside the wheel under
`openviking/web_studio/dist/` (mirroring the old `openviking/console/static/`
layout). The dockerfile copies `--from=web-studio-builder /web-studio/dist`
into the source tree before `uv sync`, so the wheel produced by the
default docker build always carries the SPA. Building the wheel without
running `npm run build` first leaves the directory empty, which gracefully
degrades /studio to a 404 without breaking server startup.
- Favicon assets (`favicon.ico` / `favicon-32.png` / `apple-touch-icon.png`,
~11 KB total) are duplicated into `openviking/server/static/` and shipped
via package-data so `/favicon.*` and `/mcp/favicon.*` routes are always
registered, regardless of whether the web-studio dist is bundled.
- `pyproject.toml` and `setup.py` `package-data` drop `console/static/**`
and add `server/static/**` + `web_studio/dist/**`.
- New favicons (the 16/32/180 set in both `openviking/server/static/` and
`web-studio/public/`) are downscaled from the canonical
`web-studio/public/openviking-icon.png`, so the small-icon family matches
the SPA's high-res rel="icon" target — the studio tab icon now stays
consistent whether the browser uses the HTML link tag or falls back to
auto-fetching `/favicon.ico`.
Server:
- `openviking/server/app.py` now reads `/studio` from
`Path(__file__).parent.parent / 'web_studio' / 'dist'` by default;
`OPENVIKING_WEB_STUDIO_DIR` still wins for dev mode pointing at a
repo-local build. Favicon routes are unconditionally registered and
load from `openviking/server/static/`.
- `openviking/observability/usage_audit/projection.py` drops the legacy
`/console/*` skip prefix (the BFF prefix `/api/v1/console/*` remains).
Docker:
- `web-studio-builder` stage moved earlier (Stage 2) so its dist can flow
into `py-builder` before `uv sync` runs.
- Runtime stage no longer separately copies the dist or sets
`OPENVIKING_WEB_STUDIO_DIR`; the in-package path is the default.
- Entrypoint renamed `openviking-console-entrypoint.sh` -> `openviking-entrypoint.sh`
and stripped of the `python -m openviking.console.bootstrap` launch.
- `EXPOSE 1933 8020` -> `EXPOSE 1933`.
- `docker-compose.yml` drops the openviking service's 8020 port mapping;
the caddy service stays but no longer needs port 8020 exposed.
- `Caddyfile` simplified to a single `:1934 { reverse_proxy openviking:1933 }`
— the legacy `/console/*` route to :8020 is gone.
Docs:
- en/zh quickstart updated to drop the 8020 mapping and explain that the
API server now also serves `/studio`.
- Other guides (`12-public-access.md`, `11-oauth.md`, `05-observability.md`,
`04-setup-for-agent.md`, `03-deployment.md`) are intentionally left for a
focused follow-up PR alongside the OAuth quick-authorize reintroduction.
Tests:
- Deleted `tests/misc/test_console_{proxy,static_assets}.py` (covered the
removed console package). `tests/observability/test_console_router.py`
stays — it covers the BFF, which remains.
* feat: ov add-resource (spec -L --level), ov stat (return count for dir)
* feat: ov add-resource (spec -L --level), ov stat (return count for dir)
* feat: Add VLM backup configuration for automatic failover
- Add backup field to VLMConfig with recursive backup prevention
- Implement FailoverVLM wrapper class for automatic failover
- Support rate limit, timeout, server error triggers
- Add comprehensive unit tests
* feat: Add VLM backup configuration for automatic failover
* feat: Add VLM backup configuration for automatic failover
* ov observer filesystem
* ov observer filesystem
* chore(format): align python and c++ file formatting
* chore: update urllib3 to 2.7.0 and clean test imports
1. bump urllib3 dependency from 2.6.3 to 2.7.0
2. remove unused pytest import and RoleScope import from test file
* style: format list comprehensions and lambda function for readability
Adjust the line breaks in the list comprehension in the VikingSearchTool class to follow standard Python formatting conventions, and rewrap the lambda assignment in the test case to improve code readability without changing functionality.
* style: fix line wrapping and remove extra blank line
- remove stray blank line in ov_server.py
- wrap long logger.info line in memory.py for better readability
* style: fix targeted ruff lint violations
* chore: clean up unused imports and reorder code
This commit removes unused imports, reorders import statements for better consistency,
and simplifies some test file imports. Changes include:
- Remove redundant blank lines and unused imports across multiple test files and core modules
- Reorder imports in openviking hooks module to follow standard layout
- Fix import ordering in memory isolation handler
- Simplify php parser type imports
- Move volcengine mock import to correct position in test file
* refactor(uri utils): remove extra blank lines in uri.py
clean up redundant whitespace to improve code readability
* fix(parse): preserve source_name when zip single root differs
ZipParser collapsed any zip whose top level held a single directory and
unconditionally dropped source_name. For CLI add-resource on a directory
whose only child differs in name (e.g. user uploads "access/" containing
just "access-dns/"), this silently dropped the user-named parent layer:
the resource landed at "<parent>/access-dns" instead of
"<parent>/access/access-dns".
Only collapse when source_name is absent or its stem matches the single
root dir name (the legacy "tt_b.zip" wrapping "tt_b/" case). When
source_name names a distinct outer layer, parse the extract root and
keep source_name so DirectoryParser preserves it as dir_name.
* fix(parse): compare basename before stem in zip single-root collapse
Path(source_name).stem drops trailing dotted segments, so source_name
"v1.2" against a zip whose single root is "v1.2/" would fail the
equality check and re-wrap the content, producing
"viking://resources/v1.2/v1.2/...". Compare the leaf name first so
dotted names match exactly, and keep the stem fallback for the
"tt_b.zip" wrapping "tt_b/" case.
Also adds end-to-end regression tests that drive ZipParser +
DirectoryParser + TreeBuilder.finalize_from_temp against an in-memory
VikingFS and assert the final root_uri, so future changes in
DirectoryParser or TreeBuilder can't silently drop the outer layer.
---------
Co-authored-by: jinze <yanzexu.yzx@antgroup.com>
* feat(ovpack): add v2 manifest and conflict policy
Add a portable OVPack manifest for scalar metadata and make imports validate scope, derived files, and conflicts before writing.
* fix(ovpack): remove import vectorize option
Make OVPack imports always rebuild vectors in the target environment, keep legacy packages compatible, and reject unsupported manifest versions before writing.
* fix(ovpack): remove force import alias
Use on_conflict as the single OVPack import conflict policy and reject removed force inputs.
* fix(ovpack): regenerate runtime vector metadata
Keep type portable but stop exporting or applying created_at, updated_at, and active_count from OVPack manifests.
* fix(ovpack): validate manifest contents
* fix(ovpack): require manifests for imports
* fix(ovpack): close manifest validation gaps
* fix(ovpack): defer parent creation until validation passes
* fix(ovpack): remove export size guard
* fix(ovpack): support session and scope-root restores
* docs(ovpack): document full backup migration
* feat(ovpack): add backup restore workflow
* fix(ovpack): validate import scope compatibility
* feat(tracker): persist task tracker state across instances
fix: default task tracker to memory backend
* refactor: scope persistent task paths by user
* perf(storage): optimize VikingFS grep implementation
- add ripgrep support and native local grep fallback for LocalFS mode
- normalize grep path handling and improve regression test coverage
* perf(storage): optimize VikingFS grep implementation
- add ripgrep support and native local grep fallback for LocalFS mode
- normalize grep path handling and improve regression test coverage
* perf(storage): format code
* perf(storage): support async grep
* feat(encryption): Push `exclude_uri` and `level_limit` down to the backend to ensure that the limits take effect after the same set of filtering conditions. && The default implementation also follows the same path contract as LocalFS, returning the path relative to the query root
* feat(encryption): Push `exclude_uri` and `level_limit` down to the backend to ensure that the limits take effect after the same set of filtering conditions. && The default implementation also follows the same path contract as LocalFS, returning the path relative to the query root
* fix(storage): relativize localfs exclude_path against grep query root && disable parent .gitignore inheritance in localfs rg fast path && add HTTP AGFS grep support for exclude_path and level_limit
* feat(agfs): support overriding queuefs db_path via config
Add a new optional `storage.agfs.queue_db_path` field to AGFSConfig so
the queuefs sqlite database can be relocated outside the workspace via
config. This helps when the workspace volume does not support sqlite
(e.g. some network filesystems returning disk I/O errors on PRAGMA/WAL),
allowing the queue db to be placed on a local path like `/tmp/queue.db`.
When not set, the default path `{storage.workspace}/_system/queue/queue.db`
is preserved, keeping backwards compatibility.
* fix(ragfs): load only abi3 binding artifacts
Avoid stale cpython-specific ragfs_python artifacts shadowing rebuilt abi3 extensions so queuefs sqlite persistence uses the current Rust binding.
---------
Co-authored-by: dingben.db@bytedance.com <dingben.db@bytedance.com@bytedance.com>