The except block in VikingFS.mkdir() had no re-raise, so any backend
failure that was not an already-exists error (permission denied, quota
exceeded, I/O errors, lock-lease violations) — and even already-exists
errors with exist_ok=False — was silently discarded and mkdir() returned
as if the directory had been created. Callers on the write hot path
(ovpack import, parsers, session, privacy) then write into a directory
that may not exist, and the original actionable error is lost.
Re-raise the original exception unless it is an already-exists error
tolerated by exist_ok=True.
Also update tests/misc/test_mkdir.py, which still mocked fs.agfs.mkdir
even though mkdir() now goes through the AsyncAGFSClient wrapper
(self._async_agfs) — the swallowed-attribute-error made the stale tests
pass/fail for the wrong reasons. Add regression tests covering error
propagation for both exist_ok values.
- add storage.agfs.pathlock.lock_timeout_secs
- use pathlock default timeout instead of hardcoded zero in wrapper
- map legacy storage.transaction.lock_timeout when new config is unset
- remote redolog by using persistent `session_commit` queue.
* fix(parse): distinguish mpegts from TypeScript ts
* fix(parse): tighten mpegts ts routing semantics
* fix(semantic): use file name for media summary type
---------
Co-authored-by: chenxiaobin.monkey <chenxiaobin.monkey@bytedance.com>
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* docs: fix broken links and anchors across READMEs and guides
Sweep findings: D-10, D-11, D-12, D-13, D-14, D-15. Restore valid documentation targets and stable cross-page anchors.
(cherry picked from commit e3504d633d)
* docs: correct contributor and release references
Reconstruct the factual parts of draft #3397 against current upstream: use the supported setup wizard, align the repository tree and workflow names with tracked files, document current release paths, and repair the bug-bounty link. Excludes install-policy and subjective content rewrites.
Based-on: b332e19e40
Based-on: c89afb17f2
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* build: propagate recipe failures and align CMake minimum
Keep build failures visible, use isolated temporary extraction paths, and enforce the native build's CMake 3.15 floor across all contributor guides. CMake version parsing accepts prerelease and vendor suffixes.
Based-on: 8943a12285
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(scripts): surface backfill enumeration failures
Preserve the safety fix from draft #3415 while retaining legacy no-op arguments for existing operational scripts. Deprecated arguments now remain parse-compatible, advertise their status in help, and emit explicit warnings when used.
Based-on: 5516d96048
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(rerank): support DashScope nested request/response envelope
OpenAIRerankClient sent a flat request body ({"model", "query",
"documents"}) and parsed "results" at the top level of the response.
DashScope (qwen3-rerank) requires a nested envelope:
Request: {"model", "input": {"query", "documents"}, "parameters": ...}
Response: {"output": {"results": [...]}, "request_id", "usage"}
This caused DashScope rerank to silently fail — the response had no
top-level "results" key, so the client returned None.
Changes:
- Add _is_dashscope() to detect DashScope endpoints by host marker.
- Add _build_request_body() that produces the nested envelope for
DashScope and the flat body for standard OpenAI/Cohere services.
- Add _extract_results() that reads output.results for DashScope and
top-level results for standard services.
- Accept both "relevance_score" (singular, DashScope) and
"relevance_scores" (plural, some providers) in result items.
- Add 13 tests covering host detection, body construction, response
parsing, end-to-end mocked flows for both providers, plural key
handling, empty documents, and sparse results.
Fixes#3459
* fix(rerank): detect DashScope protocol by URL path, not hostname
Reviewer noted the previous hostname-based switch broke the documented
qwen3-rerank compatible-api endpoint (/compatible-api/v1/reranks), which
must use the flat OpenAI-style body and top-level results.
Switch to path-based detection: only /api/v1/services/rerank uses the
native nested input/output envelope; everything else (including the
DashScope compatible-api and generic OpenAI/Cohere gateways) keeps the
flat protocol. Rename _is_dashscope -> _uses_nested_envelope for clarity.
Add regression tests covering the compatible-api flat path and reconcile
the existing native-path fixtures to the nested envelope.
* docs(rerank): use qwen3-rerank for compatible-api example
The compatible-api/v1/reranks endpoint uses the flat OpenAI-compatible
protocol; qwen3-vl-rerank is a native-envelope model served at
/api/v1/services/rerank. Align the example model with the endpoint the
implementation selects by URL path.
---------
Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
* fix(storage): block deleting write-protected viking://agent root
The delete-namespace guard rejected viking:// and viking://user but not the
account-shared viking://agent root that the sibling write guard already
forbids, so a non-root user could rm(viking://agent, recursive=True) and
recursively wipe account-wide agent skills/endpoints/tools/payments. Mirror
the write guard's bare-agent-root rejection; concrete sub-paths stay deletable.
Parity with #2873.
* test(storage): cover delete rejection of bare viking://agent root
#2873 added a delete guard so rm("viking://") / rm("viking://user") raise
PermissionDenied before any side effect. mv() is the sister destructive op
(copy + recursive rm of the source) but only guards the source via the write
guard _ensure_mutable_access, which does not reject the bare "viking://"
account root. So a normal user calling mv from the bare root (HTTP filesystem,
WebDAV MOVE, CLI) reaches the recursive source delete the guard was meant to
forbid.
Apply _ensure_delete_access to the mv source (it is recursively rm'd), keeping
the write guard on the destination. The guard is additive, so it can only
reject more (the protected bare root), never loosen the existing
write-namespace checks; concrete subtree URIs used by internal callers
(watch_manager, semantic_processor) are unaffected.
Adds a parametrized regression test mirroring the existing rm protected-root
test: bare-root mv sources now raise before any AGFS stat/rm.
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
* Handle poisoned known_keys mutexes in cache providers
* Add configurable cache support to ragfs python
* Add Redis cache provider support
* Prune native cache providers from default build
* Add RAGFS cache guides in Chinese and English
* delete .cargo/config.toml
* Fix stale cache invalidation across shared wrappers
* Document Yuanrong native concurrency limits
* add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
add CachedFileSystem + CacheProvider trait and Mooncake/Yuanrong Provide
* Handle poisoned known_keys mutexes in cache providers
* Add configurable cache support to ragfs python
* Add Redis cache provider support
* Prune native cache providers from default build
* Add RAGFS cache guides in Chinese and English
* delete .cargo/config.toml
* Fix stale cache invalidation across shared wrappers
* Document Yuanrong native concurrency limits
* Limit directory cache entries and add regression test
* docs: add TOS s3fs backend test design
* Add runtime cache config override
* add build guides for mooncake and yuanrong
* fix _FakeConfig test with cache
* fix(ragfs): cache encrypted data below the encryption layer
- pass runtime cache config to the Rust binding
- cache ciphertext instead of decrypted content
- disable cache for encrypted multi-write mounts
- update tests and provider build documentation
* fix(ragfs): invalidate cache after same-mount raw copy
Ensure copy_within_mount invalidates the cached destination file and
parent directory after the raw backend fast path writes data directly.
This prevents stale reads and stale directory metadata when cache is
enabled and Python cp() uses the same-mount copy optimization.
Add a regression test covering overwrite copy followed by read/list.
* fix(ragfs): preserve multi-write discovery through cache layer
Allow MountableFS::as_multiwrite() to unwrap CachedFileSystem when
discovering the underlying MultiWriteWrappedFS. This preserves
multi-write admin paths and same-mount copy behavior for cache-enabled,
unencrypted multi-write mounts.
Add a regression test covering sync status, sync retry, same-mount copy,
and unmount behavior for cached unencrypted multi-write mounts.
* feat(ragfs): add cache-aware tree traversal mode
- add configurable tree traversal mode to cache policy
- keep default tree behavior delegated to backend
- allow cached traversal to reuse read_dir directory cache
- bypass cached traversal for multi-write backends
- add regression coverage for tree cache behavior and fallbacks
* docs: design cache-aware grep traversal
* feat(ragfs): add cache-aware grep traversal
Introduce a shared cache traversal mode for recursive APIs and use it to
optionally run grep through CachedFileSystem.
- add CacheTraversalMode with backend and cached_traversal modes
- keep CacheTreeMode as a compatibility alias
- route tree and grep through cached traversal only when explicitly enabled
- reuse cached read_dir entries and full-file reads during grep traversal
- keep multi-write traversal on the backend path
- expose storage.agfs.cache.traversal_mode in Python config
- raise max cached directory entries threshold to 4096
- add regression tests for grep cache traversal and traversal config
* Optimize cached grep generation validation
* Parallelize cached grep file scanning
---------
Co-authored-by: fang <fang@fangMacBook-Air.local>
Cache async SDK clients per event loop to avoid cross-loop reuse in worker threads.
Move memory vectorization into semantic queue refresh and preserve target sync state for resource updates.
The OpenViking docker image still launched the legacy `openviking/console`
standalone service on port 8020. Now that web-studio is bundled into the OV
server itself at /studio (see #2156), that process is redundant and the
port is just a confusing artefact.
This change retires the old console (python package + 8020 + console-frontend
favicons) but **keeps the in-compose Caddy as a stable single-ingress on
port 1934**, just simplified to one upstream now that there's no 8020. The
server-side BFF at `openviking/server/routers/console.py` (under
`/api/v1/console/*`) is also kept — web-studio uses the same endpoints.
**The OAuth authorize page (`openviking/server/oauth/router.py`) is
deliberately untouched in this PR** — the console-link button and Quick
authorize same-origin panel will be re-pointed at web-studio in a focused
follow-up.
BREAKING CHANGES:
- Port 8020 is gone from the docker image and docker-compose.yml; Caddy at
1934 now forwards everything to 1933 (web-studio lives at /studio there).
Anything bookmarked at `http://host:8020/...` must migrate to
`http://host:1933/studio/`.
- `python -m openviking.console.bootstrap` no longer exists; the python
package `openviking.console` has been removed.
Pip packaging:
- web-studio dist is now shipped inside the wheel under
`openviking/web_studio/dist/` (mirroring the old `openviking/console/static/`
layout). The dockerfile copies `--from=web-studio-builder /web-studio/dist`
into the source tree before `uv sync`, so the wheel produced by the
default docker build always carries the SPA. Building the wheel without
running `npm run build` first leaves the directory empty, which gracefully
degrades /studio to a 404 without breaking server startup.
- Favicon assets (`favicon.ico` / `favicon-32.png` / `apple-touch-icon.png`,
~11 KB total) are duplicated into `openviking/server/static/` and shipped
via package-data so `/favicon.*` and `/mcp/favicon.*` routes are always
registered, regardless of whether the web-studio dist is bundled.
- `pyproject.toml` and `setup.py` `package-data` drop `console/static/**`
and add `server/static/**` + `web_studio/dist/**`.
- New favicons (the 16/32/180 set in both `openviking/server/static/` and
`web-studio/public/`) are downscaled from the canonical
`web-studio/public/openviking-icon.png`, so the small-icon family matches
the SPA's high-res rel="icon" target — the studio tab icon now stays
consistent whether the browser uses the HTML link tag or falls back to
auto-fetching `/favicon.ico`.
Server:
- `openviking/server/app.py` now reads `/studio` from
`Path(__file__).parent.parent / 'web_studio' / 'dist'` by default;
`OPENVIKING_WEB_STUDIO_DIR` still wins for dev mode pointing at a
repo-local build. Favicon routes are unconditionally registered and
load from `openviking/server/static/`.
- `openviking/observability/usage_audit/projection.py` drops the legacy
`/console/*` skip prefix (the BFF prefix `/api/v1/console/*` remains).
Docker:
- `web-studio-builder` stage moved earlier (Stage 2) so its dist can flow
into `py-builder` before `uv sync` runs.
- Runtime stage no longer separately copies the dist or sets
`OPENVIKING_WEB_STUDIO_DIR`; the in-package path is the default.
- Entrypoint renamed `openviking-console-entrypoint.sh` -> `openviking-entrypoint.sh`
and stripped of the `python -m openviking.console.bootstrap` launch.
- `EXPOSE 1933 8020` -> `EXPOSE 1933`.
- `docker-compose.yml` drops the openviking service's 8020 port mapping;
the caddy service stays but no longer needs port 8020 exposed.
- `Caddyfile` simplified to a single `:1934 { reverse_proxy openviking:1933 }`
— the legacy `/console/*` route to :8020 is gone.
Docs:
- en/zh quickstart updated to drop the 8020 mapping and explain that the
API server now also serves `/studio`.
- Other guides (`12-public-access.md`, `11-oauth.md`, `05-observability.md`,
`04-setup-for-agent.md`, `03-deployment.md`) are intentionally left for a
focused follow-up PR alongside the OAuth quick-authorize reintroduction.
Tests:
- Deleted `tests/misc/test_console_{proxy,static_assets}.py` (covered the
removed console package). `tests/observability/test_console_router.py`
stays — it covers the BFF, which remains.
* feat: ov add-resource (spec -L --level), ov stat (return count for dir)
* feat: ov add-resource (spec -L --level), ov stat (return count for dir)
* feat: Add VLM backup configuration for automatic failover
- Add backup field to VLMConfig with recursive backup prevention
- Implement FailoverVLM wrapper class for automatic failover
- Support rate limit, timeout, server error triggers
- Add comprehensive unit tests
* feat: Add VLM backup configuration for automatic failover
* feat: Add VLM backup configuration for automatic failover
* ov observer filesystem
* ov observer filesystem