* fix(ragfs): enforce mount containment and write-flag semantics in LocalFS
Sweep findings: E-01, E-04. Reject lexical traversal and honor LocalFS write contracts.
* fix(review): reject absolute localfs remainders
Addresses blocking review finding on #3402.
* fix(ragfs): close LocalFS glob mount-escape gap
LocalFS::glob_directory only called validate_virtual_path(path), which
rejects `..` but accepts an absolute remainder. A `//`-double-slash mount
path (e.g. `/local//etc`) survives normalize_path, and find_mount hands the
remainder `//etc` to glob_directory; validate_virtual_path passes it
(components are [RootDir, Normal("etc")], no ParentDir), and glob_via_walk's
resolve_virtual_path strips one slash to `/etc` and joins it over the base —
an absolute join that overrides the base and lists the host directory.
Every other op (read/write/stat/rename/remove/grep) routes through
resolve_path, which adds the is_absolute check. glob now calls the same
resolve_path guard (discarding the returned PathBuf) so it shares the
containment contract. Directory-listing info leak only; content reads were
already covered.
Adds test_localfs_glob_mount_rejects_absolute_remainder mirroring the read
regression test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RbWx1T81KkNV4nucxWsPXZ
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(bot): never silently replace corrupt config, session, or cron stores
Sweep findings: C-03, C-05, C-06. Fail fast on corrupt persisted state before any overwrite.
* test(bot): drop corrupt-state regression tests per review
The three fixes (loader raise on invalid config, session/cron refuse to
overwrite corrupt state) stay; the accompanying tests are removed as not
worth their weight.
* fix(kernel): stop conflating storage failures with not-found
Sweep findings: A-03, A-07, A-11, A-12, B-07. Preserve storage and parse failures instead of reporting missing or empty state.
* fix(review): restore archive failure handling
Addresses blocking review finding on #3417.
* fix(review): terminalize corrupt archive records
Addresses blocking review finding on #3417.
* test: adapt pending-archive-skip test to refactored archive scan
Rebase onto main (#3380 turn-aware retention) changed archive refs to carry
an archive_id; update the test mock's _list_archive_refs return so the missing
pending archive still routes through _get_uncovered_archive_messages and is
skipped (not raised).
* fix(crypto): never overwrite Vault root key on transient read failures
Sweep finding B-01
* fix(crypto): clear cached ephemeral root key when Vault persist fails
The InvalidPath create branch cached the freshly generated root key in
self._root_key before encrypting/persisting it. If _encrypt_with_vault or
the KV write failed (transient transit/KV outage), get_root_key raised but
left the never-persisted key cached, so a retry returned it via the
`self._root_key is not None` fast path and encrypted data with a key that
vanishes on restart — the data-loss class this provider guards.
Wrap encrypt+persist in one try and null the cache before raising; add a
regression asserting the cache is cleared and the next call re-reads Vault.
Addresses the blocking review finding on #3423.
Session commit Phase-2 memory extraction runs in the SESSION_COMMIT queue
worker, which never bound a root observability context. VLM/embedding token
events emitted during extraction therefore read identity from an empty root
context and were recorded under account_id=__unknown__, invisible to the
account-filtered usage dashboard (which showed 0 VLM tokens despite active
extraction).
Bind a QUEUE root observability context with the committing account/user at
the start of SessionCommitProcessor._process and reset it in finally, mirroring
SemanticProcessor.on_dequeue. The binding must live inside the coroutine
because on_dequeue hops event loops via run_coroutine_threadsafe, so a context
bound there would not propagate. The extraction call chain inherits the
contextvar, so one bind covers all Phase-2 VLM/embedding calls.
Add tests asserting the worker binds the committing identity onto the root
context and resets it afterwards.
* feat(embedder): support extra_body passthrough in OpenAI embedder config
Adds optional `extra_body` (dict) to the OpenAI dense embedder, merged
into every embeddings.create call. Motivating use case: OpenRouter
provider routing ({"provider": {"sort": "latency"}}) — default routing
shows p90=35s/max=127s tail latency that kills interactive recall
(A/B: sorted routing is consistently sub-second).
Explicit query_param/document_param keys still take precedence on
conflict.
* feat(config): wire extra_body through embedding config layer
Add optional extra_body field to EmbeddingModelConfig and pass it to
OpenAIDenseEmbedder for the openai/azure providers, including the
multi-credential failover merge (parent-level model-behavior field).
* docs(config): document embedding extra_body with OpenRouter routing example
* docs(config): restore concrete host/cors_origins values in EN full schema
* fix(server): attribute MCP traffic in observability (route + identity)
MCP requests were audited as route=/__unmatched__, account_id=__unknown__
because (1) the /mcp app is registered as a plain Starlette Route so
scope["route"] is never set, and (2) _IdentityASGIMiddleware resolved
identity without calling update_root_span_identity.
Register /mcp via a _ScopedRoute subclass that sets child_scope["route"]
on match (mirroring APIRoute.matches) and stamp root-span identity after
resolution. 404 fallbacks and middleware code are untouched.
* test(server): cover MCP scope route resolution and root-span identity stamping
Assert the /mcp route sets scope["route"] on match (and that unmatched
paths still fall back without it), and that _IdentityASGIMiddleware stamps
the resolved account/user onto the root span attributes.
* fix(server): attribute MCP traffic in observability (route + identity)
MCP requests were audited as route=/__unmatched__, account_id=__unknown__
because (1) the /mcp app is registered as a plain Starlette Route so
scope["route"] is never set, and (2) _IdentityASGIMiddleware resolved
identity without calling update_root_span_identity.
Register /mcp via a _ScopedRoute subclass that sets child_scope["route"]
on match (mirroring APIRoute.matches) and stamp root-span identity after
resolution. 404 fallbacks and middleware code are untouched.
* fix(parse): handle parentheses in Markdown image paths
The image regex !\[([^\]]*)\]\(([^)]+)\) used [^)]+ for the path capture
group, which truncates at the first ) character. When document titles
or filenames contain balanced parentheses (e.g. "文档_17 (17号项目)"), the
generated image paths include ) and the regex captures a truncated,
non-existent path. This causes _resolve_image_path() to fail silently
(WARNING only), and the image is never copied to VikingFS or sent to
VLM for understanding.
Fix: replace the path capture group with (?:[^()]|\([^()]*\))+, which
allows one level of balanced parentheses inside the path while still
terminating at the correct closing ) of the Markdown image syntax.
Add focused tests covering balanced parens in directory and filename
components, URLs with parens, multiple images on one line, and
non-matching of plain links.
Fixes#3455
* fix(test): exercise MarkdownParser._image_pattern directly, remove unused import
Address review feedback on #3462:
1. Tests now import and instantiate MarkdownParser to access the
production _image_pattern regex, instead of compiling an independent
copy. Tests fail if the production regex regresses.
2. Remove unused `import pytest` (Ruff F401).
* fix(parse): rewrite parenthesized image paths
* test: remove extra image rewrite regression case
* fix(parse): share markdown image parsing for rewrite
* refactor(parse): keep markdown image fix minimal
---------
Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(ov): redact gateway secrets in config show and create root key with 0600
Sweep findings: E-02, E-03. Prevent credential disclosure in output and at key creation.
* test(ov): drop trivial init-key permission test
The 0600 fix is a one-line OpenOptions::mode; a dedicated tokio+tempfile
test module for a single mode assertion is not worth its weight. The
redaction test in store.rs (a real multi-case security behavior) stays.
* fix(bot): pass sender_name in all channel adapters
Sweep findings: C-01. Supply display-name fallbacks so inbound messages reach the bus.
* fix(bot): make sender_name optional and wire real WhatsApp pushName
Root-cause guard: _handle_message required sender_name positionally, but
InboundMessage.sender_name is str|None=None and context.py already falls back
to sender_id, so the required-ness was an accidental signature/contract
mismatch. Make it optional (reordered after the still-required chat_id/content;
all 11 call sites use keyword args) so no future adapter can crash on it.
WhatsApp: the previous call read data.get("senderName")/data.get("pushName"),
neither of which the bridge ever sends, so it silently always fell back to the
numeric id. Forward baileys' msg.pushName through the bridge payload and read it
in Python, so WhatsApp group chats show real display names like other channels.
* fix(rerank): support DashScope nested request/response envelope
OpenAIRerankClient sent a flat request body ({"model", "query",
"documents"}) and parsed "results" at the top level of the response.
DashScope (qwen3-rerank) requires a nested envelope:
Request: {"model", "input": {"query", "documents"}, "parameters": ...}
Response: {"output": {"results": [...]}, "request_id", "usage"}
This caused DashScope rerank to silently fail — the response had no
top-level "results" key, so the client returned None.
Changes:
- Add _is_dashscope() to detect DashScope endpoints by host marker.
- Add _build_request_body() that produces the nested envelope for
DashScope and the flat body for standard OpenAI/Cohere services.
- Add _extract_results() that reads output.results for DashScope and
top-level results for standard services.
- Accept both "relevance_score" (singular, DashScope) and
"relevance_scores" (plural, some providers) in result items.
- Add 13 tests covering host detection, body construction, response
parsing, end-to-end mocked flows for both providers, plural key
handling, empty documents, and sparse results.
Fixes#3459
* fix(rerank): detect DashScope protocol by URL path, not hostname
Reviewer noted the previous hostname-based switch broke the documented
qwen3-rerank compatible-api endpoint (/compatible-api/v1/reranks), which
must use the flat OpenAI-style body and top-level results.
Switch to path-based detection: only /api/v1/services/rerank uses the
native nested input/output envelope; everything else (including the
DashScope compatible-api and generic OpenAI/Cohere gateways) keeps the
flat protocol. Rename _is_dashscope -> _uses_nested_envelope for clarity.
Add regression tests covering the compatible-api flat path and reconcile
the existing native-path fixtures to the nested envelope.
* docs(rerank): use qwen3-rerank for compatible-api example
The compatible-api/v1/reranks endpoint uses the flat OpenAI-compatible
protocol; qwen3-vl-rerank is a native-envelope model served at
/api/v1/services/rerank. Align the example model with the endpoint the
implementation selects by URL path.
---------
Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
- Add directory_marker_mode: none to all S3 config examples
- Add S3-compatible storage notes with required fields table
- Add Docker networking guidance for Linux vs macOS/Windows
- Remove private IP addresses from examples (use localhost)
- Apply changes to both English and Chinese versions
* docs: revamp README for readability, route detail to docs.openviking.ai
The README had grown to ~850 lines, 60% of it provider-config JSON that
duplicates the deployed configuration guide. Rewritten to ~253 lines:
- Lead with what the product is, a Studio screenshot, and five feature
bullets, each outlinked to docs.openviking.ai
- Move benchmarks (LoCoMo, tau2-bench, HotpotQA) above the fold; drop two
derived tables in favor of one-sentence summaries + ./benchmark links
- Collapse install to the init/doctor golden path; all provider JSON,
ov.conf templates, env vars, and Windows setup now route to the
configuration guide (verified live)
- Add the previously missing "Use it with your agent" section linking all
10 integration docs
- README_CN (zh docs links) and README_JA (en docs links; ja docs not
deployed) rewritten to mirror section-for-section
- New hero screenshot docs/images/studio-playground.png from
openviking.ai/studio
All 76 external URLs and every repo-relative link verified.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4oDxmJojygyomsz9BBhhn
* docs: fix review findings — restore uncovered detail, drop false pointers
Adversarial review of the revamp found four claims pointing at coverage
that does not exist:
- Restore `cargo install --git ... ov_cli` build-from-source path (was
deleted with no docs destination; docset has no cargo install anywhere)
- Restore `ov reindex` mode documentation (vectors_only /
semantic_and_vectors / prune_orphans / --dry-run / no alias warning) —
covered by no linked doc
- Remove "per-agent breakdown is in ./benchmark" (benchmark/ holds
reproduction scripts, not result tables)
- Remove "Reproduce it from ./benchmark" on the 5-dataset RAG summary
(adapters exist for only 3 of 5 datasets)
Also: EN/JA quick-start grep example now targets docs/en instead of
docs/zh. Applied identically to README.md, README_CN.md, README_JA.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4oDxmJojygyomsz9BBhhn
* docs: apply review feedback — blog philosophy link, deployed doc links, drop dead widgets
- Link the design-philosophy essay (The Database Paradigm for Context
Engineering, blog.openviking.ai) from the Why section so the old
README's design narrative has a durable home; add Blog to community
- Switch remaining ./docs about-us links (header + community, incl. QR
anchors) to docs.openviking.ai; zh anchors verified against deployed
page ids (#飞书群 / #微信群)
- Remove the star-history chart (service currently renders nothing) and
the stale "May 2026 Update" banner line
- Caption now states the Studio link is a live demo, no install needed
Applied identically to README.md, README_CN.md, README_JA.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4oDxmJojygyomsz9BBhhn
* docs: benchmark charts, easier quick start, logo padding
- Replace the three benchmark tables with one theme-aware SVG chart
(light/dark via <picture>): LoCoMo and tau2-bench as grouped bars,
gray = without OpenViking, blue = with. HotpotQA leaves the README;
full results link to the benchmark report on blog.openviking.ai.
Hand-written SVG, exact numbers from the tables — no generated images.
- Rework Quick start reading flow: nohup folds into the install block,
note that pip install already ships the ov client CLI, close with a
two-link "Next steps" (CLI setup, Deployment). ov reindex modes and
the cargo source install move to the CLI setup doc (en+zh) so the
README stays an easy entry.
- Shrink logo artwork to 0.7 inside the same 842x842 canvas for
breathing room.
Applied to README.md, README_CN.md, README_JA.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4oDxmJojygyomsz9BBhhn
* docs: enlarge logo artwork 1.1x within same canvas
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4oDxmJojygyomsz9BBhhn
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>