* feat(storage): optimize glob func
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* fix(localfs): offload blocking fs operations to spawn_blocking
* feat(glob): cap glob api default node_limit at 256
* feat(sdk): add node_limit options for glob in python and go SDKs
Ollama defaults to a 4096-token context window and silently truncates any
longer prompt to fit. OV's memory-extraction prompt is ~5.4k tokens even for a
short session, so the conversation (which sits at the top of the prompt) is
dropped and the model receives only the format spec. It then returns an empty
`{"memories": []}` with no error, so local Ollama VLMs extract nothing
regardless of model size. Thinking models compound this by emitting only
reasoning and stalling.
- litellm_vlm: for `ollama/` and `ollama_chat/` models, default `num_ctx`
to 16384 and disable thinking via `extra_body`. Both are overridable through
`extra_request_body`, and the change is gated to Ollama routes so other
providers are untouched. This fixes extraction, intent analysis, and the
query planner for every local Ollama config, not just wizard-generated ones.
- setup_wizard: write the same `extra_request_body` explicitly in the Ollama
VLM config block so the setting is visible and tunable in ov.conf.
- setup_wizard: drop the qwen3.5:2b VLM preset. 2B models "extract" the
prompt's few-shot examples as fabricated memories; qwen3.5:4b is the smallest
model that extracts cleanly. RAM-tier defaults are reindexed accordingly.
Verified end to end: with num_ctx raised, prompt_eval goes from 4096
(truncated) to the full 5441 tokens and qwen3.5:4b extracts the correct
memories; without it, extraction returns empty.
* feat(init): TUI-style setup wizard with separate embedding/VLM config
Rework `openviking-server init` into an arrow-key TUI (rich chrome +
termios raw-mode menus, with a numbered-input fallback for non-TTY).
- Two-step main flow: Step 1 embedding, Step 2 VLM, so mixing cloud and
local (one online, one local) is visible in the main flow rather than
hidden in submenus.
- Recommended all-Ollama one-shot config (embedding + VLM + query
planner) sized by available RAM.
- Refreshed VLM presets (qwen3.6:27b / qwen3.6:35b).
- Show current configuration at every level: overview panel, menu-item
"(now: ...)" descriptions, and per-section "Current: ..." lines.
- Seed every prompt default from the existing config; offer to keep
existing secrets (masked); show an old -> new diff before saving.
- Server & auth step now covers host + port, and switching to a
non-local binding requires (and can generate) a root_api_key.
- Left/right arrow navigation: <- back a step, -> confirm.
- Auto-offer init on a bare `openviking-server` when no config exists;
disk-space check before pulling models; post-save doctor validation
with optional server start.
* fix(init): preserve existing config in wizard update flow
Address Copilot review feedback on the setup wizard:
- Cloud embedding update now keeps a non-default api_base (proxy /
OpenAI-compatible endpoint) when the provider is unchanged, instead
of always resetting to the provider preset default.
- Custom OpenAI-compatible VLM branch seeds API base and model from the
existing config and routes key collection through _prompt_vlm_api_key,
so keep-existing-key / env-var reuse applies there too.
- server_bootstrap init auto-offer now catches OSError alongside
EOFError, matching the wizard's own prompt helpers, so non-standard
pipe/terminal contexts fall through silently instead of crashing.
* chore: clear unused files
* fix(tests): fix unit test
* refactor(auth): introduce plugin-based authentication architecture
Replace the monolithic `openviking/server/auth.py` with an extensible
plugin-based auth system. This refactor extracts the three built-in modes
(`dev`, `api_key`, `trusted`) into separate `AuthPlugin` implementations,
adds a registry for third-party plugins, and preserves all existing behavior
while enabling custom authentication backends (e.g. LDAP, OIDC, mTLS).
Key changes:
- **New public API**: `AuthPlugin` (ABC) and `register_auth_plugin` decorator.
- **New registry**: `AuthPluginRegistry` supports runtime registration.
- **Built-in plugins**: `DevAuthPlugin`, `ApiKeyAuthPlugin`, `TrustedAuthPlugin`.
- **Config change**: `auth_mode` widened from `Literal` to `str` for custom modes.
- **Validation delegated**: `validate_server_config()` now delegates to the active
plugin's `validate_config()`, preserving existing validation semantics.
- **Router compatibility**: All existing `require_*` decorators and `resolve_identity`
/ `get_request_context` dependencies remain unchanged. Routers import the same
symbols from `openviking.server.auth`.
- **Tests**: `conftest.py` manually wires the DevAuthPlugin in ASGI tests (lifespan
not triggered). `test_auth.py` expanded with plugin registration and validation tests.
- **Docs**: `04-authentication.md` (en/zh) updated with plugin registration examples.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): fix trusted mode test
* fix(tests): fix unit test
* fix(cli): remove unexisted transaction observer
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* fix(skills): now we allow viking://agent/skills again, and optimize CLI for skills
* docs(skills): use -p instead of --parent in agent skills examples
Align the `ov skills add` examples in the context-types and viking-uri
docs with the short flag `-p` introduced for `ov skills list/find/show`,
so all four user-facing examples consistently demonstrate the short form
when targeting `viking://agent/skills`.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): error check for api key
* fix(tests): unit test wait until resource not busy
* fix(tests): unit test wait until resource not busy
* fix(sdk): args form in skills find
* fix(skills): pass target uri in request body
---------
Co-authored-by: claude-sonnet-4-6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
Promote ov_intent_analysis_sft:v7_q8 to the recommended local Ollama
query-planner model. Add the bundled retrieval.ov_intent_analysis_sft_v7
prompt and map v7_q8 to it (v4_q8 mapping kept). Update the setup wizard
presets (v7 recommended, v4 retained, v1 dropped) and the configuration
docs (EN/ZH). Extend tests to cover the v7 mapping.
Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(cli): add query planner setup to init wizard
Let `openviking-server init` configure the optional lightweight
query_planner model. The wizard pulls the chosen Ollama model and writes
the query_planner config; the IntentAnalyzer selects the matching prompt
at retrieval time via a model->prompt-id mapping, so no prompt files are
copied and no prompts.templates_dir override is needed.
- intent_analyzer: QUERY_PLANNER_PROMPT_BY_MODEL maps the fine-tuned SFT
models to their bundled prompt id; unmapped models keep the default
retrieval.intent_analysis prompt.
- bundle retrieval/ov_intent_analysis_sft_v4.yaml (loaded by its own id).
- ollama detection + doctor now recognize query_planner Ollama usage.
- docs: describe the init flow and runtime prompt selection.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(cli): offer query planner on all paths, recommend it with an Ollama VLM
The init wizard offers the lightweight query planner after model setup. When the
chosen setup already uses an Ollama VLM (`ollama_running` is not None) the
planner rides on that running Ollama at near-zero extra cost, so the enable
prompt is tagged "(recommended)" and defaults to yes. For cloud / non-Ollama VLM
setups it is still offered, but defaults to no and drops the recommendation;
opting in there runs the Ollama install flow.
The Ollama state established during model setup is threaded through the wizard so
the planner reuses it instead of re-running the install dialog:
- `_wizard_ollama` / `_wizard_llamacpp` return `(config, ollama_running)`.
- `run_init` forwards that state to `_wizard_query_planner`.
Docs (zh/en) and tests updated accordingly.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: guoxuter <j7azwflq4h@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
CLI passes --timeout as query param but server only reads from request body,
causing infinite wait when queue is not empty. Skip until the bug is fixed.
- Add CLI integration tests (10 test files under tests/cli/)
- Extend api_test.yml with CLI install + test steps
- Run filesystem + scenarios/resources_retrieval serially to avoid 409 conflicts
- Other tests parallel with -n 4
- Add release prereleased trigger to api_test.yml and api_test_effect.yml
- Deduplicate oc2ov_test P0 cases (20→12, ~30-55min saved):
- Delete test_memory_write.py (covered by V2 suite)
- Remove events/tools from V2 suite (structurally identical to entities/skills)
- Remove test_memory_read_verify (covered by V2 suite)
- Remove test_cross_session_recall (overlaps with recall_explicit_search)
- Add ensure_resources_dir fixture to prevent NOT_FOUND on fresh environments
- Add retry logic for 429/500/403 rate-limit in api_client.py
- Add retry for commit when task_id is None in test_memory_v2_full_suite.py
- Add exponential backoff retry for GitHub platform test 5xx errors
BytePlus was configured with provider="openai" but the BytePlus
endpoint uses the Volcengine API (multimodal_embeddings at
/embeddings/multimodal). Using the OpenAI SDK sends requests to the
wrong path (/embeddings with string input), causing 500 errors or
hangs.
Also adds a "Custom (OpenAI-compatible)" VLM option to both the cloud
and local wizard flows, so users can point to any OpenAI-compatible
endpoint (e.g., MiMo, vLLM, LiteLLM proxies).
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(setup-wizard): add server binding step and reorder cloud providers
Add a Local/Remote host binding prompt with required root_api_key for
remote (0.0.0.0) so the wizard is usable inside Docker. Reorder cloud
providers to surface VolcEngine (火山引擎) first as the default, add
BytePlus as the second option (OpenAI-compatible ARK endpoint), and
drop OpenAI to third.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(setup-wizard): rotate .bak backups and mask API key input
When the previous .bak already exists, fall back to .bak.1, .bak.2, ...
so re-running init never silently overwrites an earlier backup.
For API key prompts, echo `*` per character at typing/paste time
(termios raw mode on Unix; getpass fallback on Windows; plain input
when stdin/stdout aren't TTYs so tests and pipes still work). Secrets
are no longer visible in scrollback or screen recordings.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(setup-wizard): make Cloud API the first and default setup mode
Most users land on this wizard from a remote/cloud deployment path, so
Cloud API now appears first in the setup-mode prompt and is the default
selection. Local llama.cpp drops to second, Ollama to third, Custom
stays last.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(docker): collapse persistent state into /app/.openviking
The Docker image previously crashed on startup when ov.conf was missing,
making it impossible to docker exec in to fix the configuration. This
also forced two separate volume mounts (config + data), which is awkward
on managed platforms that only offer one persistent volume.
Changes:
- Set HOME=/app and put all persistent state under /app/.openviking, so
the in-container layout mirrors the host's ~/.openviking.
- docker-compose.yml mounts ~/.openviking -> /app/.openviking as a single
volume that captures ov.conf, ovcli.conf, and the workspace.
- Entrypoint bootstraps ov.conf from OPENVIKING_CONF_CONTENT when set;
otherwise prints a fix-it message and sleeps until the file exists, so
the container stays up for docker exec.
- openviking-server init honors OPENVIKING_CONFIG_FILE and derives the
workspace from its parent directory, so a single env var places init
output where the server will read it.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(docker): describe single-mount layout and no-mount fallbacks
Switch all Docker examples to the new single-volume layout
(~/.openviking -> /app/.openviking) and document the two ways to
configure when bind mounts aren't available: pass the full ov.conf JSON
through OPENVIKING_CONF_CONTENT, or docker exec in and run
openviking-server init while the entrypoint waits for the file.
Updates docs/{en,zh}/getting-started/02-quickstart.md and
docs/{en,zh}/guides/03-deployment.md, including the macOS socat block.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: link
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(session): add account namespace policy and shared sessions
Unify namespace resolution across filesystem, indexing, and session storage.
Add account-shared session paths, role_id auth semantics, and an HTTP demo
script for the four namespace-policy combinations.
* space
* fix(pack): skip derived semantic files in ovpack transfer
Keep ovpack imports resilient to stale sidecars and rebuild semantics through the normal queue instead of restoring derived files verbatim.
* Revert "fix(pack): skip derived semantic files in ovpack transfer"
This reverts commit f4e4db8401.
* fix(namespace): default legacy accounts to agent-shared policy
Clarify that memory.agent_scope_mode is deprecated and document the supported agent memory migration paths.
* refactor: all env are listed in openviking_cli/utils/config/consts.py
* refactor: all env are listed in openviking_cli/utils/config/consts.py
* refactor: all env are listed in openviking_cli/utils/config/consts.py
* refactor: all env are listed in openviking_cli/utils/config/consts.py
* refactor: all env are listed in openviking_cli/utils/config/consts.py
---------
Co-authored-by: openviking <openviking@example.com>
* feat: add `ov init` interactive setup wizard for local model deployment
Add an interactive CLI wizard that guides users through configuring
OpenViking with local Ollama models, especially targeting macOS/Apple
Silicon beginners. The wizard auto-detects and installs Ollama,
recommends models based on system RAM, pulls selected models, and
generates a valid ov.conf.
Supported models:
- Embedding: qwen3-embedding (0.6b/4b/8b), embeddinggemma:300m
- VLM: qwen3.5 (2b-122b), gemma4 (e2b/e4b/26b/31b)
Also fixes `ov doctor` to recognize Ollama providers as valid without
requiring an API key.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* feat: add Ollama lifecycle management for server startup and health checks
Extract Ollama utilities into shared module (openviking_cli/utils/ollama.py)
so both `ov init` and `openviking-server` can reuse them. Server now
auto-detects Ollama from config and ensures it's running at startup
("ensure running, never stop" pattern). Adds Ollama connectivity to
`/ready` health check and `ov doctor` diagnostics.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor: move init and doctor from ov to openviking-server subcommands
These are server-side configuration commands (generate/validate ov.conf),
not client operations. Having them under `ov` (the client CLI) was
confusing. Now:
openviking-server init # setup wizard
openviking-server doctor # diagnostics
ov <subcommand> # client operations only
* restore rust_cli.py
* refactor: update command references to use 'openviking-server doctor'
* fix: add existence check for example config in custom configuration wizard
* refactor: update references from 'ov init' to 'openviking-server init' in setup wizard and tests
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
---------
Co-authored-by: openviking <openviking@example.com>
Apply os.path.expandvars() to raw config text before JSON parsing in all
config-loading code paths, so that $VAR and ${VAR} placeholders in ov.conf
are resolved from the environment. This is especially useful for container
deployments where secrets are injected as environment variables.
Files changed:
- openviking_cli/utils/config/config_loader.py: load_json_config()
- openviking_cli/utils/config/open_viking_config.py: _load_from_file()
- openviking_cli/doctor.py: check_config(), check_embedding(), check_vlm(), check_disk()
- bot/vikingbot/config/loader.py: load_config() (already had this, no change)
The server config path (openviking/server/config.py -> load_server_config)
delegates to load_json_config() and is covered by the fix above.
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(cli): add ov doctor diagnostic command
Adds a new `ov doctor` command that validates all OpenViking subsystems
and reports actionable diagnostics without requiring a running server.
Checks: config file, Python version, native vector engine (PersistStore),
AGFS, embedding provider, VLM provider, and disk space. Each check is
isolated so one failure doesn't block others, and every failure includes
a specific fix suggestion.
This addresses a real pain point: when the native engine is missing from
a pip wheel (e.g., Python 3.13), the only feedback is 50+ ERROR log
lines with no actionable guidance. `ov doctor` catches this immediately:
Native Engine: FAIL No compatible engine variant
Fix: pip install openviking --upgrade --force-reinstall
Alt: Use vectordb.backend = "volcengine" instead of "local"
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* docs: add ov doctor screenshot examples
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(doctor): reuse resolve_config_path and drop hardcoded OPENAI_API_KEY fallback
Address review feedback from @qin-ctx:
- Replace _CONFIG_SEARCH_PATHS and _find_config() with resolve_config_path()
from config_loader.py to avoid two sources of truth for config discovery
- Remove hardcoded OPENAI_API_KEY env var fallback from embedding and VLM
checks - only check the api_key field in config
- Renumber inline comments in rust_cli.py (1/2/3 matching docstring)
---------
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
agent_space_name() computed md5(user_id + agent_id) without a separator,
so different (user_id, agent_id) pairs could produce the same hash when
their concatenation matched (e.g. ("alice","bot") vs ("aliceb","ot")).
Add ":" separator between user_id and agent_id in the hash input. The ":"
character is safe because the validation regex [a-zA-Z0-9_-] prevents
either field from containing it.
Fix applied to all three implementations:
- openviking_cli/session/user_id.py (Python SDK)
- bot/vikingbot/openviking_mount/ov_server.py (bot server)
- examples/openclaw-memory-plugin/client.ts (TypeScript example)
Note: This is a breaking change for existing agent spaces. Existing data
directories were named using the old hash and will not be found with the
new hash. A migration script or fallback lookup may be needed.
Fixes#595
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
* feat: remove python cli and disable python -m openviking
* feat: ls, tree, find, search, grep all use --node-limit as the result limiting arg
* feat: add ls -n
* feat: update agfs to support grep -n
* feat: update agfs to support grep -n
* docs: cancel modify
---------
Co-authored-by: openviking <openviking@example.com>
* fix: remove await asyncio and call agfs directly
* feat: mv cli out of openviking
* refactor: mv cli out of openviking
* refactor: mv cli out of openviking