* fix(models): route multimodal inputs to DashScope embed_content
Sweep findings: A-09. Preserve image parts through the shared embedding entrypoints.
(cherry picked from commit 50b20d7ca9)
* fix(models): constrain DashScope multimodal routing
Sweep findings: A-09. Preserve text mode and require a single fused vector for multipart input.
(cherry picked from commit 77784429d4)
* fix(models): respect DashScope Qwen fusion parameters
Sweep finding A-09
(cherry picked from commit aafc702b20)
* fix(review): preserve tongyi multipart compatibility
Addresses blocking review finding on #3400.
(cherry picked from commit 96844047c5)
* fix(eval): flush queued records on stop and adapt RAG pipeline to FindResult
Sweep findings: B-05, B-12. Drain recorder queues through the sentinel and consume current retrieval result objects.
(cherry picked from commit fecdbeb6d4)
* fix(sdk/python): support sync client inside a running event loop
Sweep findings: F-09. Run sync wrappers on one persistent worker loop with result and exception propagation.
(cherry picked from commit ba0faadc87)
* fix(sdk/python): preserve cancellation and fork safety
Sweep finding: F-09. Preserve original cancellation errors and reset worker synchronization after fork.
(cherry picked from commit a19e3c95e9)
* feat(sdk/go): add tags filter and relations API for parity
Sweep findings: F-11, F-12. Expose existing server capabilities consistently to Go callers.
(cherry picked from commit 8af0c2cb36)
* fix(sdk/python): runnable quickstarts, correct migrate payload, explicit timeout precedence
Sweep findings: F-01, F-02, F-14. Initialize documented clients and preserve Python SDK request/config semantics.
(cherry picked from commit d9a64f5665)
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* feat(embedder): support extra_body passthrough in OpenAI embedder config
Adds optional `extra_body` (dict) to the OpenAI dense embedder, merged
into every embeddings.create call. Motivating use case: OpenRouter
provider routing ({"provider": {"sort": "latency"}}) — default routing
shows p90=35s/max=127s tail latency that kills interactive recall
(A/B: sorted routing is consistently sub-second).
Explicit query_param/document_param keys still take precedence on
conflict.
* feat(config): wire extra_body through embedding config layer
Add optional extra_body field to EmbeddingModelConfig and pass it to
OpenAIDenseEmbedder for the openai/azure providers, including the
multi-credential failover merge (parent-level model-behavior field).
* docs(config): document embedding extra_body with OpenRouter routing example
* docs(config): restore concrete host/cors_origins values in EN full schema
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* feat: implement multi-credential priority call design
Add OrderedCredentialSwitcher for N-credential failover, MultiCredentialVLM and FailoverEmbedder for credential switching, support automatic migration from legacy backup config
* refactor: deduplicate AllCredentialsFailedError and fix logging in switcher
* chore: fix code formatting for lint compliance
* chore: fix remaining code formatting
* fix: update FailoverEmbedder for multimodal API compatibility
* fix
* fix
* refactor: split error classes and fix fail-fast in credential failover
Separate the monolithic PERMANENT error class so the switcher reacts to
the actual root cause:
- 400 (request-level parameter error) -> PERMANENT, fail-fast: same
request fails on every credential of the same model.
- 401/403/unauthorized/accountoverdue -> new AUTH class: credential-level,
advances to the next credential in multi-credential mode.
- new CONTENT_SAFETY class (moderation rejections) and INPUT_TOO_LARGE ->
fail-fast: switching credentials cannot help.
classify_api_error now checks CONTENT_SAFETY before PERMANENT so a
moderation message containing "400" is not misclassified. On fail-fast the
failover wrappers re-raise the original exception (preserving type/info)
instead of wrapping it, so callers can react (e.g. truncate on
input_too_large). AllCredentialsFailedError is reserved for chain
exhaustion. The legacy PrimaryBackupSwitcher also switches on AUTH to keep
existing backup behavior.
* fix: make get_active_index side-effect free in credential switcher
get_active_index() previously mutated state: when a failback threshold was
met it would decrement the active index. Because observability properties
(active_credential_index / active_credential_id) call it, merely reading the
current credential for logging or metrics could accidentally advance the
failback state machine.
Split the concern: get_active_index() is now a pure read, and a new public
maybe_failback() performs the one-step failback (and logs an info line when
the active credential index changes). The request loops in MultiCredentialVLM
and FailoverEmbedder call maybe_failback() at the top of each attempt, so the
failback behavior is unchanged while pure reads no longer have side effects.
* fix: drop global total_max_retries cap from credential failover
The failover loops exited on `idx >= n OR total_attempts >= total_max_retries`
(default 10). With more than 10 credentials, or when failback churn inflated
the attempt count, this could raise AllCredentialsFailedError before every
credential had actually been tried, leaving lower-priority credentials unused.
Remove total_max_retries entirely from MultiCredentialVLM and FailoverEmbedder:
credential exhaustion is now decided solely by reaching the end of the chain
(idx >= n), and per-credential retries remain the responsibility of each
underlying instance via its own max_retries. The aggregated error tuple now
records the failing credential index instead of the attempt counter.
Adds a regression test covering more than 10 credentials all being tried.
* refactor: move model-behavior fields off EmbeddingCredential
encoding_format, model_path, cache_dir, enable_fusion, res_level and
max_video_frames describe how a model runs, not which credential is used.
All credentials of a single embedding model share the same model, so these
belong on the parent EmbeddingModelConfig, not on each credential.
Keeping them on the credential forced a `cred.X or config.X` merge in
_create_failover_embedder, which silently dropped explicit falsy values
(enable_fusion=False, res_level=0, max_video_frames=0) and fell back to the
parent value.
Remove these six fields from EmbeddingCredential and read them directly from
the parent config when building per-credential embedders. id/provider/model/
api_key/api_base/api_version/ak/sk/region/host/extra_headers remain
credential-level.
* fix: raise instead of guessing dimension in FailoverEmbedder
FailoverEmbedder.get_dimension() returned a hardcoded 2048 when the first
embedder had no get_dimension(). That path is reached only when wrapping
sparse embedders, which have no fixed dense dimension; returning a fabricated
2048 silently feeds a wrong dimension to callers (e.g. schema creation).
Delegate to the first embedder and raise AttributeError when it has no
get_dimension(), surfacing the misuse instead of hiding it.
* fix: token usage aggregation in failover wrappers
Two issues in the cross-instance token usage merge:
1. Encapsulation: FailoverVLM / MultiCredentialVLM / FailoverEmbedder reached
into other instances' private _token_tracker. Add a public token_tracker
accessor on VLMBase and use it in the VLM mergers.
2. Double counting in FailoverEmbedder: embedders share a process-wide
singleton token tracker (_get_token_tracker), so all wrapped embedders point
at the same object. Merging N identical trackers inflated usage N-fold.
Return a single instance's usage directly instead of merging.
* fix: trip circuit breaker on AUTH errors after error-class split
Splitting 401/403/unauthorized/accountoverdue out of PERMANENT into the new
AUTH class (commit b477b9fd) regressed the circuit breaker: it only tripped
immediately on PERMANENT/QUOTA_EXCEEDED, so auth errors no longer opened the
breaker right away.
For a single embedding instance an auth failure (key invalid / no permission /
overdue) is persistent and retrying is pointless, so the breaker should still
trip immediately. Add ERROR_CLASS_AUTH to the immediate-trip set and update the
classification tests to assert the new AUTH class (403 still trips the breaker).
* fix
* test: bump _last_switch_time when forcing active_idx in ring tests
Without setting _last_switch_time, maybe_failback() retreats to idx 0
immediately because the default 0 timestamp is always older than the
600s timeout, so the unavailable last credential never actually gets
exercised.
* format
* fix
* fix: add dimension valid
* format
* fix: resolve VLM legacy backup primary via _match_provider()
When a legacy config uses ``providers: {openai: {api_key: ...}}`` together
with a ``backup`` VLMConfig, the previous backup-migration branch only
read top-level ``self.provider/self.api_key`` to build legacy-primary,
yielding (provider=None, api_key=None) and an unavailable VLMConfig.
Both primary and backup migration now go through _match_provider() so
``providers``/``default_provider`` based legacy configs are migrated
into VLMCredential with the correct provider/api_key/api_base/etc.
Add regression tests covering primary-providers-dict + backup,
backup-providers-dict, default_provider on backup, and propagation of
extra fields.
* refactor: drop misleading wrapper.is_exhausted from failover wrappers
The ring-retry rewrite of MultiCredentialVLM / FailoverEmbedder does not
call OrderedCredentialSwitcher.on_failure(); each request loops locally
and raises AllCredentialsFailedError on full failure. As a result the
underlying _active_idx is rarely advanced to n, so wrapper-level
is_exhausted would have returned False even when every credential just
failed.
There are no production callers of either property, so remove them
(YAGNI) rather than synthesize an exhausted state from the wrapper side.
The switcher's own is_exhausted stays as a state-machine observation
point used by tests.
Cache async SDK clients per event loop to avoid cross-loop reuse in worker threads.
Move memory vectorization into semantic queue refresh and preserve target sync state for resource updates.
* feat(embedder): expose encoding_format for OpenAI/Azure providers
The OpenAI Python SDK 2.x defaults to encoding_format="base64" so the
client can decode embeddings into native float arrays locally. Some
self-hosted or vendor-fronted OpenAI-compatible gateways cannot
deserialize base64 embedding payloads coming back from upstream models
and silently hang for tens of seconds before returning HTTP 500 (e.g.
gateways that wrap providers like Qwen, GLM, Doubao, etc. behind a
strongly-typed Java SDK).
Add an optional `encoding_format` field on EmbeddingModelConfig that
gets forwarded to OpenAIDenseEmbedder. The field is unset by default,
so existing deployments keep the SDK's default behavior. Users hitting
the base64 incompatibility can set:
"embedding": {
"dense": {
"provider": "openai",
"encoding_format": "float",
...
}
}
Wiring is intentionally limited to provider="openai" and
provider="azure" — the only two factory branches that route to
OpenAIDenseEmbedder for an actual upstream HTTP gateway. Other
providers either don't expose this knob (volcengine/vikingdb/jina/...)
or run against local stacks where the issue cannot occur (ollama).
* test(embedder): improve encoding_format validation error handling
- Add ValidationError import from pydantic for explicit exception handling
- Update test_rejects_unknown_value to assert ValidationError instead of generic Exception
- Improve test specificity by catching the exact validation error type raised by pydantic models
* docs(embedder): complete encoding_format configuration guide
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(security): clean up code scanning and runtime findings
Harden path and logging boundaries, remove noisy cleanup issues,
and keep observability failures from breaking runtime flows.
* fix(security): close werewolf and feishu validation gaps
Block the remaining path traversal bypass in the werewolf demo,
and validate Feishu hosts on the main parse() entry point.
VikingDB embedders accepted is_query but ignored it. Now
VikingDBDenseEmbedder and VikingDBHybridEmbedder accept
query_param/document_param and pass input_type to the API
when non-symmetric mode is configured.
- Add query_param/document_param to VikingDB Dense and Hybrid constructors
- Add _resolve_input_type() to select query vs document param
- Pass input_type in _call_api data items when set
- Wire factory entries to pass config params through
- Sparse embedder unchanged (sparse models are symmetric)
Closes#655
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
* fix(embedder): report configured provider in slow-call logs
Preserve the configured embedding provider in slow-call warnings so OpenAI-compatible backends like Ollama do not show up as unknown or as their transport mode.
* fix log
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
Introduce native async embedding paths across providers, switch async
retrieval/session hotspots to use them, and add a standalone mixed-load
benchmark plus before/after benchmark evidence for the regression.
* fix(bot): respect OPENVIKING_CONFIG_FILE even when file doesn't exist
Previously, bot's config path resolution would fallback to
~/.openviking/ov.conf when OPENVIKING_CONFIG_FILE was set but the
file didn't exist. This was inconsistent with server's behavior,
which treats a missing env-specified config as an error.
In container deployments with OPENVIKING_CONFIG_FILE=/app/ov.conf,
this caused bot to potentially write auto-generated config to
/root/.openviking/ov.conf instead of the intended /app/ov.conf,
leading to config file path mismatches.
Now bot respects the environment variable unconditionally, matching
server's behavior and ensuring both components use the same config
path.
Fixes#1242🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(embedder): support dimension truncation for OpenAI-compatible models
Previously, when users configured a dimension (e.g., 1024) for OpenAI-compatible
embedding models that don't support the 'dimensions' API parameter, the system
would fail with "dimensions is currently not supported" error.
Additionally, when dimension was not configured, the config layer would return
a hardcoded fallback of 2048, but the actual model might return a different
dimension (e.g., 1024), causing dimension validation failures.
This fix implements vector truncation in OpenAIDenseEmbedder:
- Removes the 'dimensions' parameter from API calls (not supported by all models)
- If user configures dimension=1024 and model returns 2048, truncates to 1024
- If no dimension is configured, uses model's native dimension without truncation
- Applies truncation to both single and batch embedding operations
This allows users to:
1. Use OpenAI-compatible models with custom dimensions via truncation
2. Control vector dimensions for storage optimization
3. Avoid dimension mismatch errors between config and actual embeddings
Fixes#1238🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: https://github.com/volcengine/OpenViking/issues/1232
---------
Co-authored-by: openviking <openviking@example.com>
Co-authored-by: Claude <noreply@anthropic.com>
* fix: add models observer info for embedder and rerank
* fix: make build deps
* fix: ov observer
* fix: ov observer
---------
Co-authored-by: openviking <openviking@example.com>
Single-text embedding calls now forward the optional dimension setting to the
OpenAI embeddings API, matching embed_batch() behavior.
Made-with: Cursor
* refactor(model): unify config-driven retry across VLM and embedding
Move retry behavior into shared model-call utilities and config defaults so VLM and embedding providers handle transient failures consistently.
Co-Authored-By: Claude Opus 4.6
* fix
* docs(config): document model retry settings
* feat(retry): add unified transient retry module (#922)
Implements `openviking/models/retry.py` with `is_transient_error`,
`transient_retry`, and `transient_retry_async` — a single config-driven
retry layer replacing scattered per-backend implementations.
Adds 50 unit tests covering classification, backoff, jitter, exhaustion,
and custom predicates.
* feat(vlm): integrate unified retry into VLM backends (#922)
- VLMBase: change max_retries default 2→3, remove max_retries param from
get_completion_async abstract signature
- OpenAI backend: wrap all 4 methods with transient_retry/transient_retry_async,
disable SDK retry (max_retries=0 in client constructors), remove manual
for-loop retry
- VolcEngine backend: same pattern — transient_retry for all methods,
remove manual for-loop retry
- LiteLLM backend: same pattern — transient_retry for all methods,
remove manual for-loop retry
* refactor(vlm): migrate call sites to kwargs, remove max_retries params (#922)
- VLMConfig: default max_retries 2→3, remove max_retries from
get_completion_async signature, switch all wrappers to kwargs
- StructuredVLM (llm.py): remove max_retries from complete_json_async and
complete_model_async, switch all internal calls to kwargs
- memory_react.py: remove max_retries=self.vlm.max_retries (now handled
internally by backend)
- Update test stubs to match new signatures (remove max_retries=0)
* test(vlm): add VLM retry integration tests (#922)
Tests cover OpenAI backend as representative:
- Completion retries on 429, does NOT retry on 401
- Vision completion now retries (was zero before)
- Config max_retries is used (default=3)
- max_retries removed from get_completion_async signature (all backends)
- OpenAI SDK retry disabled (max_retries=0 in client constructors)
* feat(embedding): добавить max_retries в EmbeddingConfig и EmbedderBase
- EmbeddingConfig: новое поле max_retries (default=3) для конфигурации retry
- EmbeddingConfig._create_embedder(): инжектирует max_retries в params["config"]
- EmbedderBase.__init__(): извлекает max_retries из config dict
* feat(embedding): перевести все embedding провайдеры на transient_retry
- OpenAI: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- Volcengine: заменить exponential_backoff_retry на transient_retry, убрать is_429_error
- VikingDB: добавить transient_retry (ранее retry отсутствовал)
- Gemini: отключить SDK HttpRetryOptions (attempts=1), обернуть embed/embed_batch
- MiniMax: отключить urllib3 Retry (total=0), обернуть embed/embed_batch
- Jina: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- Voyage: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- LiteLLM: обернуть litellm.embedding() вызовы
Все провайдеры теперь используют единый transient_retry с is_transient_error
для классификации ошибок. Wrapper размещён ВНУТРИ метода вокруг raw API call,
ДО try/except который конвертирует в RuntimeError.
* test(embedding): тесты retry для embedding провайдеров и backward compat
- test_embedding_retry_integration: OpenAI и VikingDB retry на transient/permanent ошибки
- test_retry_config: VLMConfig и EmbeddingConfig max_retries поля и defaults
- test_backward_compat: exponential_backoff_retry importable, signature unchanged, time-based
* style: ruff format для всех изменённых файлов
* style: fix ruff lint errors (import sorting, unused imports)
* style: format test_embedding_retry_integration.py
* style: remove unused transient_retry import from volcengine_vlm
After upstream refactored sync methods to use run_async(),
only transient_retry_async is needed in volcengine_vlm.py.
* style: fix ruff lint errors in volcengine_vlm.py
- Rename unused loop vars i, item → _i, _item (B007)
- Suppress unused has_images assignment (F841) — upstream code, kept for clarity
* ci: trigger re-run for flaky Windows API integration test
test_fs_tree intermittently returns 500 on windows-latest when
HAS_SECRETS=false — the resource processing pipeline retries
failed embeddings (401 auth errors) in the background, causing
server load that affects the fs/tree endpoint.
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
---------
Co-authored-by: openviking <openviking@example.com>
* feat(embedder): add Cohere dense embedder with embed-v4.0 support
Adds CohereDenseEmbedder using Cohere's Embed API v2.
- Supports embed-v4.0, embed-english-v3.0, embed-multilingual-v3.0
- Server-side dimension reduction for embed-v4.0 (256/512/1024/1536)
- Client-side truncation + renormalization fallback for v3 models
- Asymmetric search via input_type (search_query/search_document)
- Batch embedding with 96-item chunking (Cohere API limit)
- Full factory integration: provider validation, dimension resolution
none
* test(embedder): add unit tests for Cohere embedder
16 tests covering:
- Init validation (api_key required, defaults, model dimensions)
- Dimension handling (v4 server-side, v3 client-side truncation, invalid dims)
- Embedding calls (single, batch, query vs document input_type)
- output_dimension sent for embed-v4.0
- Error handling (API errors → RuntimeError)
- Resource cleanup (close)
none
* feat(rerank): add Cohere rerank-v3.5 support
Extends RerankConfig with provider field and api_key for Cohere.
Adds CohereRerankClient with same interface as VikingDB RerankClient.
HierarchicalRetriever auto-selects rerank backend based on provider.
Config example:
"rerank": {"provider": "cohere", "api_key": "...", "threshold": 0.15}
Quality improvement: META tokenomics query 0.55 → 0.77 relevance score.
none
* test(rerank): add unit tests for Cohere reranker
9 tests covering:
- Rerank batch scoring with index-to-order mapping
- Empty input handling
- API error graceful fallback (returns None)
- Original order preservation from Cohere's sorted response
- Resource cleanup
- RerankConfig provider auto-detection (cohere/vikingdb/empty)
none
* perf(retrieve): increase GLOBAL_SEARCH_TOPK from 5 to 10
More vector candidates for reranker to evaluate = better precision.
With Cohere rerank-v3.5, 10 candidates gives the cross-encoder enough
material to find the best match without excessive latency.
none
* refactor: unify rerank dispatch — route all providers through RerankClient.from_config()
Cohere was special-cased in hierarchical_retriever.py while openai/litellm
went through the centralized RerankClient.from_config() dispatch. This commit
adds CohereRerankClient.from_config() and routes it through the same path.
Also fixes a bug where from_config() used config.provider directly instead
of _effective_provider(), which meant auto-detected providers (e.g. api_key
without explicit provider="cohere") would not dispatch correctly.
none
Add jina-code-embeddings-1.5b and jina-code-embeddings-0.5b to
JINA_MODEL_DIMENSIONS so dimension validation works correctly.
Add _raise_task_error() helper that catches 422 validation errors
mentioning 'task' and produces a user-friendly error message guiding
users to set query_param and document_param in their embedding config.
Fixes#912.
* feat(embedding): add litellm as embedding provider
Adds LiteLLM as a new embedding provider, bringing embedding parity with
the VLM layer which already supports litellm. This enables users to route
embedding requests through OpenRouter, Ollama, vLLM, and any other
OpenAI-compatible endpoint via litellm's unified interface.
Closes#847
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address review feedback on litellm embedding provider
1. Move os.environ.setdefault from module-level to __init__ to avoid
mutating the process environment on import
2. Require dimension as mandatory — removes the probe API call during
construction that caused surprise billable requests and silent fallbacks
3. Add None guard for LiteLLMDenseEmbedder in factory to give a clear
error when litellm is not installed
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(gemini): add GeminiDenseEmbedder with is_query routing per #702 pattern
- GeminiDenseEmbedder: accept query_param/document_param, use is_query
in embed() and embed_batch() to select task_type at call time
- EmbeddingConfig: add Gemini provider, factory, validation, dimension
- No get_query_embedder/get_document_embedder/_get_contextual_embedder
(removed in #702; embed(is_query=True/False) is the pattern)
- Tests use embed(text, is_query=True/False) pattern throughout
- Rebased onto current upstream/main
* fix(gemini): remove task_type config field, fix conditional import for CI
- Remove task_type from EmbeddingModelConfig (query_param/document_param suffice)
- Wrap GeminiDenseEmbedder import in try/except (google-genai is optional)
- Update tests for removed field
* feat: add Azure OpenAI support for embedding and VLM
Add `azure` as a first-class provider for both embedding models and VLM,
using the official `openai.AzureOpenAI` / `openai.AsyncAzureOpenAI` clients.
Changes:
- embedder: OpenAIDenseEmbedder now accepts `provider` and `api_version`
params; when provider is "azure", initializes AzureOpenAI client with
azure_endpoint and api_version
- VLM: OpenAIVLM switches between OpenAI and AzureOpenAI clients based
on the provider field; VLMFactory maps "azure" to OpenAIVLM
- config: EmbeddingModelConfig and VLMConfig gain `api_version` field;
embedding_config adds ("azure", "dense") factory entry with validation
- registry: add "azure" to VALID_PROVIDERS
- docs: update README_CN.md with Azure provider table entry, config
template fields, usage examples, and full config sample
Made-with: Cursor
* refactor: address PR review — extract helper, add validation, shared constant
- Extract `_build_openai_client_kwargs()` helper in openai_vlm.py to
eliminate duplicated Azure/OpenAI client construction across
get_client() and get_async_client()
- Add `api_base` validation for Azure provider in VLM client methods
(was missing, unlike the embedder which already validated it)
- Extract `DEFAULT_AZURE_API_VERSION` constant in registry.py and
reference it from both embedder and VLM, so future Azure API version
bumps only require a single change
- Note: removing the hardcoded `self.provider = "openai"` override in
OpenAIVLM.__init__ (done in the previous commit) is an intentional
bug fix — it allows VLMBase.provider to correctly reflect the
configured value (e.g. "azure"), which fixes token usage tracking
Made-with: Cursor
* feat: enable minimax embedding and adapt to separate query/document parameter configuration
* feat: enable minimax embedding and adapt to separate query/document parameter configuration
---------
Co-authored-by: xiaogang.zhou <xiaogang.zhou@bytedance.com>
* feat(storage): add transaction support with journal, undo, and crash recovery
Implement a full transaction system for VikingFS storage operations including
write-ahead journal, path locking, undo/rollback, context manager API, and
crash recovery. Includes comprehensive tests and documentation.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* test(transaction): add e2e rollback tests for mv and multi-step operations
Add end-to-end tests covering rollback scenarios that were missing:
- mv rollback: file moved back to original location on failure
- mv commit: file persists at new location
- Multi-step rollback: mkdir + write + mkdir all reversed in order
- Partial step rollback: only completed entries are reversed
- Nested directory rollback: child removed before parent
- Best-effort rollback: single step failure does not block others
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* feat(storage): add transaction support with path locking and journal
Implement transaction system for VikingFS with ACID-like guarantees:
- TransactionManager with configurable lock timeout and journal-based recovery
- PathLock supporting point, subtree, and mv lock modes
- Refactor VikingFS mv to use cp+rm to prevent lock files from being carried
- Fix stale lock detection returning false for missing lock files
- Update ragas eval to use LangchainLLMWrapper
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: tests
* fix(transaction): fix rollback and race condition bugs
- Reconstruct RequestContext from undo params for vectordb_delete/update_uri
rollback (previously skipped silently due to missing ctx)
- Serialize ctx fields into undo params in rm/mv operations
- Fix Phase 1 undo path to target archive dir instead of session root
- Remove Phase 2 fs_write_new undo (overwrites are idempotent, checkpoint
handles recovery)
- Add ancestor SUBTREE recheck after lock creation in acquire_subtree
- Move _collect_uris inside TransactionContext in rm/mv to close race window
- Log journal persistence failures instead of silently swallowing
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor(transaction): make TransactionManager required and rewrite tests with real backends
Remove all optional/fallback code paths where tx_manager could be None. get_transaction_manager()
now raises RuntimeError if not initialized. Fix undo rollback to reconstruct ctx for vectordb_upsert
and use correct agent_id default. Replace mock-based transaction tests with integration tests using
real AGFS and VectorDB backends.
* refactor(transaction): make rollback fully async and unify session commit path
- Convert execute_rollback/rollback_entry to async, removing sync run_async wrappers
- Unify Session.commit() to delegate to commit_async(), removing duplicate phase methods
- Fix SUBTREE lock to conflict with ancestor SUBTREE locks (was previously missing)
- Fix mv lock mode: directory moves now use SUBTREE on both source and destination
- Replace deprecated asyncio.get_event_loop() with get_running_loop()
- Remove max_parallel_locks config option
- Update docs (en/zh) and tests to match new async rollback signatures
* fix: tests
* refactor(transaction): simplify session commit and add redo-based crash recovery
Session commit no longer wraps archive phase in a transaction. Phase 2 uses
redo semantics so crashed memory-extraction can be replayed from archive.
PathLock stale-lock cleanup no longer redundantly re-checks timeout.
Semantic processor vectorization runs concurrently via asyncio.gather.
* fix: transaction
* fix: UserIdentifier
* refactor(transaction): replace undo-based transaction manager with lightweight lock + redo-log
Remove the heavyweight TransactionManager/Journal/UndoEntry system (~4000 lines) and
replace it with a simpler architecture: LockManager for path locking, LockContext as
the async context manager, LockHandle/LockOwner protocol, and a RedoLog for crash
recovery of session_memory operations. VikingFS rm/mv now use inline error handling
instead of rollback semantics. Updated docs, observers, and tests accordingly.
Co-Authored-By: Claude Opus 4.6
* fix(transaction): remove checkpoint dead code, fix TOCTOU race, clarify mv lock param
- Remove unused _write_checkpoint/_write_checkpoint_async/_read_checkpoint
from Session (superseded by redo-log)
- Re-resolve URI inside lock in resource_processor Phase 3.5 to prevent
concurrent add_resource calls from resolving to the same final_uri
- Rename acquire_mv dst_path to dst_parent_path with docstring to clarify
that callers pass the destination parent directory
* fix: path
* fix: resource lock
* fix: test
* docs: update
* fix: tests
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
* feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF)
Native text + multimodal (image, video, audio, PDF) embedding via `gemini-embedding-2-preview` (google-genai 1.67.0). Additive provider pattern — Volcengine remains the default; Gemini is opt-in via `provider: "gemini"` in `ov.conf`.
- **Model**: `gemini-embedding-2-preview`
- **Input**: text, image, video, audio, PDF (17 MIME types)
- **Output dimension**: 128–3072 (default: **3072**, recommended: 768 / 1536 / 3072)
- **Input token limit**: **8,192 tokens**
- **Supported MIME types**: `image/jpeg`, `image/png`, `image/gif`, `image/webp`, `audio/mpeg`, `audio/mp3`, `audio/wav`, `audio/ogg`, `audio/flac`, `video/mp4`, `video/mpeg`, `video/mov`, `video/avi`, `video/webm`, `video/wmv`, `video/3gpp`, `application/pdf`
- Gemini Embedding 2 Multimodal Support: Introduced a new GeminiDenseEmbedder to support native text and multimodal (image, video, audio, PDF) embedding using the gemini-embedding-2-preview model. This is an opt-in provider via configuration.
- Extended Queue Pipeline for Multimodal Content: The EmbeddingMsg now carries media_uri and media_mime_type to facilitate multimodal content processing. The TextEmbeddingHandler.on_dequeue() method was updated to read raw bytes from viking_fs and call embed_multimodal() when applicable.
- End-to-End Configuration and Security: The EmbeddingConfig now registers the 'gemini' provider with a task_type field. A critical security validation was added to ensure media_uri matches context_data['uri'] before file reads, preventing forged queue messages from accessing arbitrary files. If validation fails or multimodal embedding fails, it falls back to text embedding.
- Multimodal Content Representation: A new ModalContent dataclass was introduced to represent media references, including MIME type, URI, and optional raw data, enabling the Vectorize object to encapsulate both text and media for embedding
* feat: Add asynchronous batch embedding with concurrency control to Gemini embedder.
* Reduce scope to use GeminiDenseEmbedder as only text embed
Add `extra_headers` field to `EmbeddingModelConfig` and `OpenAIDenseEmbedder`,
enabling custom HTTP headers (e.g., `HTTP-Referer`, `X-Title` required by
OpenRouter) to be passed as `default_headers` to the OpenAI client.
Also fix a dead-code bug in `OpenAIDenseEmbedder.__init__` where an
unconditional `raise ValueError("api_key is required")` prevented the
more nuanced check that allows missing `api_key` when `api_base` is set
(e.g., local OpenAI-compatible servers).
Closes#675
* feat(embedding): add voyage dense embedder
Add first-class Voyage dense embedding support with a dedicated embedder,
provider validation, model-aware default dimensions, and focused tests.
Keep the configuration surface intentionally narrow:
- use the existing dimension field and map it to Voyage's
output_dimension request field
- do not expose Voyage-only output_dtype
- do not expose query/document mode until OpenViking has separate
index/query embedder configuration
This keeps the PR aligned with the current OpenViking architecture,
which stores and retrieves dense float vectors through a single dense
embedder configuration.
Verification:
- .venv/bin/python -m pytest tests/unit/test_voyage_embedder.py tests/unit/test_embedding_config_voyage.py --noconftest -o addopts='' -q
- .venv/bin/python -m pytest tests/misc/test_config_validation.py -o addopts='' -q
- .venv/bin/ruff check openviking/models/embedder/voyage_embedders.py openviking_cli/utils/config/embedding_config.py tests/unit/test_embedding_config_voyage.py tests/unit/test_voyage_embedder.py
* refactor(embedder): inline Voyage extra body
Remove the trivial Voyage-specific helper and inline the output_dimension payload construction at the two call sites.
This keeps the request shape unchanged while making the embedder implementation more direct.
* feat(embedding): add nonsymmetric embedding support for query/document
- Add input_type parameter to OpenAIDenseEmbedder
- Add get_query_embedder() and get_document_embedder() to EmbeddingConfig
- Support different embedding strategies for queries and documents
- Improve error handling and default value processing
- Fix import ordering and remove unused imports
* fix(service): implement PR feedback for core.py
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
* fix: address PR review comments - add telemetry tests and inline VikingDBManager init
* refactor(embedder): move context param handling into embedder classes
Per reviewer feedback, each embedder now decides its own context-specific
parameter mapping (input_type for OpenAI, task for Jina) based on a
'context' arg ('query'/'document'/None), instead of having EmbeddingConfig
pass pre-computed values.
- Unified config fields: replaced 4 separate fields (input_type_query,
input_type_document, task_query, task_document) with 2 unified fields
(query_param, document_param) that work for all providers
- OpenAI is symmetric by default; non-symmetric mode is activated
implicitly when query_param or document_param is set
- Jina is non-symmetric by default; task is always sent unless context=None
- Updated tests to use new API
Note: Official OpenAI models are symmetric and do not support input_type.
Non-symmetric mode is only supported by OpenAI-compatible third-party
models (e.g., BGE-M3, Jina, Cohere).
---------
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
When tiktoken is unavailable, the fallback `len(text) // 3` severely
underestimates tokens for CJK text (Chinese/Japanese/Korean characters
are ~1-2 tokens each, not 0.33). This causes text exceeding the 8192-token
API limit to bypass chunking, resulting in BadRequestError.
Use `max(len(text) // 3, len(text.encode("utf-8")) // 4)` instead, which
picks the more conservative estimate. For ASCII-heavy text the char-based
estimate still wins; for CJK text the byte-based estimate correctly
produces ~0.75 tokens per character.
Fixes#616, fixes#634
Signed-off-by: JiangNan <1394485448@qq.com>
- Add 'ollama' as a supported embedding provider
- Ollama runs locally via OpenAI-compatible API, no API key required
- Allow OpenAI provider to work without api_key when api_base is set
(supports local OpenAI-compatible servers like vLLM, LocalAI)
- Add configuration example and tests for Ollama provider
This enables fully local embedding deployment without cloud API keys.