Commit Graph
61 Commits
Author SHA1 Message Date
Hao Zheandzhiheng.liu 797406c381 fix(sdk): consolidate model, client, and evaluation correctness (#3550)
* fix(models): route multimodal inputs to DashScope embed_content

Sweep findings: A-09. Preserve image parts through the shared embedding entrypoints.

(cherry picked from commit 50b20d7ca9)

* fix(models): constrain DashScope multimodal routing

Sweep findings: A-09. Preserve text mode and require a single fused vector for multipart input.

(cherry picked from commit 77784429d4)

* fix(models): respect DashScope Qwen fusion parameters

Sweep finding A-09

(cherry picked from commit aafc702b20)

* fix(review): preserve tongyi multipart compatibility

Addresses blocking review finding on #3400.

(cherry picked from commit 96844047c5)

* fix(eval): flush queued records on stop and adapt RAG pipeline to FindResult

Sweep findings: B-05, B-12. Drain recorder queues through the sentinel and consume current retrieval result objects.

(cherry picked from commit fecdbeb6d4)

* fix(sdk/python): support sync client inside a running event loop

Sweep findings: F-09. Run sync wrappers on one persistent worker loop with result and exception propagation.

(cherry picked from commit ba0faadc87)

* fix(sdk/python): preserve cancellation and fork safety

Sweep finding: F-09. Preserve original cancellation errors and reset worker synchronization after fork.

(cherry picked from commit a19e3c95e9)

* feat(sdk/go): add tags filter and relations API for parity

Sweep findings: F-11, F-12. Expose existing server capabilities consistently to Go callers.

(cherry picked from commit 8af0c2cb36)

* fix(sdk/python): runnable quickstarts, correct migrate payload, explicit timeout precedence

Sweep findings: F-01, F-02, F-14. Initialize documented clients and preserve Python SDK request/config semantics.

(cherry picked from commit d9a64f5665)

---------

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
2026-07-28 18:11:30 +08:00
Colter Dahlberg e97b74d2e5 feat(embedder): support extra_body passthrough in OpenAI embedder config (#3495)
* feat(embedder): support extra_body passthrough in OpenAI embedder config

Adds optional `extra_body` (dict) to the OpenAI dense embedder, merged
into every embeddings.create call. Motivating use case: OpenRouter
provider routing ({"provider": {"sort": "latency"}}) — default routing
shows p90=35s/max=127s tail latency that kills interactive recall
(A/B: sorted routing is consistently sub-second).

Explicit query_param/document_param keys still take precedence on
conflict.

* feat(config): wire extra_body through embedding config layer

Add optional extra_body field to EmbeddingModelConfig and pass it to
OpenAIDenseEmbedder for the openai/azure providers, including the
multi-credential failover merge (parent-level model-behavior field).

* docs(config): document embedding extra_body with OpenRouter routing example

* docs(config): restore concrete host/cors_origins values in EN full schema
2026-07-24 15:07:44 +08:00
Qin Haojie d47f2106ee refactor: remove unused and deprecated APIs (#3272)
Delete dead compatibility paths and test-only helpers so unsupported APIs do not remain as accidental contracts.
2026-07-16 10:49:56 +08:00
Qin Haojie ca70bc0649 refactor(embedding): remove unused batch APIs (#3260) 2026-07-15 16:44:57 +08:00
Qin Haojie 07708225cb fix(volcengine): forward model request headers (#2909)
Ensure Volcengine VLM and embedding calls propagate configured headers and default the client request id header for service traffic.
2026-07-01 14:31:39 +08:00
87329714dd feat(grep): integrate VikingDB bm25 keyword search for grep engine (#2144)
* feat(grep): integrate VikingDB bm25 keyword search for grep engine

* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)

* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison

* fix(schema): upsert data to vikingdb lack of content

* chore: add benchmark for retrieval

* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs

* fix(benchmark): sub uri args; add report

* refactor: code format by ruff

* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf

* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search

* fix: adjust benchmark scripts

* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls

* refactor: new benchmark

* fix: step1 add resource by real code data

* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex

* optimize (benchmark): adjust keywords and ground truth for testing

* fix: truncate 64KB for content field

* optimize: effectiveness add resource plainly

* optimize: change param use of SearchByKeywords from "keywords" to "query"

* optimize(benchmark): refactor effectiveness scripts

* optimize: ensure raw data for content field

* optimize: fulltext analyzer's stop-words only use symbols

* fix: adapt to new ov cli for benchmark

* optimize: reuse file content to avoid re-read AGFS file

* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts

* optimize: benchmark client timeout

* update README

* fix: rm unused param

* fix: default values in docs

* optimize: increase truncate byte size to 1MB for content field for VikingDB

* fix(logger): harden queued stream logging (#2786)

* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock

When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.

During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.

Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.

Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.

Closes: #2752

* fix(logger): harden queued stream logging

---------

Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
2026-06-24 18:46:02 +08:00
dingben be77b2dac1 feat: implement multi-credential priority call design (#2468)
* feat: implement multi-credential priority call design

Add OrderedCredentialSwitcher for N-credential failover, MultiCredentialVLM and FailoverEmbedder for credential switching, support automatic migration from legacy backup config

* refactor: deduplicate AllCredentialsFailedError and fix logging in switcher

* chore: fix code formatting for lint compliance

* chore: fix remaining code formatting

* fix: update FailoverEmbedder for multimodal API compatibility

* fix

* fix

* refactor: split error classes and fix fail-fast in credential failover

Separate the monolithic PERMANENT error class so the switcher reacts to
the actual root cause:

- 400 (request-level parameter error) -> PERMANENT, fail-fast: same
  request fails on every credential of the same model.
- 401/403/unauthorized/accountoverdue -> new AUTH class: credential-level,
  advances to the next credential in multi-credential mode.
- new CONTENT_SAFETY class (moderation rejections) and INPUT_TOO_LARGE ->
  fail-fast: switching credentials cannot help.

classify_api_error now checks CONTENT_SAFETY before PERMANENT so a
moderation message containing "400" is not misclassified. On fail-fast the
failover wrappers re-raise the original exception (preserving type/info)
instead of wrapping it, so callers can react (e.g. truncate on
input_too_large). AllCredentialsFailedError is reserved for chain
exhaustion. The legacy PrimaryBackupSwitcher also switches on AUTH to keep
existing backup behavior.

* fix: make get_active_index side-effect free in credential switcher

get_active_index() previously mutated state: when a failback threshold was
met it would decrement the active index. Because observability properties
(active_credential_index / active_credential_id) call it, merely reading the
current credential for logging or metrics could accidentally advance the
failback state machine.

Split the concern: get_active_index() is now a pure read, and a new public
maybe_failback() performs the one-step failback (and logs an info line when
the active credential index changes). The request loops in MultiCredentialVLM
and FailoverEmbedder call maybe_failback() at the top of each attempt, so the
failback behavior is unchanged while pure reads no longer have side effects.

* fix: drop global total_max_retries cap from credential failover

The failover loops exited on `idx >= n OR total_attempts >= total_max_retries`
(default 10). With more than 10 credentials, or when failback churn inflated
the attempt count, this could raise AllCredentialsFailedError before every
credential had actually been tried, leaving lower-priority credentials unused.

Remove total_max_retries entirely from MultiCredentialVLM and FailoverEmbedder:
credential exhaustion is now decided solely by reaching the end of the chain
(idx >= n), and per-credential retries remain the responsibility of each
underlying instance via its own max_retries. The aggregated error tuple now
records the failing credential index instead of the attempt counter.

Adds a regression test covering more than 10 credentials all being tried.

* refactor: move model-behavior fields off EmbeddingCredential

encoding_format, model_path, cache_dir, enable_fusion, res_level and
max_video_frames describe how a model runs, not which credential is used.
All credentials of a single embedding model share the same model, so these
belong on the parent EmbeddingModelConfig, not on each credential.

Keeping them on the credential forced a `cred.X or config.X` merge in
_create_failover_embedder, which silently dropped explicit falsy values
(enable_fusion=False, res_level=0, max_video_frames=0) and fell back to the
parent value.

Remove these six fields from EmbeddingCredential and read them directly from
the parent config when building per-credential embedders. id/provider/model/
api_key/api_base/api_version/ak/sk/region/host/extra_headers remain
credential-level.

* fix: raise instead of guessing dimension in FailoverEmbedder

FailoverEmbedder.get_dimension() returned a hardcoded 2048 when the first
embedder had no get_dimension(). That path is reached only when wrapping
sparse embedders, which have no fixed dense dimension; returning a fabricated
2048 silently feeds a wrong dimension to callers (e.g. schema creation).

Delegate to the first embedder and raise AttributeError when it has no
get_dimension(), surfacing the misuse instead of hiding it.

* fix: token usage aggregation in failover wrappers

Two issues in the cross-instance token usage merge:

1. Encapsulation: FailoverVLM / MultiCredentialVLM / FailoverEmbedder reached
   into other instances' private _token_tracker. Add a public token_tracker
   accessor on VLMBase and use it in the VLM mergers.

2. Double counting in FailoverEmbedder: embedders share a process-wide
   singleton token tracker (_get_token_tracker), so all wrapped embedders point
   at the same object. Merging N identical trackers inflated usage N-fold.
   Return a single instance's usage directly instead of merging.

* fix: trip circuit breaker on AUTH errors after error-class split

Splitting 401/403/unauthorized/accountoverdue out of PERMANENT into the new
AUTH class (commit b477b9fd) regressed the circuit breaker: it only tripped
immediately on PERMANENT/QUOTA_EXCEEDED, so auth errors no longer opened the
breaker right away.

For a single embedding instance an auth failure (key invalid / no permission /
overdue) is persistent and retrying is pointless, so the breaker should still
trip immediately. Add ERROR_CLASS_AUTH to the immediate-trip set and update the
classification tests to assert the new AUTH class (403 still trips the breaker).

* fix

* test: bump _last_switch_time when forcing active_idx in ring tests

Without setting _last_switch_time, maybe_failback() retreats to idx 0
immediately because the default 0 timestamp is always older than the
600s timeout, so the unavailable last credential never actually gets
exercised.

* format

* fix

* fix: add dimension valid

* format

* fix: resolve VLM legacy backup primary via _match_provider()

When a legacy config uses ``providers: {openai: {api_key: ...}}`` together
with a ``backup`` VLMConfig, the previous backup-migration branch only
read top-level ``self.provider/self.api_key`` to build legacy-primary,
yielding (provider=None, api_key=None) and an unavailable VLMConfig.

Both primary and backup migration now go through _match_provider() so
``providers``/``default_provider`` based legacy configs are migrated
into VLMCredential with the correct provider/api_key/api_base/etc.

Add regression tests covering primary-providers-dict + backup,
backup-providers-dict, default_provider on backup, and propagation of
extra fields.

* refactor: drop misleading wrapper.is_exhausted from failover wrappers

The ring-retry rewrite of MultiCredentialVLM / FailoverEmbedder does not
call OrderedCredentialSwitcher.on_failure(); each request loops locally
and raises AllCredentialsFailedError on full failure. As a result the
underlying _active_idx is rarely advanced to n, so wrapper-level
is_exhausted would have returned False even when every credential just
failed.

There are no production callers of either property, so remove them
(YAGNI) rather than synthesize an exhausted state from the wrapper side.
The switcher's own is_exhausted stays as a state-machine observation
point used by tests.
2026-06-15 15:37:03 +08:00
Haoyu Zhangandzhanghaoyu.la c675b1732e feat: 计算图片embed向量时支持图文embed,通过配置参数image_vectorization决定embed模式(summary_only,image_only,image_and_summary) (#2460)
Co-authored-by: zhanghaoyu.la <zhanghaoyu.la@bytedance.com>
2026-06-08 10:25:38 +08:00
Qin Haojie c72c8a87e8 fix: apply embedding input limits in embedder (#2266) 2026-05-27 20:23:40 +08:00
Qin Haojie cfb19a50ca fix(storage): isolate async clients and semantic refresh (#2168)
Cache async SDK clients per event loop to avoid cross-loop reuse in worker threads.
Move memory vectorization into semantic queue refresh and preserve target sync state for resource updates.
2026-05-21 17:24:52 +08:00
Simon Shiandqin-ctx 72b407ffa8 feat(embedder): expose encoding_format for OpenAI/Azure providers (#2092)
* feat(embedder): expose encoding_format for OpenAI/Azure providers

The OpenAI Python SDK 2.x defaults to encoding_format="base64" so the
client can decode embeddings into native float arrays locally. Some
self-hosted or vendor-fronted OpenAI-compatible gateways cannot
deserialize base64 embedding payloads coming back from upstream models
and silently hang for tens of seconds before returning HTTP 500 (e.g.
gateways that wrap providers like Qwen, GLM, Doubao, etc. behind a
strongly-typed Java SDK).

Add an optional `encoding_format` field on EmbeddingModelConfig that
gets forwarded to OpenAIDenseEmbedder. The field is unset by default,
so existing deployments keep the SDK's default behavior. Users hitting
the base64 incompatibility can set:

    "embedding": {
      "dense": {
        "provider": "openai",
        "encoding_format": "float",
        ...
      }
    }

Wiring is intentionally limited to provider="openai" and
provider="azure" — the only two factory branches that route to
OpenAIDenseEmbedder for an actual upstream HTTP gateway. Other
providers either don't expose this knob (volcengine/vikingdb/jina/...)
or run against local stacks where the issue cannot occur (ollama).

* test(embedder): improve encoding_format validation error handling

- Add ValidationError import from pydantic for explicit exception handling
- Update test_rejects_unknown_value to assert ValidationError instead of generic Exception
- Improve test specificity by catching the exact validation error type raised by pydantic models

* docs(embedder): complete encoding_format configuration guide

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-05-18 14:13:03 +08:00
Jiahui Zhou 3bcefb298d fix: unify runtime loggers with openviking logger (#1981) 2026-05-12 11:26:30 +08:00
chenjw 44d3cc41b1 Feat/memory isolation 支持群聊模式 (#1711) 2026-05-06 10:45:06 +08:00
baojun-zhangandMaojiaSheng 17d2c5603e feat(observability): unify observability context && support otel && etc. (#1666)
* feat(observability): unify OTLP metrics export, log/trace context, and telemetry bridging
- - Add OTLP metrics http/grpc exporter that pushes MetricRegistry snapshots
- - Decouple telemetry response payload from telemetry collection; always finish() and bridge summary to metrics
- - Unify observability config under server.observability (metrics/traces/logs siblings); update ov.conf.example and docs (zh/en)
- - Improve log/trace correlation via structured context injection
- - Add/adjust tests for exporter lifecycle, config loader, metrics/telemetry runtime
- BREAKING CHANGE: remove legacy telemetry.* config path; use server.observability.*

* feat(observability): import Status/StatusCode for LogToSpanEventFilter

* feat(observability): fix check issue

* feat(observability): format code

---------

Co-authored-by: MaojiaSheng <shengmaojia@bytedance.com>
2026-04-24 21:32:25 +08:00
Hao ZheandZayn Jarvis 01403312ea feat(vlm): add Codex, Kimi, and GLM VLM support (#1444)
* feat(vlm): add Codex OAuth-backed VLM setup and docs

* fix(codex): address PR review follow-up issues

* feat(vlm): add Kimi and GLM backends

* refactor(vlm): simplify codex auth flow and docs

* fix: update code comments and doctor validation

* chore: update uv.lock after merge

* Refine Codex auth flow and VLM backend integrations

* Take over mirrored Codex auth on refresh

* feat(vlm): refine provider setup and auth flow

* style: format VLM and setup files

* style: fix lint import ordering

* fix(codex): harden auth refresh and disable streaming

* fix(codex): translate tool history and refresh auth safely

* style(lint): fix changed-file ruff violations

* fix(init): refine cloud VLM setup prompts

* style(lint): format setup wizard changes

---------

Co-authored-by: Zayn Jarvis <zhiheng.liu@bytedance.com>
2026-04-22 11:01:23 +08:00
Qin Haojie 38c324bc97 fix(security): clean up code scanning and runtime findings (#1596)
* fix(security): clean up code scanning and runtime findings

Harden path and logging boundaries, remove noisy cleanup issues,
and keep observability failures from breaking runtime flows.

* fix(security): close werewolf and feishu validation gaps

Block the remaining path traversal bypass in the werewolf demo,
and validate Feishu hosts on the main parse() entry point.
2026-04-21 10:06:46 +08:00
A0nameless0man 147f1e3bd3 feat(embedder): add DashScope embedding provider (#1535)
* feat(embedder): 实现 DashScopeDenseEmbedder 双模式嵌入器

* test(embedder): 添加 DashScope 单元测试和集成测试

* feat(config): 添加 DashScope provider 支持、验证和 factory 注册

* docs(embedding): add DashScope provider configuration guide

* fix(embedder): 关闭 DashScope embedder 的 sync/async OpenAI 客户端避免资源泄漏

close() 方法原先仅关闭 httpx 客户端,遗漏了 sync OpenAI 客户端和
async OpenAI 客户端。将两个 async 客户端的清理逻辑合并为单个协程,
复用 event-loop-aware 模式统一处理运行中/无事件循环两种场景。
2026-04-17 18:09:43 +08:00
baojun-zhang 629fc241e4 feat(metric): add token-full-cycle metric (#1488)
* feat(metric): add token full-cycle metric && support token dashboard && optimize metric guide && deprecated 'server.telemetry.prometheus.enabled' configuration && change metric config to observability.metric

* feat(metric): add token full-cycle metric && support token dashboard && optimize metric guide && deprecated 'server.telemetry.prometheus.enabled' configuration && change metric config to observability.metric

* feat(metric): add debug log

* feat(metric): format code

* feat(metric): format code
2026-04-16 18:22:40 +08:00
066b4aac5e feat(embedding): surface non-symmetric embedding config for VikingDB provider (#1110)
VikingDB embedders accepted is_query but ignored it. Now
VikingDBDenseEmbedder and VikingDBHybridEmbedder accept
query_param/document_param and pass input_type to the API
when non-symmetric mode is configured.

- Add query_param/document_param to VikingDB Dense and Hybrid constructors
- Add _resolve_input_type() to select query vs document param
- Pass input_type in _call_api data items when set
- Wire factory entries to pass config params through
- Sparse embedder unchanged (sparse models are symmetric)

Closes #655

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-04-16 17:33:21 +08:00
Mingjian QueandGPT-5.4 3b0ac8a0d4 feat: add local llama-cpp embedding support (#1388)
* feat: local llama.cpp embedding support, setup wizard, lazy storage imports, and config singleton deadlock fix

Made-with: Cursor
Co-authored-by: GPT-5.4 <noreply@openai.com>

* fix: remove unsupported custom local gguf setup

---------

Co-authored-by: GPT-5.4 <noreply@openai.com>
2026-04-15 12:03:40 +08:00
bot-of-qin-ctxandqin-ctx 044e259163 fix(embedder): report configured provider in slow-call logs (#1403)
* fix(embedder): report configured provider in slow-call logs

Preserve the configured embedding provider in slow-call warnings so OpenAI-compatible backends like Ollama do not show up as unknown or as their transport mode.

* fix log

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-04-13 11:44:52 +08:00
Roccoon f9867beffd fix(embedder): initialize async client state in VolcengineSparseEmbedder (#1362) 2026-04-11 20:53:05 +08:00
MaojiaShengandopenviking 3e6d7ffe49 fix: openai like embedding models fix, no more matryoshka error (#1350)
Co-authored-by: openviking <openviking@example.com>
2026-04-10 15:44:42 +08:00
MaojiaShengandopenviking adc2276bb3 fix litellm embedding dimension issues (#1323)
Co-authored-by: openviking <openviking@example.com>
2026-04-09 12:49:40 +08:00
Qin Haojie 804de2a2ad fix(embedder): reduce async contention in session flows (#1301)
Introduce native async embedding paths across providers, switch async
retrieval/session hotspots to use them, and add a standalone mixed-load
benchmark plus before/after benchmark evidence for the regression.
2026-04-08 16:31:14 +08:00
4c175f2536 fix: #1238 and #1242 and #1232 (#1243)
* fix(bot): respect OPENVIKING_CONFIG_FILE even when file doesn't exist

Previously, bot's config path resolution would fallback to
~/.openviking/ov.conf when OPENVIKING_CONFIG_FILE was set but the
file didn't exist. This was inconsistent with server's behavior,
which treats a missing env-specified config as an error.

In container deployments with OPENVIKING_CONFIG_FILE=/app/ov.conf,
this caused bot to potentially write auto-generated config to
/root/.openviking/ov.conf instead of the intended /app/ov.conf,
leading to config file path mismatches.

Now bot respects the environment variable unconditionally, matching
server's behavior and ensuring both components use the same config
path.

Fixes #1242

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(embedder): support dimension truncation for OpenAI-compatible models

Previously, when users configured a dimension (e.g., 1024) for OpenAI-compatible
embedding models that don't support the 'dimensions' API parameter, the system
would fail with "dimensions is currently not supported" error.

Additionally, when dimension was not configured, the config layer would return
a hardcoded fallback of 2048, but the actual model might return a different
dimension (e.g., 1024), causing dimension validation failures.

This fix implements vector truncation in OpenAIDenseEmbedder:
- Removes the 'dimensions' parameter from API calls (not supported by all models)
- If user configures dimension=1024 and model returns 2048, truncates to 1024
- If no dimension is configured, uses model's native dimension without truncation
- Applies truncation to both single and batch embedding operations

This allows users to:
1. Use OpenAI-compatible models with custom dimensions via truncation
2. Control vector dimensions for storage optimization
3. Avoid dimension mismatch errors between config and actual embeddings

Fixes #1238

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

* fix: https://github.com/volcengine/OpenViking/issues/1232

---------

Co-authored-by: openviking <openviking@example.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-04-06 13:53:24 +08:00
MaojiaShengandopenviking d739f742d7 fix: ov status shows embedding and rerank models usage (#1191)
* fix: add models observer info for embedder and rerank

* fix: make build deps

* fix: ov observer

* fix: ov observer

---------

Co-authored-by: openviking <openviking@example.com>
2026-04-03 09:00:38 +08:00
LinQiang391 a932b51eeb fix(embedder): pass dimensions in OpenAIDenseEmbedder.embed() (#1183)
Single-text embedding calls now forward the optional dimension setting to the
OpenAI embeddings API, matching embed_batch() behavior.

Made-with: Cursor
2026-04-02 19:35:10 +08:00
Qin HaojieandClaude Opus 4.6 cad9ec1408 refactor(model): unify config-driven retry across VLM and embedding (#926)
* refactor(model): unify config-driven retry across VLM and embedding

Move retry behavior into shared model-call utilities and config defaults so VLM and embedding providers handle transient failures consistently.

Co-Authored-By: Claude Opus 4.6

* fix

* docs(config): document model retry settings
2026-03-31 16:17:28 +08:00
bot-of-qin-ctx fe6c465d09 Revert "feat: unify config-driven retry across VLM and embedding (#1049)" (#1113)
This reverts commit a092d64867.
2026-03-31 15:01:04 +08:00
Sergey Nemesh a092d64867 feat: unify config-driven retry across VLM and embedding (#1049)
* feat(retry): add unified transient retry module (#922)

Implements `openviking/models/retry.py` with `is_transient_error`,
`transient_retry`, and `transient_retry_async` — a single config-driven
retry layer replacing scattered per-backend implementations.
Adds 50 unit tests covering classification, backoff, jitter, exhaustion,
and custom predicates.

* feat(vlm): integrate unified retry into VLM backends (#922)

- VLMBase: change max_retries default 2→3, remove max_retries param from
  get_completion_async abstract signature
- OpenAI backend: wrap all 4 methods with transient_retry/transient_retry_async,
  disable SDK retry (max_retries=0 in client constructors), remove manual
  for-loop retry
- VolcEngine backend: same pattern — transient_retry for all methods,
  remove manual for-loop retry
- LiteLLM backend: same pattern — transient_retry for all methods,
  remove manual for-loop retry

* refactor(vlm): migrate call sites to kwargs, remove max_retries params (#922)

- VLMConfig: default max_retries 2→3, remove max_retries from
  get_completion_async signature, switch all wrappers to kwargs
- StructuredVLM (llm.py): remove max_retries from complete_json_async and
  complete_model_async, switch all internal calls to kwargs
- memory_react.py: remove max_retries=self.vlm.max_retries (now handled
  internally by backend)
- Update test stubs to match new signatures (remove max_retries=0)

* test(vlm): add VLM retry integration tests (#922)

Tests cover OpenAI backend as representative:
- Completion retries on 429, does NOT retry on 401
- Vision completion now retries (was zero before)
- Config max_retries is used (default=3)
- max_retries removed from get_completion_async signature (all backends)
- OpenAI SDK retry disabled (max_retries=0 in client constructors)

* feat(embedding): добавить max_retries в EmbeddingConfig и EmbedderBase

- EmbeddingConfig: новое поле max_retries (default=3) для конфигурации retry
- EmbeddingConfig._create_embedder(): инжектирует max_retries в params["config"]
- EmbedderBase.__init__(): извлекает max_retries из config dict

* feat(embedding): перевести все embedding провайдеры на transient_retry

- OpenAI: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- Volcengine: заменить exponential_backoff_retry на transient_retry, убрать is_429_error
- VikingDB: добавить transient_retry (ранее retry отсутствовал)
- Gemini: отключить SDK HttpRetryOptions (attempts=1), обернуть embed/embed_batch
- MiniMax: отключить urllib3 Retry (total=0), обернуть embed/embed_batch
- Jina: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- Voyage: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- LiteLLM: обернуть litellm.embedding() вызовы

Все провайдеры теперь используют единый transient_retry с is_transient_error
для классификации ошибок. Wrapper размещён ВНУТРИ метода вокруг raw API call,
ДО try/except который конвертирует в RuntimeError.

* test(embedding): тесты retry для embedding провайдеров и backward compat

- test_embedding_retry_integration: OpenAI и VikingDB retry на transient/permanent ошибки
- test_retry_config: VLMConfig и EmbeddingConfig max_retries поля и defaults
- test_backward_compat: exponential_backoff_retry importable, signature unchanged, time-based

* style: ruff format для всех изменённых файлов

* style: fix ruff lint errors (import sorting, unused imports)

* style: format test_embedding_retry_integration.py

* style: remove unused transient_retry import from volcengine_vlm

After upstream refactored sync methods to use run_async(),
only transient_retry_async is needed in volcengine_vlm.py.

* style: fix ruff lint errors in volcengine_vlm.py

- Rename unused loop vars i, item → _i, _item (B007)
- Suppress unused has_images assignment (F841) — upstream code, kept for clarity

* ci: trigger re-run for flaky Windows API integration test

test_fs_tree intermittently returns 500 on windows-latest when
HAS_SECRETS=false — the resource processing pipeline retries
failed embeddings (401 auth errors) in the background, causing
server load that affects the fs/tree endpoint.
2026-03-31 13:47:46 +08:00
MaojiaShengandopenviking ce998873f9 lisence: change the main lisence to AGPL-3.0 (#1085)
* lisence: change the main lisence from Apache-2.0 to AGPL-v3

* lisence: change the main lisence from Apache-2.0 to AGPL-v3

* lisence: change the main lisence from Apache-2.0 to AGPL-v3

---------

Co-authored-by: openviking <openviking@example.com>
2026-03-30 14:37:42 +08:00
Dico Angelo 9bcdaf473e feat(embedder): add Cohere dense embedder (#941)
* feat(embedder): add Cohere dense embedder with embed-v4.0 support

Adds CohereDenseEmbedder using Cohere's Embed API v2.
- Supports embed-v4.0, embed-english-v3.0, embed-multilingual-v3.0
- Server-side dimension reduction for embed-v4.0 (256/512/1024/1536)
- Client-side truncation + renormalization fallback for v3 models
- Asymmetric search via input_type (search_query/search_document)
- Batch embedding with 96-item chunking (Cohere API limit)
- Full factory integration: provider validation, dimension resolution

none

* test(embedder): add unit tests for Cohere embedder

16 tests covering:
- Init validation (api_key required, defaults, model dimensions)
- Dimension handling (v4 server-side, v3 client-side truncation, invalid dims)
- Embedding calls (single, batch, query vs document input_type)
- output_dimension sent for embed-v4.0
- Error handling (API errors → RuntimeError)
- Resource cleanup (close)

none

* feat(rerank): add Cohere rerank-v3.5 support

Extends RerankConfig with provider field and api_key for Cohere.
Adds CohereRerankClient with same interface as VikingDB RerankClient.
HierarchicalRetriever auto-selects rerank backend based on provider.

Config example:
  "rerank": {"provider": "cohere", "api_key": "...", "threshold": 0.15}

Quality improvement: META tokenomics query 0.55 → 0.77 relevance score.

none

* test(rerank): add unit tests for Cohere reranker

9 tests covering:
- Rerank batch scoring with index-to-order mapping
- Empty input handling
- API error graceful fallback (returns None)
- Original order preservation from Cohere's sorted response
- Resource cleanup
- RerankConfig provider auto-detection (cohere/vikingdb/empty)

none

* perf(retrieve): increase GLOBAL_SEARCH_TOPK from 5 to 10

More vector candidates for reranker to evaluate = better precision.
With Cohere rerank-v3.5, 10 candidates gives the cross-encoder enough
material to find the best match without excessive latency.

none

* refactor: unify rerank dispatch — route all providers through RerankClient.from_config()

Cohere was special-cased in hierarchical_retriever.py while openai/litellm
went through the centralized RerankClient.from_config() dispatch. This commit
adds CohereRerankClient.from_config() and routes it through the same path.

Also fixes a bug where from_config() used config.provider directly instead
of _effective_provider(), which meant auto-detected providers (e.g. api_key
without explicit provider="cohere") would not dispatch correctly.

none
2026-03-30 14:15:08 +08:00
Deepak Dev Panwar df8ba97846 fix(embedder): add code model dimensions and actionable 422 error for Jina (#928)
Add jina-code-embeddings-1.5b and jina-code-embeddings-0.5b to
JINA_MODEL_DIMENSIONS so dimension validation works correctly.

Add _raise_task_error() helper that catches 422 validation errors
mentioning 'task' and produces a user-friendly error message guiding
users to set query_param and document_param in their embedding config.

Fixes #912.
2026-03-24 21:10:29 +08:00
Su Hangand3kyou1 79cc248140 fix: use code task defaults for jina (#914)
Co-authored-by: 3kyou1 <3kyou1@users.noreply.github.com>
2026-03-24 14:37:47 +08:00
Jannes StubbemannandClaude Opus 4.6 49a50fdcc4 feat(embedding): add litellm as embedding provider (#853)
* feat(embedding): add litellm as embedding provider

Adds LiteLLM as a new embedding provider, bringing embedding parity with
the VLM layer which already supports litellm. This enables users to route
embedding requests through OpenRouter, Ollama, vLLM, and any other
OpenAI-compatible endpoint via litellm's unified interface.

Closes #847

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review feedback on litellm embedding provider

1. Move os.environ.setdefault from module-level to __init__ to avoid
   mutating the process environment on import
2. Require dimension as mandatory — removes the probe API call during
   construction that caused surprise billable requests and silent fallbacks
3. Add None guard for LiteLLMDenseEmbedder in factory to give a clear
   error when litellm is not installed

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 17:52:00 +08:00
chethanuk 33f1a8f42b feat(gemini): add GeminiDenseEmbedder text embedding provider (#751)
* feat(gemini): add GeminiDenseEmbedder with is_query routing per #702 pattern

- GeminiDenseEmbedder: accept query_param/document_param, use is_query
  in embed() and embed_batch() to select task_type at call time
- EmbeddingConfig: add Gemini provider, factory, validation, dimension
- No get_query_embedder/get_document_embedder/_get_contextual_embedder
  (removed in #702; embed(is_query=True/False) is the pattern)
- Tests use embed(text, is_query=True/False) pattern throughout
- Rebased onto current upstream/main

* fix(gemini): remove task_type config field, fix conditional import for CI

- Remove task_type from EmbeddingModelConfig (query_param/document_param suffice)
- Wrap GeminiDenseEmbedder import in try/except (google-genai is optional)
- Update tests for removed field
2026-03-20 21:26:23 +08:00
zeng121 83a1cc44b0 feat: add Azure OpenAI support for embedding and VLM (#808)
* feat: add Azure OpenAI support for embedding and VLM

Add `azure` as a first-class provider for both embedding models and VLM,
using the official `openai.AzureOpenAI` / `openai.AsyncAzureOpenAI` clients.

Changes:
- embedder: OpenAIDenseEmbedder now accepts `provider` and `api_version`
  params; when provider is "azure", initializes AzureOpenAI client with
  azure_endpoint and api_version
- VLM: OpenAIVLM switches between OpenAI and AzureOpenAI clients based
  on the provider field; VLMFactory maps "azure" to OpenAIVLM
- config: EmbeddingModelConfig and VLMConfig gain `api_version` field;
  embedding_config adds ("azure", "dense") factory entry with validation
- registry: add "azure" to VALID_PROVIDERS
- docs: update README_CN.md with Azure provider table entry, config
  template fields, usage examples, and full config sample

Made-with: Cursor

* refactor: address PR review — extract helper, add validation, shared constant

- Extract `_build_openai_client_kwargs()` helper in openai_vlm.py to
  eliminate duplicated Azure/OpenAI client construction across
  get_client() and get_async_client()
- Add `api_base` validation for Azure provider in VLM client methods
  (was missing, unlike the embedder which already validated it)
- Extract `DEFAULT_AZURE_API_VERSION` constant in registry.py and
  reference it from both embedder and VLM, so future Azure API version
  bumps only require a single change
- Note: removing the hardcoded `self.provider = "openai"` override in
  OpenAIVLM.__init__ (done in the previous commit) is an intentional
  bug fix — it allows VLMBase.provider to correctly reflect the
  configured value (e.g. "azure"), which fixes token usage tracking

Made-with: Cursor
2026-03-20 16:39:35 +08:00
Xiaogang Zhouandxiaogang.zhou 1770a1792a feat(embedder): minimax embeding (#624)
* feat: enable minimax embedding and adapt to separate query/document parameter configuration

* feat: enable minimax embedding and adapt to separate query/document parameter configuration

---------

Co-authored-by: xiaogang.zhou <xiaogang.zhou@bytedance.com>
2026-03-20 00:36:51 +08:00
Qin Haojie d30d96cadf revert-embedding-chunking (#741) 2026-03-18 20:20:06 +08:00
Qin HaojieandClaude Opus 4.6 1823a7c4f7 feat(storage): add path locking and selective crash recovery for write operations (#431)
* feat(storage): add transaction support with journal, undo, and crash recovery

Implement a full transaction system for VikingFS storage operations including
write-ahead journal, path locking, undo/rollback, context manager API, and
crash recovery. Includes comprehensive tests and documentation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test(transaction): add e2e rollback tests for mv and multi-step operations

Add end-to-end tests covering rollback scenarios that were missing:
- mv rollback: file moved back to original location on failure
- mv commit: file persists at new location
- Multi-step rollback: mkdir + write + mkdir all reversed in order
- Partial step rollback: only completed entries are reversed
- Nested directory rollback: child removed before parent
- Best-effort rollback: single step failure does not block others

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(storage): add transaction support with path locking and journal

Implement transaction system for VikingFS with ACID-like guarantees:
- TransactionManager with configurable lock timeout and journal-based recovery
- PathLock supporting point, subtree, and mv lock modes
- Refactor VikingFS mv to use cp+rm to prevent lock files from being carried
- Fix stale lock detection returning false for missing lock files
- Update ragas eval to use LangchainLLMWrapper

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: tests

* fix(transaction): fix rollback and race condition bugs

- Reconstruct RequestContext from undo params for vectordb_delete/update_uri
  rollback (previously skipped silently due to missing ctx)
- Serialize ctx fields into undo params in rm/mv operations
- Fix Phase 1 undo path to target archive dir instead of session root
- Remove Phase 2 fs_write_new undo (overwrites are idempotent, checkpoint
  handles recovery)
- Add ancestor SUBTREE recheck after lock creation in acquire_subtree
- Move _collect_uris inside TransactionContext in rm/mv to close race window
- Log journal persistence failures instead of silently swallowing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(transaction): make TransactionManager required and rewrite tests with real backends

Remove all optional/fallback code paths where tx_manager could be None. get_transaction_manager()
now raises RuntimeError if not initialized. Fix undo rollback to reconstruct ctx for vectordb_upsert
and use correct agent_id default. Replace mock-based transaction tests with integration tests using
real AGFS and VectorDB backends.

* refactor(transaction): make rollback fully async and unify session commit path

- Convert execute_rollback/rollback_entry to async, removing sync run_async wrappers
- Unify Session.commit() to delegate to commit_async(), removing duplicate phase methods
- Fix SUBTREE lock to conflict with ancestor SUBTREE locks (was previously missing)
- Fix mv lock mode: directory moves now use SUBTREE on both source and destination
- Replace deprecated asyncio.get_event_loop() with get_running_loop()
- Remove max_parallel_locks config option
- Update docs (en/zh) and tests to match new async rollback signatures

* fix: tests

* refactor(transaction): simplify session commit and add redo-based crash recovery

Session commit no longer wraps archive phase in a transaction. Phase 2 uses
redo semantics so crashed memory-extraction can be replayed from archive.
PathLock stale-lock cleanup no longer redundantly re-checks timeout.
Semantic processor vectorization runs concurrently via asyncio.gather.

* fix: transaction

* fix: UserIdentifier

* refactor(transaction): replace undo-based transaction manager with lightweight lock + redo-log

Remove the heavyweight TransactionManager/Journal/UndoEntry system (~4000 lines) and
replace it with a simpler architecture: LockManager for path locking, LockContext as
the async context manager, LockHandle/LockOwner protocol, and a RedoLog for crash
recovery of session_memory operations. VikingFS rm/mv now use inline error handling
instead of rollback semantics. Updated docs, observers, and tests accordingly.

Co-Authored-By: Claude Opus 4.6

* fix(transaction): remove checkpoint dead code, fix TOCTOU race, clarify mv lock param

- Remove unused _write_checkpoint/_write_checkpoint_async/_read_checkpoint
  from Session (superseded by redo-log)
- Re-resolve URI inside lock in resource_processor Phase 3.5 to prevent
  concurrent add_resource calls from resolving to the same final_uri
- Rename acquire_mv dst_path to dst_parent_path with docstring to clarify
  that callers pass the destination parent directory

* fix: path

* fix: resource lock

* fix: test

* docs: update

* fix: tests

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-18 14:50:19 +08:00
Zayn Jarvis 3536531883 feat: add key=value parameter parsing to OpenAI embedder (#711)
* feat: add key=value parameter parsing to OpenAI embedder

Clean implementation on latest main (post PR #702 merge):
- Add _parse_param_string() method for key=value parsing
- Enhance _build_extra_body() to support multiple parameters
- Maintain backward compatibility with simple string format
- Update docstrings and examples

Usage:
- Simple: query_param='query' → {'input_type': 'query'}
- Enhanced: query_param='input_type=query,task=search' → {'input_type': 'query', 'task': 'search'}

Supports OpenAI-compatible servers with custom parameters while
maintaining clean integration with PR #702's is_query API.

* fix: lint issues
2026-03-18 11:51:13 +08:00
Xiaogang Zhouandxiaogang.zhou 9800b4001b feat(embedding): simplify the asymmetric embedder (#702)
Co-authored-by: xiaogang.zhou <xiaogang.zhou@bytedance.com>
2026-03-18 00:18:44 +08:00
Qin Haojie ae35f46c75 Revert "feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF) (#607)" (#703)
This reverts commit 95bd19797f.
2026-03-17 21:29:21 +08:00
chethanuk 95bd19797f feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF) (#607)
* feat(embedder): Gemini Embedding 2 multimodal support (text + image/video/audio/PDF)

Native text + multimodal (image, video, audio, PDF) embedding via `gemini-embedding-2-preview` (google-genai 1.67.0). Additive provider pattern — Volcengine remains the default; Gemini is opt-in via `provider: "gemini"` in `ov.conf`.
    - **Model**: `gemini-embedding-2-preview`
    - **Input**: text, image, video, audio, PDF (17 MIME types)
    - **Output dimension**: 128–3072 (default: **3072**, recommended: 768 / 1536 / 3072)
    - **Input token limit**: **8,192 tokens**
    - **Supported MIME types**: `image/jpeg`, `image/png`, `image/gif`, `image/webp`, `audio/mpeg`, `audio/mp3`, `audio/wav`, `audio/ogg`, `audio/flac`, `video/mp4`, `video/mpeg`, `video/mov`, `video/avi`, `video/webm`, `video/wmv`, `video/3gpp`, `application/pdf`
- Gemini Embedding 2 Multimodal Support: Introduced a new GeminiDenseEmbedder to support native text and multimodal (image, video, audio, PDF) embedding using the gemini-embedding-2-preview model. This is an opt-in provider via configuration.
- Extended Queue Pipeline for Multimodal Content: The EmbeddingMsg now carries media_uri and media_mime_type to facilitate multimodal content processing. The TextEmbeddingHandler.on_dequeue() method was updated to read raw bytes from viking_fs and call embed_multimodal() when applicable.
- End-to-End Configuration and Security: The EmbeddingConfig now registers the 'gemini' provider with a task_type field. A critical security validation was added to ensure media_uri matches context_data['uri'] before file reads, preventing forged queue messages from accessing arbitrary files. If validation fails or multimodal embedding fails, it falls back to text embedding.
- Multimodal Content Representation: A new ModalContent dataclass was introduced to represent media references, including MIME type, URI, and optional raw data, enabling the Vectorize object to encapsulate both text and media for embedding

* feat: Add asynchronous batch embedding with concurrency control to Gemini embedder.

* Reduce scope to use GeminiDenseEmbedder as only text embed
2026-03-17 21:15:12 +08:00
AstroHan de6cf2715f feat(embedder): support extra HTTP headers for OpenAI-compatible providers (#694)
Add `extra_headers` field to `EmbeddingModelConfig` and `OpenAIDenseEmbedder`,
enabling custom HTTP headers (e.g., `HTTP-Referer`, `X-Title` required by
OpenRouter) to be passed as `default_headers` to the OpenAI client.

Also fix a dead-code bug in `OpenAIDenseEmbedder.__init__` where an
unconditional `raise ValueError("api_key is required")` prevented the
more nuanced check that allows missing `api_key` when `api_base` is set
(e.g., local OpenAI-compatible servers).

Closes #675
2026-03-17 16:26:07 +08:00
kfiramar cdd0f795c3 feat(embedding): add Voyage dense embedding support (#635)
* feat(embedding): add voyage dense embedder

Add first-class Voyage dense embedding support with a dedicated embedder,
provider validation, model-aware default dimensions, and focused tests.

Keep the configuration surface intentionally narrow:
- use the existing dimension field and map it to Voyage's
  output_dimension request field
- do not expose Voyage-only output_dtype
- do not expose query/document mode until OpenViking has separate
  index/query embedder configuration

This keeps the PR aligned with the current OpenViking architecture,
which stores and retrieves dense float vectors through a single dense
embedder configuration.

Verification:
- .venv/bin/python -m pytest tests/unit/test_voyage_embedder.py tests/unit/test_embedding_config_voyage.py --noconftest -o addopts='' -q
- .venv/bin/python -m pytest tests/misc/test_config_validation.py -o addopts='' -q
- .venv/bin/ruff check openviking/models/embedder/voyage_embedders.py openviking_cli/utils/config/embedding_config.py tests/unit/test_embedding_config_voyage.py tests/unit/test_voyage_embedder.py

* refactor(embedder): inline Voyage extra body

Remove the trivial Voyage-specific helper and inline the output_dimension payload construction at the two call sites.

This keeps the request shape unchanged while making the embedder implementation more direct.
2026-03-16 21:46:37 +08:00
Wong ChihungandSisyphus 5260b817a0 feat(embedder): add non-symmetric embedding support for query/document (#608)
* feat(embedding): add nonsymmetric embedding support for query/document

- Add input_type parameter to OpenAIDenseEmbedder
- Add get_query_embedder() and get_document_embedder() to EmbeddingConfig
- Support different embedding strategies for queries and documents
- Improve error handling and default value processing
- Fix import ordering and remove unused imports

* fix(service): implement PR feedback for core.py

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* fix: address PR review comments - add telemetry tests and inline VikingDBManager init

* refactor(embedder): move context param handling into embedder classes

Per reviewer feedback, each embedder now decides its own context-specific
parameter mapping (input_type for OpenAI, task for Jina) based on a
'context' arg ('query'/'document'/None), instead of having EmbeddingConfig
pass pre-computed values.

- Unified config fields: replaced 4 separate fields (input_type_query,
  input_type_document, task_query, task_document) with 2 unified fields
  (query_param, document_param) that work for all providers
- OpenAI is symmetric by default; non-symmetric mode is activated
  implicitly when query_param or document_param is set
- Jina is non-symmetric by default; task is always sent unless context=None
- Updated tests to use new API

Note: Official OpenAI models are symmetric and do not support input_type.
Non-symmetric mode is only supported by OpenAI-compatible third-party
models (e.g., BGE-M3, Jina, Cohere).

---------

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-16 19:53:59 +08:00
jnMetaCode 6471868c90 Fix token underestimation for CJK text in _estimate_tokens fallback (#661)
When tiktoken is unavailable, the fallback `len(text) // 3` severely
underestimates tokens for CJK text (Chinese/Japanese/Korean characters
are ~1-2 tokens each, not 0.33). This causes text exceeding the 8192-token
API limit to bypass chunking, resulting in BadRequestError.

Use `max(len(text) // 3, len(text.encode("utf-8")) // 4)` instead, which
picks the more conservative estimate. For ASCII-heavy text the char-based
estimate still wins; for CJK text the byte-based estimate correctly
produces ~0.75 tokens per character.

Fixes #616, fixes #634

Signed-off-by: JiangNan <1394485448@qq.com>
2026-03-16 19:30:40 +08:00
ChenXiaofei e5abe88f04 feat(embedding): add Ollama provider support for local embedding (#644)
- Add 'ollama' as a supported embedding provider

- Ollama runs locally via OpenAI-compatible API, no API key required

- Allow OpenAI provider to work without api_key when api_base is set

  (supports local OpenAI-compatible servers like vLLM, LocalAI)

- Add configuration example and tests for Ollama provider

This enables fully local embedding deployment without cloud API keys.
2026-03-16 17:55:50 +08:00