* feat: add audio and video understanding via VLM
* docs: design media resource guards
* fix: bound media staging concurrency
* fix: cap unknown-size media staging
* test: stage media in routing fake
* test: exercise media staging callbacks
* test: trim media understanding coverage
* chore: 清理实现计划文档
* fix: 修复多凭证切换问题
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
* fix(rerank): support DashScope nested request/response envelope
OpenAIRerankClient sent a flat request body ({"model", "query",
"documents"}) and parsed "results" at the top level of the response.
DashScope (qwen3-rerank) requires a nested envelope:
Request: {"model", "input": {"query", "documents"}, "parameters": ...}
Response: {"output": {"results": [...]}, "request_id", "usage"}
This caused DashScope rerank to silently fail — the response had no
top-level "results" key, so the client returned None.
Changes:
- Add _is_dashscope() to detect DashScope endpoints by host marker.
- Add _build_request_body() that produces the nested envelope for
DashScope and the flat body for standard OpenAI/Cohere services.
- Add _extract_results() that reads output.results for DashScope and
top-level results for standard services.
- Accept both "relevance_score" (singular, DashScope) and
"relevance_scores" (plural, some providers) in result items.
- Add 13 tests covering host detection, body construction, response
parsing, end-to-end mocked flows for both providers, plural key
handling, empty documents, and sparse results.
Fixes#3459
* fix(rerank): detect DashScope protocol by URL path, not hostname
Reviewer noted the previous hostname-based switch broke the documented
qwen3-rerank compatible-api endpoint (/compatible-api/v1/reranks), which
must use the flat OpenAI-style body and top-level results.
Switch to path-based detection: only /api/v1/services/rerank uses the
native nested input/output envelope; everything else (including the
DashScope compatible-api and generic OpenAI/Cohere gateways) keeps the
flat protocol. Rename _is_dashscope -> _uses_nested_envelope for clarity.
Add regression tests covering the compatible-api flat path and reconcile
the existing native-path fixtures to the nested envelope.
* docs(rerank): use qwen3-rerank for compatible-api example
The compatible-api/v1/reranks endpoint uses the flat OpenAI-compatible
protocol; qwen3-vl-rerank is a native-envelope model served at
/api/v1/services/rerank. Align the example model with the endpoint the
implementation selects by URL path.
---------
Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
* chore(format): align python and c++ file formatting
* chore: update urllib3 to 2.7.0 and clean test imports
1. bump urllib3 dependency from 2.6.3 to 2.7.0
2. remove unused pytest import and RoleScope import from test file
* style: format list comprehensions and lambda function for readability
Adjust the line breaks in the list comprehension in the VikingSearchTool class to follow standard Python formatting conventions, and rewrap the lambda assignment in the test case to improve code readability without changing functionality.
* style: fix line wrapping and remove extra blank line
- remove stray blank line in ov_server.py
- wrap long logger.info line in memory.py for better readability
* style: fix targeted ruff lint violations
* chore: clean up unused imports and reorder code
This commit removes unused imports, reorders import statements for better consistency,
and simplifies some test file imports. Changes include:
- Remove redundant blank lines and unused imports across multiple test files and core modules
- Reorder imports in openviking hooks module to follow standard layout
- Fix import ordering in memory isolation handler
- Simplify php parser type imports
- Move volcengine mock import to correct position in test file
* refactor(uri utils): remove extra blank lines in uri.py
clean up redundant whitespace to improve code readability
* feat(vlm): expose timeout as a config field and thread it through
PR #1208 added a timeout parameter to _build_openai_client_kwargs with a
60.0s default but left it unexposed:
- VLMConfig has no timeout field (extra='forbid' rejects user overrides).
- OpenAIVLM.get_client / get_async_client call the builder without
passing timeout, so 60.0s is always used.
- LiteLLM backend never set timeout on its acompletion / completion
calls either.
On slow-response endpoints (DashScope, local inference servers,
large-prompt overview generation), the 60s ceiling causes spurious
timeout-retry loops. Fix by:
- Adding timeout: float = 60.0 to VLMConfig with gt=0 validation.
- Reading timeout into VLMBase.__init__.
- Threading self.timeout into _build_openai_client_kwargs callers.
- Including timeout in the LiteLLM acompletion / completion kwargs.
* style(vlm): apply ruff format + narrow pytest.raises to ValidationError
Align docs, examples, the setup wizard, and telemetry tests with the
current Doubao multimodal embedding model so generated configs keep
referencing the supported default consistently.
* refactor(model): unify config-driven retry across VLM and embedding
Move retry behavior into shared model-call utilities and config defaults so VLM and embedding providers handle transient failures consistently.
Co-Authored-By: Claude Opus 4.6
* fix
* docs(config): document model retry settings
* feat(retry): add unified transient retry module (#922)
Implements `openviking/models/retry.py` with `is_transient_error`,
`transient_retry`, and `transient_retry_async` — a single config-driven
retry layer replacing scattered per-backend implementations.
Adds 50 unit tests covering classification, backoff, jitter, exhaustion,
and custom predicates.
* feat(vlm): integrate unified retry into VLM backends (#922)
- VLMBase: change max_retries default 2→3, remove max_retries param from
get_completion_async abstract signature
- OpenAI backend: wrap all 4 methods with transient_retry/transient_retry_async,
disable SDK retry (max_retries=0 in client constructors), remove manual
for-loop retry
- VolcEngine backend: same pattern — transient_retry for all methods,
remove manual for-loop retry
- LiteLLM backend: same pattern — transient_retry for all methods,
remove manual for-loop retry
* refactor(vlm): migrate call sites to kwargs, remove max_retries params (#922)
- VLMConfig: default max_retries 2→3, remove max_retries from
get_completion_async signature, switch all wrappers to kwargs
- StructuredVLM (llm.py): remove max_retries from complete_json_async and
complete_model_async, switch all internal calls to kwargs
- memory_react.py: remove max_retries=self.vlm.max_retries (now handled
internally by backend)
- Update test stubs to match new signatures (remove max_retries=0)
* test(vlm): add VLM retry integration tests (#922)
Tests cover OpenAI backend as representative:
- Completion retries on 429, does NOT retry on 401
- Vision completion now retries (was zero before)
- Config max_retries is used (default=3)
- max_retries removed from get_completion_async signature (all backends)
- OpenAI SDK retry disabled (max_retries=0 in client constructors)
* feat(embedding): добавить max_retries в EmbeddingConfig и EmbedderBase
- EmbeddingConfig: новое поле max_retries (default=3) для конфигурации retry
- EmbeddingConfig._create_embedder(): инжектирует max_retries в params["config"]
- EmbedderBase.__init__(): извлекает max_retries из config dict
* feat(embedding): перевести все embedding провайдеры на transient_retry
- OpenAI: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- Volcengine: заменить exponential_backoff_retry на transient_retry, убрать is_429_error
- VikingDB: добавить transient_retry (ранее retry отсутствовал)
- Gemini: отключить SDK HttpRetryOptions (attempts=1), обернуть embed/embed_batch
- MiniMax: отключить urllib3 Retry (total=0), обернуть embed/embed_batch
- Jina: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- Voyage: отключить SDK retry (max_retries=0), обернуть embed/embed_batch
- LiteLLM: обернуть litellm.embedding() вызовы
Все провайдеры теперь используют единый transient_retry с is_transient_error
для классификации ошибок. Wrapper размещён ВНУТРИ метода вокруг raw API call,
ДО try/except который конвертирует в RuntimeError.
* test(embedding): тесты retry для embedding провайдеров и backward compat
- test_embedding_retry_integration: OpenAI и VikingDB retry на transient/permanent ошибки
- test_retry_config: VLMConfig и EmbeddingConfig max_retries поля и defaults
- test_backward_compat: exponential_backoff_retry importable, signature unchanged, time-based
* style: ruff format для всех изменённых файлов
* style: fix ruff lint errors (import sorting, unused imports)
* style: format test_embedding_retry_integration.py
* style: remove unused transient_retry import from volcengine_vlm
After upstream refactored sync methods to use run_async(),
only transient_retry_async is needed in volcengine_vlm.py.
* style: fix ruff lint errors in volcengine_vlm.py
- Rename unused loop vars i, item → _i, _item (B007)
- Suppress unused has_images assignment (F841) — upstream code, kept for clarity
* ci: trigger re-run for flaky Windows API integration test
test_fs_tree intermittently returns 500 on windows-latest when
HAS_SECRETS=false — the resource processing pipeline retries
failed embeddings (401 auth errors) in the background, causing
server load that affects the fs/tree endpoint.
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
---------
Co-authored-by: openviking <openviking@example.com>
Reasoning-capable LLMs (MiniMax-M2.5, DeepSeek-R1, QwQ) served via vLLM
embed <think>...</think> blocks in message.content. These were stored
verbatim in file summaries and directory overviews, polluting semantic
search and wasting token budget during context loading.
Add _clean_response() in VLMBase that strips <think> tags, and call it
in all three backends (OpenAI, VolcEngine, LiteLLM) at every return point.
Closes#685
Co-Authored-By: Claude Opus 4.6