Commit Graph
12 Commits
Author SHA1 Message Date
Yu Zhangandzhangyu.34 a57adb951f fix(rerank): support DashScope nested request/response envelope (#3463)
* fix(rerank): support DashScope nested request/response envelope

OpenAIRerankClient sent a flat request body ({"model", "query",
"documents"}) and parsed "results" at the top level of the response.
DashScope (qwen3-rerank) requires a nested envelope:

  Request:  {"model", "input": {"query", "documents"}, "parameters": ...}
  Response: {"output": {"results": [...]}, "request_id", "usage"}

This caused DashScope rerank to silently fail — the response had no
top-level "results" key, so the client returned None.

Changes:
- Add _is_dashscope() to detect DashScope endpoints by host marker.
- Add _build_request_body() that produces the nested envelope for
  DashScope and the flat body for standard OpenAI/Cohere services.
- Add _extract_results() that reads output.results for DashScope and
  top-level results for standard services.
- Accept both "relevance_score" (singular, DashScope) and
  "relevance_scores" (plural, some providers) in result items.
- Add 13 tests covering host detection, body construction, response
  parsing, end-to-end mocked flows for both providers, plural key
  handling, empty documents, and sparse results.

Fixes #3459

* fix(rerank): detect DashScope protocol by URL path, not hostname

Reviewer noted the previous hostname-based switch broke the documented
qwen3-rerank compatible-api endpoint (/compatible-api/v1/reranks), which
must use the flat OpenAI-style body and top-level results.

Switch to path-based detection: only /api/v1/services/rerank uses the
native nested input/output envelope; everything else (including the
DashScope compatible-api and generic OpenAI/Cohere gateways) keeps the
flat protocol. Rename _is_dashscope -> _uses_nested_envelope for clarity.

Add regression tests covering the compatible-api flat path and reconcile
the existing native-path fixtures to the nested envelope.

* docs(rerank): use qwen3-rerank for compatible-api example

The compatible-api/v1/reranks endpoint uses the flat OpenAI-compatible
protocol; qwen3-vl-rerank is a native-envelope model served at
/api/v1/services/rerank. Align the example model with the endpoint the
implementation selects by URL path.

---------

Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
2026-07-22 19:30:41 +08:00
19ca274a24 fix(retrieve): bound reranker input size (#3289)
* fix(retrieve): bound reranker input size

* fix(retrieve): make rerank input limit opt-in

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-16 17:19:38 +08:00
michaeltarletonandMichael Tarleton 3b7c61d266 fix(rerank): support Voyage via litellm (plain-string documents + dict results) (#3291)
LiteLLMRerankClient.rerank_batch wrapped each document as {"text": d} and read
result items via getattr(item, ...). That works for Cohere-style object results
but breaks Voyage through litellm: Voyage's rerank API rejects dict-wrapped
documents (400: 'documents' is not a valid string), and litellm returns Voyage
results as plain dicts, so getattr(item, "index") misses and rerank silently
falls back to a no-op.

- Pass documents as plain strings (litellm.rerank expects List[str]).
- Add _result_field() to read index/relevance_score from dict- or object-shaped
  result items.

Verified end-to-end against voyage/rerank-2.5 (scores now applied). Adds
tests/unit/models/rerank/test_litellm_rerank.py covering the plain-string
documents contract and both dict- and object-shaped results.

Co-authored-by: Michael Tarleton <mtarleton@istation.com>
2026-07-16 15:34:55 +08:00
huangruitengandhuangruiteng 2f5b2e27e1 fix(rerank): accept sparse indexed results (#3121)
* fix(rerank): accept sparse indexed results

* fix(rerank): warn on sparse provider results

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-11 10:00:56 +08:00
Dico Angeloandqin-ctx 9ec15c8e07 feat(rerank): add configurable HTTP timeout for OpenAI-compatible client (#2784)
* feat(rerank): add configurable HTTP timeout for OpenAI-compatible client

OpenAIRerankClient hardcoded a 30s HTTP timeout, which is insufficient for
local LLM servers (e.g. llama.cpp on ROCm) that incur model cold-start
latency on the first request after inactivity, causing ReadTimeout errors.

Add a `timeout` field to RerankConfig (default 30.0, backwards-compatible)
and thread it through OpenAIRerankClient.__init__, from_config, and the
requests.post call in rerank_batch. The timeout can now be set per-environment
in ov.conf, e.g. "timeout": 120.

Closes #2732

* docs: document rerank timeout config

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-06-23 18:23:23 +08:00
Jiahui Zhou 3bcefb298d fix: unify runtime loggers with openviking logger (#1981) 2026-05-12 11:26:30 +08:00
agent dd7a222c65 fix: rerank (#1933) 2026-05-09 14:51:56 +08:00
baojun-zhangandMaojiaSheng 17d2c5603e feat(observability): unify observability context && support otel && etc. (#1666)
* feat(observability): unify OTLP metrics export, log/trace context, and telemetry bridging
- - Add OTLP metrics http/grpc exporter that pushes MetricRegistry snapshots
- - Decouple telemetry response payload from telemetry collection; always finish() and bridge summary to metrics
- - Unify observability config under server.observability (metrics/traces/logs siblings); update ov.conf.example and docs (zh/en)
- - Improve log/trace correlation via structured context injection
- - Add/adjust tests for exporter lifecycle, config loader, metrics/telemetry runtime
- BREAKING CHANGE: remove legacy telemetry.* config path; use server.observability.*

* feat(observability): import Status/StatusCode for LogToSpanEventFilter

* feat(observability): fix check issue

* feat(observability): format code

---------

Co-authored-by: MaojiaSheng <shengmaojia@bytedance.com>
2026-04-24 21:32:25 +08:00
baojun-zhang 629fc241e4 feat(metric): add token-full-cycle metric (#1488)
* feat(metric): add token full-cycle metric && support token dashboard && optimize metric guide && deprecated 'server.telemetry.prometheus.enabled' configuration && change metric config to observability.metric

* feat(metric): add token full-cycle metric && support token dashboard && optimize metric guide && deprecated 'server.telemetry.prometheus.enabled' configuration && change metric config to observability.metric

* feat(metric): add debug log

* feat(metric): format code

* feat(metric): format code
2026-04-16 18:22:40 +08:00
baojun-zhang b441622ee6 feat(metric): add metric system (#1357)
* feat(metric): add metric system

* feat(metric): add metric system

* feat(metric): add metric system

* fix(metric): do not cancel refresh tasks on deadline; trust only authenticated account id; avoid per-scrape rerank clients

* doc(metric): add metric guide doc

* doc(metric): add metric guide doc

* doc(metric): add metric guide doc

* doc(metric): add metric guide doc

* feat(metric): fix bug & change account dimension switch to default true
2026-04-14 20:55:02 +08:00
caisiriusandClaude Opus 4.6 37e106a958 feat: rerank support extra headers (#1359)
* feat(rerank): add extra_headers field to RerankConfig

Add support for extra_headers configuration in RerankConfig, following the same pattern as EmbeddingConfig and VLMConfig. This allows users to specify custom HTTP headers for OpenAI-compatible rerank providers.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: add copyright header to test file

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(rerank): OpenAIRerankClient accepts and stores extra_headers

- Add extra_headers parameter to __init__ method
- Add extra_headers to from_config classmethod
- Default to empty dict when extra_headers is None
- Add comprehensive unit tests

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(rerank): merge extra_headers into API requests

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add `extra_headers` param in doc

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-13 14:24:52 +08:00
MaojiaShengandopenviking d739f742d7 fix: ov status shows embedding and rerank models usage (#1191)
* fix: add models observer info for embedder and rerank

* fix: make build deps

* fix: ov observer

* fix: ov observer

---------

Co-authored-by: openviking <openviking@example.com>
2026-04-03 09:00:38 +08:00