* feat(pdf): refactor MinerU parsing to the official file_parse API
* feat(pdf): remove mineru_api_key from configuration and examples
* feat(pdf): preflight MinerU /health during service initialization
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
Frontmatter was parsed into ParseResult.meta and removed from the body, but
that metadata is never persisted, so every ingested markdown file lost its
frontmatter with no way to read the fields back.
Parse frontmatter into meta unconditionally (it still drives doc_title) and
only remove it from the stored body when explicitly configured; that removal
is now off by default.
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix: remove heima partner, clean up auth docs, add web-studio unsupported auth banner
- Remove heima from partner list in README (en/zh/ja)
- Remove unsupported env var references (OPENVIKING_AUTH_MODE, OPENVIKING_USERNAME,
OPENVIKING_PASSWORD) from LDAP auth docs
- Remove temporary switch bash snippets from auth docs
- Fix ldap_password description
- Add web-studio unsupported-auth-mode banner for oidc/ldap servers
* fix: address OIDC/LDAP review comments on auth plugin design
Key changes driven by PR review:
- **Role mapping**: OIDC and LDAP external identities always resolve to
USER role. Removed map_role() calls and group_membership-based role
mapping. Admin access is gated by the root API key mechanism only.
- **LDAP credential extraction**: Removed query-parameter-based username/
password extraction (security concern — passwords in URLs can leak via
shell history, proxy logs, and monitoring). Clients must use Basic Auth
header or form data.
- **OIDC identifier sanitization**: Auth0 and other providers may include
characters like "|" in the `sub` claim. These are now replaced with "_"
to produce valid OpenViking user identifiers.
- **Dead code removal**: Removed _extract_groups, memberof_attribute,
require_root_api_key_for_admin, _initialize_api_key_manager, and
get_request_context_checks from both plugins since they are no longer
needed.
- **Docs**: Removed query-parameter curl example, memberof_attribute and
require_root_api_key_for_admin config references.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix(auth): bind lazy OIDC imports at module scope
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"
Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.
This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.
New assembly kernel under openviking/retrieve/context_assembler/:
- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
readable floor, then overview, then full for high-scoring entries. An
oversized tier falls back to the previous one instead of being truncated,
bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
Summary section, code files reuse code_outline signatures, long documents use
a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
directories carry no stored abstract; their full tier stays capped at
overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
every harness inherits cross-turn dedup; exclude_uris remains as the
stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
element per entry. Every tier carries its URI, so the model can always drill
down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
timeout fuses, and a failed rewrite still returns the unrewritten block.
Retrieval failures are counted into stats rather than silently yielding an
empty block.
Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.
* refactor(retrieval): give context tiers a per-category default
The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.
Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.
Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".
Assembly fixes found alongside:
- Removing the abstract cap exemption would turn an oversized abstract into
a bare URI, so it now falls back to overview first — for memory that is a
cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
`asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
such attribute; usage was structurally always null. It now reads the model
instance's tracker and reports only when the call count moved by exactly
one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
fail, and the file was never rewritten, so it could not heal. Records are
now coerced on read and dropped on the next write, along with records left
ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
`viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
resolved a different profile than `POST /recall`; an unknown `detail` value
raised `KeyError` through the whole call instead of degrading.
* feat(codex): inject profile context on session start
Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): raise rewrite timeout default to 30s
* docs(agents): document low-latency recall settings
* fix(codex): prefer luna as recall compressor fallback
* refactor(plugins): unify recall compression setting
* feat(plugins): enable recall compression by default
* docs(agents): use absolute links in image docs
* fix(retrieval): address context assembly review feedback
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* test: trim redundant context assembly coverage
* fix(retrieval): address second-round context assembly review
- Drop the backticked `/search` from the deprecated-recall row in both API
overviews. The reference checker scans the whole row after the method cell
for backticked paths, so it read the description as a route named
`POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
belongs to the Rust CLI, which writes `root_api_key`, `output`,
`echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
two Python readers had drifted into stricter subsets, so the shipped example
already failed to load in both. Adding the new `plugin` section to a working
ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
Retrieval validates query and image_url before searching, and the gather
fuse swallowed that rejection along with genuine scope failures, so a body
of `{"mode":"context"}` came back 200 with an empty block instead of the
documented parameter error. Runtime failures still degrade into
`stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
than the 30s fuse, so a rewrite that finished inside its own budget was
aborted client-side, discarding the whole response — including the
uncompressed block the server returns when a rewrite fails — and falling
back to `/recall`. The deadline is only extended when the body actually
requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
`plugin.recallContextTimeoutMs` pins it.
* chore(plugins): sync shared modules into the zcode snapshot
* fix(retrieval): align context quotas and plugin defaults
Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* fix(retrieval): preserve recall compatibility
Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(config): add enable_watch_scheduler toggle for read-only replicas
Add a top-level `enable_watch_scheduler` config flag (default true) and gate
`WatchScheduler.start()` on it in service initialization.
Read-only replicas that share a writer's data must not run the periodic
watch/refresh background loop: it would execute writes against shared storage
(vectordb / backups) and race the writer when persisting watch task state,
which has no cross-process lock. The scheduler instance is still created and
wired for on-demand read paths; only its background loop is skipped.
Defaults to true, so existing single-instance deployments are unaffected.
* fix
* fix
* feat: add audio and video understanding via VLM
* docs: design media resource guards
* fix: bound media staging concurrency
* fix: cap unknown-size media staging
* test: stage media in routing fake
* test: exercise media staging callbacks
* test: trim media understanding coverage
* chore: 清理实现计划文档
* fix: 修复多凭证切换问题
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
- add storage.agfs.pathlock.lock_timeout_secs
- use pathlock default timeout instead of hardcoded zero in wrapper
- map legacy storage.transaction.lock_timeout when new config is unset
- remote redolog by using persistent `session_commit` queue.
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(connector): support more git like platform
* feat(embedder): support extra_body passthrough in OpenAI embedder config
Adds optional `extra_body` (dict) to the OpenAI dense embedder, merged
into every embeddings.create call. Motivating use case: OpenRouter
provider routing ({"provider": {"sort": "latency"}}) — default routing
shows p90=35s/max=127s tail latency that kills interactive recall
(A/B: sorted routing is consistently sub-second).
Explicit query_param/document_param keys still take precedence on
conflict.
* feat(config): wire extra_body through embedding config layer
Add optional extra_body field to EmbeddingModelConfig and pass it to
OpenAIDenseEmbedder for the openai/azure providers, including the
multi-credential failover merge (parent-level model-behavior field).
* docs(config): document embedding extra_body with OpenRouter routing example
* docs(config): restore concrete host/cors_origins values in EN full schema
* feat(parse): add large image processing for image parser
- Add large_image_processor.py: detect large images (>10MB or >4096px),
create low-res previews, split into grid tiles, and generate grid
overlay images with tile labels
- Refactor ImageParser.parse() to integrate large image processing pipeline
- Enable SVG-to-PNG conversion in utils.py (cairosvg/wand)
- Rename ImageConfig.max_dimension to preview_max_dimension and add new
config fields: max_file_size_mb, max_tile_size_mb, max_tile_dimension_px,
tile_overlap_px, large_image_threshold_dimension
- Update ov.conf.example with new image config options
* fix(parse): correct tile dimension comment from 1024px to 2048px
* fix(parse): fix tile label path in grid overlay to include tiles/ directory
* fix(parse): register missing image extensions for ImageParser
TIFF, ICO, DIB, ICNS, SGI, JP2 were not in IMAGE_EXTENSIONS, causing
them to fallback to TextParser. All are supported by PIL.
* fix(parse): preserve PNG format for tiles instead of always converting to JPEG
* fix(parse): address review feedback for large image processing
- Wire config.image to ImageParser in ParserRegistry (was missing)
- Remove unnecessary preview creation for small images (broke LA mode PNG)
- Enforce max_tile_size_mb on tiles with quality reduction and resize fallback
- Remove 64-tile hard cap that conflicted with max_tile_dimension_px
- Add comment explaining why original file is not saved for large images
* refactor(parse): remove max_tile_size_mb as it is a soft suggestion
max_tile_size_mb was a soft constraint that was not enforced
consistently. Remove it from config, constants, and all enforcement
logic. Tile dimension (max_tile_dimension_px) remains the sole constraint.
* fix(parse): use CJK-capable font for grid overlay labels
The old font loading only tried macOS-specific paths and fell back to
PIL's default bitmap font, which cannot render CJK characters in
filenames. Add a cross-platform CJK font lookup that covers Linux
(Noto/Droid/WQY/DejaVu), macOS (PingFang), and Windows (MSYH/SimSun).
* fix(parse): convert non-VLM-supported image formats to PNG on save
Image formats like TIFF, ICO, DIB, ICNS, SGI, JP2 are not recognized
by VLM backends (OpenAI/LiteLLM/VolcEngine only support PNG/JPEG/GIF/
WebP/BMP) or by embedding_utils for image vectorization. When a file
with one of these extensions is parsed, convert it to PNG and use a
.png extension so that downstream pipelines see consistent data.
SVG files (already PNG-converted via cairosvg) also get the .png
extension for the same reason.
* fix(parse): import io for SVG conversion
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(connector): delegate add_resource imports to external Connector
Opt-in integration that routes add_resource data fetching and parsing
to external Connector service; the Connector stages source data and
calls back into OV through the standard add_resource pipeline.
- add ConnectorClient wrapping the control plane's inner doc/add and
task/info endpoints
- add [connector] config section: enable, connector/tracker endpoint
URLs, timeout_seconds, poll_interval_ms, allowed_add_types
- route add_resource via Connector when enabled and args.add_type is
in allowed_add_types; otherwise fall back to the standard pipeline
with an info log
- track imports as connector_import TaskRecords and poll Connector
task status in the background until terminal state or timeout
* fix(connector): delegate add_resource imports to external Connector
* feat(feishu): persist imported document images
Download Feishu image tokens into local import temp trees so Markdown image references can be ingested alongside the document.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(feishu): harden inline-image download (async, extension, user token)
Address review feedback on the inline-image import path:
- Run the synchronous lark-oapi media download via asyncio.to_thread so a
slow Feishu request no longer blocks unrelated async work on the event loop,
matching the existing _fetch_document() pattern.
- Infer the image file extension from the downloaded bytes (byte-magic
sniffing) and fall back to the response Content-Type via the existing
mime_types.get_preferred_extension helper, instead of hardcoding .png. This
stops JPEG/WebP/GIF bytes from being mislabeled as PNG to downstream
consumers (e.g. the data:image/... URI built during multimodal vectorization).
- Advertise AccessTokenType.USER on the media download request when a user
access token is supplied, so lark-oapi actually injects it. Previously the
request only allowed TENANT, so user-token imports read the document body
but silently dropped every image.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: wugj <wugj@g-bits.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(ingest): replay local agent-harness logs into OV sessions / 本地 agent harness 日志重放入库
Add openviking/ingest/: parse Claude Code / Codex / OpenCode / Hermes / OpenClaw conversation logs into normalized messages and replay them through OpenViking's existing session pipeline (create_session -> batch_add_messages -> commit -> async memory extraction), instead of a bespoke ETL.
Supports one-shot backfill ("存量") and cursor-driven incremental polling ("新增", WatchScheduler-style, no fs-event dependency), per-harness enable/mode/paths config, and meaningful peer_id on every turn (assistant = {harness}/{model}; user = git identity for single-user harnesses, original username for group-chat harnesses). Cursor IDE is a registered but deferred stub.
Read-position cursors persist under ~/.openviking/ingest/state.db for crash-safe, idempotent resume. New openviking-ingest CLI (backfill/watch/run/status/list-sources) and an "ingest" section on OpenVikingConfig. Verified end-to-end against a local server: 3-message fixture -> session commit -> 10 memories extracted -> idempotent re-run.
Inspired by / supersedes volcengine/OpenViking#2674.
Co-authored-by: baobaodae <2014596548@qq.com>
* docs(ingest): bilingual guide + ov.conf.example for openviking-ingest / 本地日志入库双语文档与配置示例
Add docs/{zh,en}/agent-integrations/09-log-ingestion.md (auto-registered in the VitePress sidebar) and an `ingest` section in examples/ov.conf.example (off by default).
* fix(ingest): address review — gating, crash-safe batch replay, commit recovery, single-instance lock / 修复评审问题
Fixes the merge-blockers from the adversarial review:
- master switch ingest.enabled now actually gates enabled_harnesses();
- idempotent per-batch append with a durable pending-intent reconciled against the server message count on restart (no duplicate imports after a mid-append crash);
- bounded reads (<=100 msgs/call) so huge sessions don't materialize at once;
- needs_commit flag + commit_if_needed so appended-but-uncommitted sessions still get extracted (commit even when no new source rows);
- poller keeps dirty sessions until a commit actually succeeds;
- OpenCode advances its SQLite cursor only past complete rows (late part text no longer skipped);
- single-instance file lock guards concurrent ingest processes;
- positive-value config validation (no poll busy-loop); malformed ov.conf surfaces instead of silently defaulting.
Adds 6 tests (config gating/validation, crash reconcile both ways, commit recovery).
* refactor(ingest): expose as 'openviking-server ingest' subcommand; English-only code/docs
- Route the ingest CLI through 'openviking-server ingest ...' (same dispatch as 'init'/'doctor') and drop the separate 'openviking-ingest' console_script.
- Remove mixed-in Chinese terms (存量/新增) from source docstrings, CLI help, and the English doc; the Chinese doc keeps them.
* style(ingest): ruff format + import sort (isort I)
Run ruff 0.15.16 (from the uv cache) with the repo config: fixes 5 I001 import-order errors in tests and reformats 9 files. 'ruff check' and 'ruff format --check' now pass on all added/edited files.
---------
Co-authored-by: baobaodae <2014596548@qq.com>
Add WebFeedAccessor (priority 60) that turns a single sitemap /
sitemapindex / RSS / Atom URL into ONE resource tree: it mirrors every
listed page into a temp directory and reuses the existing DirectoryParser
pipeline (the same "fetch-many -> dir -> tree" contract as GitAccessor).
A watch on the feed URL keeps the whole site refreshed (new pages added,
removed pages dropped on each rebuild).
- New openviking/parse/accessors/web_feed_accessor.py: WebFeedAccessor +
sitemap/feed extractors (nested sitemapindex recursion with depth cap,
RSS 2.0 / Atom via feedparser), bounded concurrent polite mirroring,
robots.txt, same-host / include / exclude / max_pages limits.
- args={"site": true} forces whole-site ingestion from a bare domain or
page by auto-discovering the sitemap/RSS (robots.txt, HTML
<link rel=alternate>, conventional paths); {"site": false} opts a
feed-looking URL back out to HTTPAccessor.
- Thread accessor-selection kwargs through can_handle; the registry
tolerates accessors whose can_handle lacks **kwargs (back-compatible).
- Single-page adds get a non-blocking "this site exposes a sitemap/RSS"
suggestion appended to the MCP add_resource response, gated to the
site root only; never auto-crawls.
- New WebFeedConfig (parsers.webfeed): max_pages, concurrency, politeness
delay, same_host_only, respect_robots, max_depth, suggest_feed.
- Dependencies: feedparser (robust RSS/Atom), defusedxml (XXE-safe XML).
- Docs: zh/en resources API, MCP/CLI/SDK help, ov.conf.example.
- Tests: 52 unit tests (fake httpx, no network).
* chore: clear unused files
* fix(tests): fix unit test
* refactor(auth): introduce plugin-based authentication architecture
Replace the monolithic `openviking/server/auth.py` with an extensible
plugin-based auth system. This refactor extracts the three built-in modes
(`dev`, `api_key`, `trusted`) into separate `AuthPlugin` implementations,
adds a registry for third-party plugins, and preserves all existing behavior
while enabling custom authentication backends (e.g. LDAP, OIDC, mTLS).
Key changes:
- **New public API**: `AuthPlugin` (ABC) and `register_auth_plugin` decorator.
- **New registry**: `AuthPluginRegistry` supports runtime registration.
- **Built-in plugins**: `DevAuthPlugin`, `ApiKeyAuthPlugin`, `TrustedAuthPlugin`.
- **Config change**: `auth_mode` widened from `Literal` to `str` for custom modes.
- **Validation delegated**: `validate_server_config()` now delegates to the active
plugin's `validate_config()`, preserving existing validation semantics.
- **Router compatibility**: All existing `require_*` decorators and `resolve_identity`
/ `get_request_context` dependencies remain unchanged. Routers import the same
symbols from `openviking.server.auth`.
- **Tests**: `conftest.py` manually wires the DevAuthPlugin in ASGI tests (lifespan
not triggered). `test_auth.py` expanded with plugin registration and validation tests.
- **Docs**: `04-authentication.md` (en/zh) updated with plugin registration examples.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): fix trusted mode test
* fix(tests): fix unit test
* fix(cli): remove unexisted transaction observer
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* docs: update skills definition
* fix(skills): now we allow viking://agent/skills again, and optimize CLI for skills
* docs(skills): use -p instead of --parent in agent skills examples
Align the `ov skills add` examples in the context-types and viking-uri
docs with the short flag `-p` introduced for `ov skills list/find/show`,
so all four user-facing examples consistently demonstrate the short form
when targeting `viking://agent/skills`.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): error check for api key
* fix(tests): unit test wait until resource not busy
* fix(tests): unit test wait until resource not busy
* fix(sdk): args form in skills find
* fix(skills): pass target uri in request body
---------
Co-authored-by: claude-sonnet-4-6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* feat(rerank): add configurable HTTP timeout for OpenAI-compatible client
OpenAIRerankClient hardcoded a 30s HTTP timeout, which is insufficient for
local LLM servers (e.g. llama.cpp on ROCm) that incur model cold-start
latency on the first request after inactivity, causing ReadTimeout errors.
Add a `timeout` field to RerankConfig (default 30.0, backwards-compatible)
and thread it through OpenAIRerankClient.__init__, from_config, and the
requests.post call in rerank_batch. The timeout can now be set per-environment
in ov.conf, e.g. "timeout": 120.
Closes#2732
* docs: document rerank timeout config
---------
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>