mirror of
https://github.com/volcengine/OpenViking.git
synced 2026-09-30 17:28:07 +08:00
* feat(grep): integrate VikingDB bm25 keyword search for grep engine * fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10) * fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison * fix(schema): upsert data to vikingdb lack of content * chore: add benchmark for retrieval * fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs * fix(benchmark): sub uri args; add report * refactor: code format by ruff * optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf * optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search * fix: adjust benchmark scripts * fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls * refactor: new benchmark * fix: step1 add resource by real code data * feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex * optimize (benchmark): adjust keywords and ground truth for testing * fix: truncate 64KB for content field * optimize: effectiveness add resource plainly * optimize: change param use of SearchByKeywords from "keywords" to "query" * optimize(benchmark): refactor effectiveness scripts * optimize: ensure raw data for content field * optimize: fulltext analyzer's stop-words only use symbols * fix: adapt to new ov cli for benchmark * optimize: reuse file content to avoid re-read AGFS file * optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts * optimize: benchmark client timeout * update README * fix: rm unused param * fix: default values in docs * optimize: increase truncate byte size to 1MB for content field for VikingDB * fix(logger): harden queued stream logging (#2786) * fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock When log.output='stdout' (default) and the server is managed by systemd, concurrent log writes can deadlock because logging.StreamHandler holds a thread lock across stream.flush() which blocks on systemd-piped file I/O. During session.commit() phase 2, multiple async coroutines (memory extraction, summarization) concurrently call logger.info()/warning() with large payloads. The first thread's flush() blocks on the pipe, while all subsequent threads block on handler.acquire() forever. This permanently silences the server log and prevents _write_done_file() from executing, leaving phase 2 hanging without .done. Fix: use QueueHandler + QueueListener from stdlib logging.handlers (Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock or I/O, returning immediately. QueueListener has a dedicated single thread as the sole consumer touching the real StreamHandler, making lock contention impossible. Changes in _create_log_handler(): stdout/stderr branches now create a shared QueueListener with unbounded queue, returning QueueHandler instances to callers. _build_standard_handler() delegates formatter and filter setup to the real handler in the listener thread. Closes: #2752 * fix(logger): harden queued stream logging --------- Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com> --------- Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com> Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>