* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* feat(benchmark): add Claude Code LoCoMo evaluation
Add locomo benchmark for Claude Code, following the same pipeline as
openclaw/vikingbot: ingest → QA → judge → stat.
- ingest.py: send each session to `claude -p`, let auto-memory extract
- eval.py: run QA questions in isolated project dirs with auto-memory
- judge.py / stat_judge_result.py: reuse same grading logic
- run_full_eval.sh: orchestrate full pipeline
Supports custom API endpoint (ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN),
environment isolation via HOME override, and retry on timeout/error.
* feat(benchmark): add prompt-prefix support and ingest-phase statistics
- ingest.py: add --prompt-prefix flag to nudge auto-memory during ingest
- run_full_eval.sh: add --project-root, --home, --result-dir, --prompt-prefix
flags for flexible environment isolation (bare vs prompted experiments)
- stat_judge_result.py: report ingest-phase token/cost/latency alongside
QA statistics, with --ingest-csv auto-detection
* feat(benchmark): add OpenViking integration for LoCoMo eval
- eval.py: add --hooks-settings, --mcp-config, --ov-config for OV recall
during QA (sets OPENVIKING env vars, injects OV hint in prompt)
- run_full_eval.sh: wire --input, --hooks-settings, --mcp-config,
--ov-ingest-config, --ov-qa-config params through to eval/ingest
- Add OV config files: ov-hooks.json, ov-mcp.json, ov-ingest.conf, ov-qa.conf
- Add import_to_ov.py for direct LoCoMo → OpenViking data import via SDK
* feat(benchmark): update configuration files and enhance evaluation logging
* feat(benchmark): add configuration files for Claude Code LoCoMo evaluation and enhance eval script
* feat(benchmark): reorganize LoCoMo claude-code eval around 4 reproducible modes
Collapse months of rN iteration scripts into 4 entry points covering
the four published numbers:
- run_prompted.sh - vanilla CC auto-memory (no OV)
- run_sdk_iso.sh - openviking SDK pre-ingest, per-sample namespace
- run_sdk_noiso.sh - openviking SDK pre-ingest, shared namespace
- run_e2e.sh - claude -p stream-json multi-turn + auto-capture
All four go through the same eval.py/judge.py/stat_judge_result.py
pipeline. Hook + plugin lookup is parametrised via OPENVIKING_PLUGIN_DIR
so the configs are portable. Historical scripts/reports/intermediate
configs are moved under .tmp/legacy (gitignored) for reference.
* style(benchmark): apply ruff format + auto-fix F541 across LoCoMo claude-code suite
- ruff format on eval.py / ingest.py / judge.py / stat_judge_result.py
- drop extraneous f-prefixes flagged by F541 in stat_judge_result.py
and eval.py
* benchmark: add LoCoMo evaluation scripts for supermemory
* benchmark(locomo): improve supermemory ingest and eval robustness
- ingest.py: parallelize session upload/poll with ThreadPoolExecutor,
add sample-level concurrency, parse LoCoMo dates to ISO 8601,
simplify session content format
- supermemory/eval.py: force explicit supermemory_search in prompt to
work around first-turn autoRecall skip, pass question_time to gateway
- mem0/eval.py: increase gateway startup sleep from 3s to 5s
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(benchmark): remove dead code and fix potential IndexError in delete_container.py
- Remove unused variable `prefix_sanitized`
- Guard `k.split(":")[1]` access with length check to avoid IndexError
on malformed ingest record keys
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Implements a two-phase benchmark pipeline for evaluating mem0 on the
LoCoMo long-term conversation dataset (10 samples, 1540 non-adversarial QA pairs).
- ingest.py: imports LoCoMo conversation sessions into mem0, using
sample_id as the userId namespace. All messages use "user" role with
[SpeakerName]: prefix to preserve two-person dialogue structure.
Temporal context is added via a [System] prefix on each session.
- eval.py: sends QA questions to an OpenClaw agent backed by the
openclaw-mem0 plugin. Restarts the gateway per sample to switch the
active userId, verifies the correct user is loaded before running
questions, then parallelizes questions within each sample using unique
session keys. Parses session jsonl to collect accurate per-turn token
usage. Optionally judges answers with a Volcengine ARK LLM.
- delete_user.py: utility to clear mem0 memories for given user_ids.
- README.md: documents prerequisites, ingest/eval parameters, output
format, and per-sample run commands.