* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest
Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.
- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit
Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.
* fix(vectordb): reject stale index replacements
* fix(vectordb): harden bulk rebuild lifecycle
---------
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
* feat(storage): optimize glob func
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* fix(localfs): offload blocking fs operations to spawn_blocking
* feat(glob): cap glob api default node_limit at 256
* feat(sdk): add node_limit options for glob in python and go SDKs
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* feat(tau2/vikingbot): config-driven experience recall + per-domain isolation
Switch tau2 self-improvement behaviour from core-code patches to three
ov.conf flags (recall_exp_first_round_only, exp_recall_limit,
exp_recall_max_chars), so the VikingBot core is unchanged for non-tau2
users.
- context.py: when recall_exp_first_round_only=true, skip per-turn
user+agent memory retrieval and inject experience once on the first
user-turn; accepts explicit agent_id to scope retrieval per domain
- memory.py: exp_recall_limit and exp_recall_max_chars read from config
instead of hardcoded values
- schema.py: add three new OpenVikingConfig fields (all default to
existing behaviour so existing deployments are unaffected)
- ov_server.py: extract _is_session_key() helper to unify the two
places that distinguish session keys from per-domain agent ids;
local mode now respects agent_id for namespace isolation (remote mode
already supported this)
- tau2 runner: pass agent_id= instead of memory_users= to build_messages
- README: document Python >=3.12 prerequisite, correct pip install
extra, clarify that isolation works in both local and remote modes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ov_server): remove redundant underscore check in search_experiences
The "_" in self.agent_id guard was a leftover heuristic to distinguish
domain ids from session keys. Now that _is_session_key() handles that
check via "__", the extra "_" condition is unnecessary and actively
breaks agent ids without underscores (e.g. "airline", "retail").
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(benchmark/tau2): add VikingBot agent runner for tau2-bench
Adds benchmark/tau2/vikingbot/, an end-to-end harness that runs the full
VikingBot AgentLoop on tau2-bench tasks and commits trajectories back into
OpenViking memory for epoch-based self-improvement. This complements the
existing memory-retrieval harness in benchmark/tau2/ (which is retrieval-only).
Contents:
- scripts/vikingbot_tau2_runner.py: run one tau2 task through the agent loop
(tau2 tool registry swap, simulated-time patch, advisory memory scope guard).
- scripts/run_tau2_domain.sh / run_eval_reward.sh: run a domain split with
bounded concurrency and score average reward.
- scripts/commit_trajectory_to_memory.py: commit train trajectories to memory.
- scripts/stat_trajectory.py, check_openviking_tool_calls.py: analysis helpers.
- tau2_env/: tau2 environment + tool-provider integration.
- run_full_test.sh and run_{airline,retail}_*epochs.sh: full / multi-epoch runs.
- setup_env.sh, README.md, .gitignore.
tau2-bench is referenced as an external dependency (cloned + installed by the
user); no OpenViking core changes are required. The runner is API-compatible
with bot/vikingbot on current main.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* refactor(benchmark/tau2): split into llm/ and vikingbot/ subfolders
Mirror the two evaluation approaches as sibling subfolders under benchmark/tau2/:
- llm/: the existing OpenViking Memory V2 retrieval harness, moved from
benchmark/tau2/. All internal benchmark/tau2/... path references and the
REPO_ROOT depth computations (run_full_eval.sh, tau2_common.py,
run_memory_v2_eval.py) are updated for the extra directory level.
- vikingbot/: the VikingBot agent runner (added in the previous commit).
vikingbot/ cleanup:
- make memory-block extraction time-independent: anchor on the stable session
header and trailing reply instruction instead of a fixed simulated timestamp
(the sim-time patch was removed, so the current time is now system-generated).
- drop the now-removed sim-time / scope-guard notes from the README.
- remove the unused stat_trajectory.py and check_openviking_tool_calls.py helpers.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* refactor(benchmark/tau2/vikingbot): train-once/test-8x eval, drop smolagents, doc updates
- run_full_test.sh: run train once per epoch (experience extraction) and test
N times in parallel (--test-repeats, default 8), reporting the averaged test
accuracy; keep --commit/--no-commit.
- tau2_environment.py: remove the unused smolagents Tool path (CommunicateWithUser /
create_tool_from_json_schema / self.tools); communicate_with_user is handled
directly in tool_call. tau2-bench has no smolagents dependency, so it is dropped.
- README: reorder install (tau2-bench first so setup_env can derive TAU2_DATA_ROOT),
explain train-once/test-8x methodology and train-only memory extraction, document
the required bot/vikingbot core changes (agent_id isolation + agent-experience
memory), fix sibling links to ../llm/.
- Remove run_retail_3epochs.sh.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(tau2/vikingbot): one-step setup_env.sh + communicate_with_user refactor
setup_env.sh now does full environment setup in a single `source`: creates a
fresh repo-root .venv, clones tau2-bench (external dep), installs openviking +
vikingbot (pip install -e ., runs the Cargo build) + tau2-bench + smolagents,
then activates and exports the runtime env vars. Idempotent via a marker file;
supports --reinstall. README updated to document the one-step flow and the
overridable env vars.
Also move the communicate_with_user tool into a CommunicateWithUser class in
tau2_environment.py (owns both schema and execution) and drop the duplicated
inline schema from tau2_tool_provider.py.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs(tau2/vikingbot): sync setup_env.sh fixes + README port/diff clarifications
Backport the environment-setup fixes and README clarifications discovered while
running the harness end-to-end (the core bot/vikingbot code changes live on the
test/tau2-vikingbot-core-changes branch, not here):
- setup_env.sh: install the [bot] extra (prompt_toolkit/gradio/mcp/...), build +
bundle ragfs_python via maturin when the editable install skips it under pip
build isolation, and install tau2-bench with the [gym] extra (gymnasium)
- README.md: explain the server port (default 1933 vs bot.ov_server.server_url)
and show the None-safe forms of the Change-1 diffs
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* clean README message
---------
Co-authored-by: ByteDance <wenting.qi@bytedance.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* feat(benchmark): add Claude Code LoCoMo evaluation
Add locomo benchmark for Claude Code, following the same pipeline as
openclaw/vikingbot: ingest → QA → judge → stat.
- ingest.py: send each session to `claude -p`, let auto-memory extract
- eval.py: run QA questions in isolated project dirs with auto-memory
- judge.py / stat_judge_result.py: reuse same grading logic
- run_full_eval.sh: orchestrate full pipeline
Supports custom API endpoint (ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN),
environment isolation via HOME override, and retry on timeout/error.
* feat(benchmark): add prompt-prefix support and ingest-phase statistics
- ingest.py: add --prompt-prefix flag to nudge auto-memory during ingest
- run_full_eval.sh: add --project-root, --home, --result-dir, --prompt-prefix
flags for flexible environment isolation (bare vs prompted experiments)
- stat_judge_result.py: report ingest-phase token/cost/latency alongside
QA statistics, with --ingest-csv auto-detection
* feat(benchmark): add OpenViking integration for LoCoMo eval
- eval.py: add --hooks-settings, --mcp-config, --ov-config for OV recall
during QA (sets OPENVIKING env vars, injects OV hint in prompt)
- run_full_eval.sh: wire --input, --hooks-settings, --mcp-config,
--ov-ingest-config, --ov-qa-config params through to eval/ingest
- Add OV config files: ov-hooks.json, ov-mcp.json, ov-ingest.conf, ov-qa.conf
- Add import_to_ov.py for direct LoCoMo → OpenViking data import via SDK
* feat(benchmark): update configuration files and enhance evaluation logging
* feat(benchmark): add configuration files for Claude Code LoCoMo evaluation and enhance eval script
* feat(benchmark): reorganize LoCoMo claude-code eval around 4 reproducible modes
Collapse months of rN iteration scripts into 4 entry points covering
the four published numbers:
- run_prompted.sh - vanilla CC auto-memory (no OV)
- run_sdk_iso.sh - openviking SDK pre-ingest, per-sample namespace
- run_sdk_noiso.sh - openviking SDK pre-ingest, shared namespace
- run_e2e.sh - claude -p stream-json multi-turn + auto-capture
All four go through the same eval.py/judge.py/stat_judge_result.py
pipeline. Hook + plugin lookup is parametrised via OPENVIKING_PLUGIN_DIR
so the configs are portable. Historical scripts/reports/intermediate
configs are moved under .tmp/legacy (gitignored) for reference.
* style(benchmark): apply ruff format + auto-fix F541 across LoCoMo claude-code suite
- ruff format on eval.py / ingest.py / judge.py / stat_judge_result.py
- drop extraneous f-prefixes flagged by F541 in stat_judge_result.py
and eval.py
* fix(security): clean up code scanning and runtime findings
Harden path and logging boundaries, remove noisy cleanup issues,
and keep observability failures from breaking runtime flows.
* fix(security): close werewolf and feishu validation gaps
Block the remaining path traversal bypass in the werewolf demo,
and validate Feishu hosts on the main parse() entry point.