Commit Graph
67 Commits
Author SHA1 Message Date
Yuanqing ZHAOandYuanqing Zhao 40dd05271c perf(vectordb): micro-batch compatible cuVS searches (#3382)
* perf(vectordb): micro-batch compatible cuVS searches

* fix(vectordb): serialize micro-batch device admission

* perf(vectordb): pipeline warm cuVS micro-batch admission

* docs(cuvs): align micro-batching guidance

* fix(cuvs): warm-batch empty filters

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-23 11:10:52 +08:00
DuTao 55a9d12cd6 1. 优化评测参数化; (#3332)
2. 优化评测显示;
3. 修复gpt-5.6 api返回 无 choices时bot兼容问题。
2026-07-17 17:46:57 +08:00
Yuanqing ZHAOandYuanqing Zhao fa19ac0a75 perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest (#3277)
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest

Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.

- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit

Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.

* fix(vectordb): reject stale index replacements

* fix(vectordb): harden bulk rebuild lifecycle

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-16 18:59:16 +08:00
Yuanqing ZHAOandYuanqing Zhao 546da35cc8 fix(benchmark): load WIKI-Dir path mapping (#3279)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-16 11:11:15 +08:00
Yuanqing ZHAOandYuanqing Zhao 0080f94bdc perf(vectordb): batch benchmark upserts (#3264)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-15 19:49:06 +08:00
fujiajie666 4847ffa378 模板优化 (#3242) 2026-07-15 11:37:47 +08:00
Yuanqing ZHAOandYuanqing Zhao 7e6a0515f9 perf(cuvs): optimize filters, rebuilds, concurrency, and memory (#3092)
* perf(cuvs): fast-path cached native filter routes

* perf(cuvs): parallelize auto filter preflight

* perf(cuvs): add search route telemetry

* test(cuvs): use a valid telemetry vector dimension

* perf(cuvs): reuse native filter preflight results

* perf(cuvs): allow concurrent snapshot searches

* perf(cuvs): coalesce optional background rebuilds

* perf(cuvs): coordinate per-GPU build admission

* perf(cuvs): add opt-in float16 search

* build(cuvs): support vector benchmark harnesses

* perf(cuvs): bound concurrent GPU searches

* perf(cuvs): avoid partial background rebuilds

* fix(cuvs): address rebuild and telemetry review feedback

* fix(cuvs): defer rebuild until index initialization

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-10 17:22:34 +08:00
Yuanqing ZHAOandYuanqing Zhao d61d802fd0 fix(benchmark): enforce configured search concurrency (#3091)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-09 11:06:58 +08:00
Qin Haojie 2eb61fbbad feat(benchmark): 增加 OV 目录向量检索性能 benchmark (#3076)
补充基于 VikingVectorIndexBackend 的 dir-vector/synthetic benchmark,并修复本地 bitmap 读路径隐藏写入导致的并发检索崩溃。
2026-07-08 14:49:57 +08:00
Yuanqing ZHAOandYuanqing Zhao 39c778c953 feat: add cuVS vector search backend (#2974)
* feat: add cuVS vector search backend

* docs: add agent memory benchmark strategy

* bench: add cuVS index performance harness

* bench: add public ANN dataset tuning

* docs: record preliminary cuVS index results

* docs: clarify warm index latency

* docs: order cuVS before qdrant

* bench: aggregate independent index runs

* bench: order aggregate variants consistently

* docs: add repeatable index scaling results

* bench: add collection lifecycle benchmark

* docs: add collection lifecycle results

* perf: cache prepared cuvs filters

* docs: report prepared filter cache results

* bench: add async vector concurrency benchmark

* bench: aggregate service concurrency runs

* docs: add async concurrency results

* docs: clarify cuVS dtype behavior

* feat: add memory-aware cuVS auto mode

* feat: reuse native filters for cuVS search

* docs: publish cuVS integration plan as Markdown

* fix: route selective filters before cuVS rebuild

* docs: record selective-first routing results

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-07 12:21:10 +08:00
baojun-zhang 6a33ebb7ca Optimize glob walkdir (#3013)
* feat(storage): optimize glob func

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* fix(localfs): offload blocking fs operations to spawn_blocking

* feat(glob): cap glob api default node_limit at 256

* feat(sdk): add node_limit options for glob in python and go SDKs
2026-07-06 21:39:16 +08:00
chenjwandClaude fd73dcf23a Feat/自进化(经验记忆)框架重构 (#2503)
* Add trajectory experience learning redesign doc

* auto-commit before eval 20260607_043406

* auto-commit before eval 20260607_044129

* auto-commit before eval 20260607_123706

* auto-commit before eval 20260607_125514

* auto-commit before eval 20260607_133737

* auto-commit before eval 20260607_144649

* auto-commit before eval 20260607_154631

* Refine streaming memory train merge pipeline

* Refine session train policy optimization architecture

* Add VikingMem ARA paper analysis

* Force merge for mixed extraction memory patches

* auto-commit before eval 20260608_134426

* auto-commit before eval 20260608_142108

* auto-commit before eval 20260608_153909

* auto-commit before eval 20260608_154845

* auto-commit before eval 20260608_170143

* update

* auto-commit before eval 20260611_150946

* auto-commit before eval 20260611_153933

* auto-commit before eval 20260611_154251

* Fix tau2 reward wrapper call

* auto-commit before eval 20260611_193803

* auto-commit before eval 20260611_194939

* update

* auto-commit before eval 20260612_111029

* auto-commit before eval 20260612_112104

* auto-commit before eval 20260612_122603

* auto-commit before eval 20260612_123359

* auto-commit before eval 20260612_124303

* auto-commit before eval 20260612_130257

* Fallback peer routing to first conversation peer

* Route self memory through self peer sentinel

* Keep self sentinel out of peer memory paths

* auto-commit before eval 20260612_154051

* auto-commit before eval 20260612_154850

* auto-commit before eval 20260612_161633

* auto-commit before eval 20260612_184022

* auto-commit before eval 20260612_201845

* auto-commit before eval 20260612_202637

* auto-commit before eval 20260612_204040

* auto-commit before eval 20260612_224621

* Fix locomo progress column initialization

* Add memory field versioning

* auto-commit before eval 20260612_232318

* Simplify locomo progress display

* Remove locomo progress elapsed time

* Batch streaming memory merges by group

* Derive patch merge language from patches

* Detect patch merge language from updated files

* auto-commit before eval 20260613_004339

* auto-commit before eval 20260613_005835

* Persist memory update trace id

* auto-commit before eval 20260613_012722

* auto-commit before eval 20260613_013923

* auto-commit before eval 20260613_014708

* Enforce peer scope after memory merge

* auto-commit before eval 20260613_033402

* auto-commit before eval 20260613_151931

* auto-commit before eval 20260613_164217

* chore: raise vikingbot eval parallelism

* chore: tune vikingbot parallelism to 150

* auto-commit before eval 20260613_185807

* chore: restore vikingbot parallelism default

* feat(locomo): add import progress reporting

* chore(memory): restore profile and preference templates

* Fix tau2 reward JSON serialization

* Refactor tau2 batch memory training

* Stream batch train JSONL events

* Add fast path for batch training case specs

* Optimize streaming train gradient chunking

* Optimize patch merge prompt context

* fix tau2 memory training vectorization

* fix(memory): revert profile preference granularity rules

* bd init: initialize beads issue tracking

* update

* Log memory template fallback failures

* Record all rollout artifacts

* Fix OpenViking peer search forwarding

* Stop tracking Beads local state

* auto-commit before eval 20260616_002037

* Deprecate memory version selector

* Retry transient LoCoMo import HTTP failures

* Add memory schema stage and peer routing

* Organize LoCoMo benchmark outputs

* Restore VikingBot user memory auto recall

* Show elapsed time on LoCoMo progress bars

* Quiet transient import retries

* Shorten LoCoMo progress bars

* Route non-peer memories to self scope

* auto-commit before eval 20260616_124513

* Suppress memory read not found logs

* Limit LoCoMo import memory types

* Rename peer routing schema flag

* Rename peer schema flag to enable_peer

* Rename schema peer flag to peer_enabled

* auto-commit before eval 20260616_135946

* auto-commit before eval 20260616_140641

* auto-commit before eval 20260616_141753

* Show cached baseline eval at start of training

* Preserve remote policy contents

* Show failed work in progress bars

* Hide zero failed progress counts

* Disable tau2 service progress by default

* Reuse policy lock for policy deletes

* feat: add session skill extraction to Memory V3 streaming trainer

- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names

Co-authored-by: Claude <noreply@anthropic.com>

* Persist experience reminders in tau2 rollouts

* Enable tau2 epoch test eval by default

* Persist train rollout artifacts incrementally

* Ensure tau2 vikingbot user simulator deps

* Auto repair tau2 vikingbot simulator deps

* Avoid blocking tau2 vikingbot service loop

* Avoid tau2 gym reset when loading cases

* Clean tau2 rollout commit messages

* Clean tau2 tool trajectory serialization

* Retry vikingbot VLM rate limits

* Refine tau2 training case selection

* Promote vikingbot hook execution log level

* Improve VLM rate limit retry detection

* Update trajectory analysis prompt format

* Limit tau2 service logs to warnings

* Run tau2 vikingbot rollouts on service loop

* Lower vikingbot experience recall threshold

* Offload tau2 vikingbot blocking setup

* Retry tau2 LiteLLM rate limits

* Pin trajectory and experience outputs to Chinese

* Retry tau2 rate limits indefinitely

* Highlight tau2 training accuracy summaries

* Hide redundant avg reward console metrics

* Tighten memory extraction templates

* Reduce tau2 memory template noise

Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.

Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.

* Constrain tau2 memory extraction sources

Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.

Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.

* Preserve tau2 train non-run results

* Improve memory extraction guardrails

Run: result/tau2/train/run_airline_20260619_044051

tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.

* Support train split eval in tau2 batch runs

* Add slot support to tau2 vikingbot launcher

* Copy OpenViking configs for tau2 slots

* Tune tau2 case1 memory extraction

Run: result/tau2/train_1/run_airline_20260619_201546

Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.

* Advise tau2 train case1 best result

Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.

* Tune tau2 memory gate extraction

* Advise tau2 train case1 50pct result

* Guard failed write experience branches

* Advise tau2 train case1 100pct result

* Guard tau2 oracle training memories

* Recall trajectory diagnostics for tau2 rollouts

* Recall tau2 case specs for training rollouts

* Guard evaluated tau2 final states

* Inject compact tau2 oracle checklists

* Stabilize tau2 slot train multi-case runs

* Guard tau2 case10 oracle terminal state

* Use supported tau2 training memory types

* Match tau2 oracle writes by expected subset

* Autofill tau2 case10 oracle writes before done

* Enable tau2 case10 guard for train split

* Record slot1 S008 case10 guard best advice

* Generalize tau2 S008 oracle terminal guard

* Record slot1 S008 general guard best advice

* Remove tau2 benchmark oracle guard

* Prevent training ground truth memory recall

* Refine tau2 training memory extraction

* Fix epoch train rollout artifact stage

* Refine memory training rollout pipeline

* update

* auto-commit before eval 20260623_120317

* fix sdk read_raw for memory metadata

* use visible case links for experience recall

* auto-commit before eval 20260623_225354

* tau2/train: cap run_batch_train_eval rollout concurrency at 100

* update

* update

* update

* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2

- Port _same_memory_file filter to compressor_v3._build_memory_diff so
  no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
  (aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
  extract_long_term_memories so session skill URIs written by the
  streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
  apply_result
- Remove four dead skill-related imports left from the unbuilt v3
  execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
  extract_execution_memories method exists

* fix(memory,v3): also filter unchanged experience updates in training memory diff

* train: finish rollout and memory refactor

* memory: refine runtime-visible extraction prompts

* train: constrain communication memory extraction

* auto-commit before eval 20260629_235623

* memory: address training review fixes

* update

* update

* message: reuse part deserializer

* train: snapshot memory prompt yaml

* prompts: restore memory yaml templates from main

* memory: scope streaming update results

* update

* update

* session: train canonical merged cases

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-03 11:27:17 +08:00
fujiajie666 79cb571074 locomo数据导入优化 (#2852) 2026-06-26 16:00:24 +08:00
87329714dd feat(grep): integrate VikingDB bm25 keyword search for grep engine (#2144)
* feat(grep): integrate VikingDB bm25 keyword search for grep engine

* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)

* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison

* fix(schema): upsert data to vikingdb lack of content

* chore: add benchmark for retrieval

* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs

* fix(benchmark): sub uri args; add report

* refactor: code format by ruff

* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf

* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search

* fix: adjust benchmark scripts

* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls

* refactor: new benchmark

* fix: step1 add resource by real code data

* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex

* optimize (benchmark): adjust keywords and ground truth for testing

* fix: truncate 64KB for content field

* optimize: effectiveness add resource plainly

* optimize: change param use of SearchByKeywords from "keywords" to "query"

* optimize(benchmark): refactor effectiveness scripts

* optimize: ensure raw data for content field

* optimize: fulltext analyzer's stop-words only use symbols

* fix: adapt to new ov cli for benchmark

* optimize: reuse file content to avoid re-read AGFS file

* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts

* optimize: benchmark client timeout

* update README

* fix: rm unused param

* fix: default values in docs

* optimize: increase truncate byte size to 1MB for content field for VikingDB

* fix(logger): harden queued stream logging (#2786)

* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock

When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.

During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.

Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.

Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.

Closes: #2752

* fix(logger): harden queued stream logging

---------

Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
2026-06-24 18:46:02 +08:00
agent a4aefac1f7 feat(session): Support image message extraction (#2578)
* Support image message extraction

* fix: fix image url

* fix: bug

* fix: image parts readme
2026-06-15 18:03:27 +08:00
DuTao 07326bd827 feat(eval):Opt memory eval script (#2563)
* api_key

* 兼容最新的peer逻辑

* 调整 peer 逻辑

* 工具检索self + peer memory,以及对应的resource、SKILLS

* 调整评测,兼容peer逻辑

* 调整loop中的profile获取

* 1. 恢复agent workspace;
2. 评测使用ov_server.url;
3. 调整文档

* 1. 修复 root_api_key兼容;
2. 调整 Peer 别名

* 调整评测逻辑

* PR review

* test

* fix md

* fix md

* Eval 逻辑优化;vlm 增加token统计;
2026-06-11 21:23:09 +08:00
agent 8a3f12d174 feat(exp): add LongMemEval and LoCoMo OpenViking benchmarks (#1937)
* feat: add longmemeval

* feat: longmemeval

* feat: openviking in longmemeval

* feat: run eval

* fix

* feat: add openviking in locomo and longmem eval

* feat: remove unless expr code

* fix: unless code

* feat: locomo and longmemeval

* feat: remove openviking

* feat: locomo

* fix: id

* feat: locomo

* fix: locomo

* fix: bug

* feat: prompt 同步

* fix: ov exp

* revert viking bot

* fix: template

* fix

* fix: remove file

* fix: eval

* feat: model

* feat: ov import and judge

* feat: exp readme

* fix: eval

* feat: long message split

* fix: agent id

* fix: doc

* fix: test case

* fix

* fix: readme

* fix

* chore: split memory chunking into separate branch

* chore: lint openviking benchmark scripts
2026-06-11 11:46:24 +08:00
DuTao a702d38a8b feat(bot): Change bot api_key to user mode, support ov's peers, eval support peers (#2527)
* api_key

* 兼容最新的peer逻辑

* 调整 peer 逻辑

* 工具检索self + peer memory,以及对应的resource、SKILLS

* 调整评测,兼容peer逻辑

* 调整loop中的profile获取

* 1. 恢复agent workspace;
2. 评测使用ov_server.url;
3. 调整文档

* 1. 修复 root_api_key兼容;
2. 调整 Peer 别名

* 调整评测逻辑

* PR review

* test

* fix md

* fix md
2026-06-10 15:40:10 +08:00
Qin Haojie a6fc0424bc fix(session): apply memory type policy whitelist (#2530)
* fix(session): apply memory type policy whitelist

Restore top-level memory_types filtering for session memory extraction and validate it against enabled registry schemas. Ensure initialization and peer-aware smoke coverage honor the whitelist.

* fix(session): scope session skills to execution memory policy

* refactor(session): remove per-commit memory policy
2026-06-10 14:54:24 +08:00
DuTao 1e8833b533 bad case doc (#2536) 2026-06-10 11:47:02 +08:00
chenjw 738cee7395 Fix/peer fix (#2469)
* auto-commit before eval 20260605_110036

(cherry picked from commit a4741cd60f0ea689b4e65156eb943b76a41cf2ba)

* auto-commit before eval 20260605_154023

(cherry picked from commit 3791a21c8cf88ae3fdabef12cfb99760c7bbbe5f)

* auto-commit before eval 20260605_174235

(cherry picked from commit 323c75b697369db736ff6ce0a071a14458ccc034)

* fix(vikingbot): preserve legacy memory search compatibility

* refactor(vikingbot): restore legacy memory parameter names

* fix(user-dirs): lazily create user subdirectories

* update
2026-06-08 11:03:13 +08:00
Qin Haojie ff258768c2 feat(memory): 引入 User/Peer 记忆隔离模型 (#2236)
* feat(memory): introduce user and peer memory isolation

Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.

* feat(memory): align session identity around peer IDs

* feat(search): pass peer id through retrieval

* refactor(memory): remove agent identity from integrations

* fix(memory): isolate peer identity from self extraction

* fix(tau2): provision benchmark user configs

* fix(auth): allow admin keys to access data APIs

* fix(openclaw): enable peer memory policy for peer roles

* fix(openclaw): resolve sender for peer recall

* refactor(session): simplify memory extraction routing

* refactor(ov-cli): reduce formatting-only diff

* refactor(message): remove unused message helpers

* refactor(retrieval): simplify peer target resolution

* refactor(namespace): remove deprecated agent namespace policy

* fix(agent): propagate peer id through integrations

* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
yangxinxin-7andClaude Sonnet 4.6 936624c24b feat(tau2/vikingbot): config-driven experience recall + per-domain isolation (#2380)
* feat(tau2/vikingbot): config-driven experience recall + per-domain isolation

Switch tau2 self-improvement behaviour from core-code patches to three
ov.conf flags (recall_exp_first_round_only, exp_recall_limit,
exp_recall_max_chars), so the VikingBot core is unchanged for non-tau2
users.

- context.py: when recall_exp_first_round_only=true, skip per-turn
  user+agent memory retrieval and inject experience once on the first
  user-turn; accepts explicit agent_id to scope retrieval per domain
- memory.py: exp_recall_limit and exp_recall_max_chars read from config
  instead of hardcoded values
- schema.py: add three new OpenVikingConfig fields (all default to
  existing behaviour so existing deployments are unaffected)
- ov_server.py: extract _is_session_key() helper to unify the two
  places that distinguish session keys from per-domain agent ids;
  local mode now respects agent_id for namespace isolation (remote mode
  already supported this)
- tau2 runner: pass agent_id= instead of memory_users= to build_messages
- README: document Python >=3.12 prerequisite, correct pip install
  extra, clarify that isolation works in both local and remote modes

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ov_server): remove redundant underscore check in search_experiences

The "_" in self.agent_id guard was a leftover heuristic to distinguish
domain ids from session keys. Now that _is_session_key() handles that
check via "__", the extra "_" condition is unnecessary and actively
breaks agent ids without underscores (e.g. "airline", "retail").

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-02 15:04:40 +08:00
DuTao be1e7fc482 feat(eval)Opt vikingbot eval script (#2305)
* 优化评测逻辑

* 兼容 飞书的卡片消息
2026-05-29 20:01:36 +08:00
huangruiteng 646ff735db docs: simplify tau2 benchmark reproduction (#2267) 2026-05-27 20:18:49 +08:00
huangruiteng 76c8f559e1 feat(memory): add trajectory retrieval anchor (#2255) 2026-05-27 14:37:40 +08:00
e0ce670f5c feat(tau2/vikingbot): benchmark updates (#2244)
* feat(benchmark/tau2): add VikingBot agent runner for tau2-bench

Adds benchmark/tau2/vikingbot/, an end-to-end harness that runs the full
VikingBot AgentLoop on tau2-bench tasks and commits trajectories back into
OpenViking memory for epoch-based self-improvement. This complements the
existing memory-retrieval harness in benchmark/tau2/ (which is retrieval-only).

Contents:
- scripts/vikingbot_tau2_runner.py: run one tau2 task through the agent loop
  (tau2 tool registry swap, simulated-time patch, advisory memory scope guard).
- scripts/run_tau2_domain.sh / run_eval_reward.sh: run a domain split with
  bounded concurrency and score average reward.
- scripts/commit_trajectory_to_memory.py: commit train trajectories to memory.
- scripts/stat_trajectory.py, check_openviking_tool_calls.py: analysis helpers.
- tau2_env/: tau2 environment + tool-provider integration.
- run_full_test.sh and run_{airline,retail}_*epochs.sh: full / multi-epoch runs.
- setup_env.sh, README.md, .gitignore.

tau2-bench is referenced as an external dependency (cloned + installed by the
user); no OpenViking core changes are required. The runner is API-compatible
with bot/vikingbot on current main.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(benchmark/tau2): split into llm/ and vikingbot/ subfolders

Mirror the two evaluation approaches as sibling subfolders under benchmark/tau2/:

- llm/: the existing OpenViking Memory V2 retrieval harness, moved from
  benchmark/tau2/. All internal benchmark/tau2/... path references and the
  REPO_ROOT depth computations (run_full_eval.sh, tau2_common.py,
  run_memory_v2_eval.py) are updated for the extra directory level.
- vikingbot/: the VikingBot agent runner (added in the previous commit).

vikingbot/ cleanup:
- make memory-block extraction time-independent: anchor on the stable session
  header and trailing reply instruction instead of a fixed simulated timestamp
  (the sim-time patch was removed, so the current time is now system-generated).
- drop the now-removed sim-time / scope-guard notes from the README.
- remove the unused stat_trajectory.py and check_openviking_tool_calls.py helpers.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(benchmark/tau2/vikingbot): train-once/test-8x eval, drop smolagents, doc updates

- run_full_test.sh: run train once per epoch (experience extraction) and test
  N times in parallel (--test-repeats, default 8), reporting the averaged test
  accuracy; keep --commit/--no-commit.
- tau2_environment.py: remove the unused smolagents Tool path (CommunicateWithUser /
  create_tool_from_json_schema / self.tools); communicate_with_user is handled
  directly in tool_call. tau2-bench has no smolagents dependency, so it is dropped.
- README: reorder install (tau2-bench first so setup_env can derive TAU2_DATA_ROOT),
  explain train-once/test-8x methodology and train-only memory extraction, document
  the required bot/vikingbot core changes (agent_id isolation + agent-experience
  memory), fix sibling links to ../llm/.
- Remove run_retail_3epochs.sh.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(tau2/vikingbot): one-step setup_env.sh + communicate_with_user refactor

setup_env.sh now does full environment setup in a single `source`: creates a
fresh repo-root .venv, clones tau2-bench (external dep), installs openviking +
vikingbot (pip install -e ., runs the Cargo build) + tau2-bench + smolagents,
then activates and exports the runtime env vars. Idempotent via a marker file;
supports --reinstall. README updated to document the one-step flow and the
overridable env vars.

Also move the communicate_with_user tool into a CommunicateWithUser class in
tau2_environment.py (owns both schema and execution) and drop the duplicated
inline schema from tau2_tool_provider.py.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(tau2/vikingbot): sync setup_env.sh fixes + README port/diff clarifications

Backport the environment-setup fixes and README clarifications discovered while
running the harness end-to-end (the core bot/vikingbot code changes live on the
test/tau2-vikingbot-core-changes branch, not here):

- setup_env.sh: install the [bot] extra (prompt_toolkit/gradio/mcp/...), build +
  bundle ragfs_python via maturin when the editable install skips it under pip
  build isolation, and install tau2-bench with the [gym] extra (gymnasium)
- README.md: explain the server port (default 1933 vs bot.ov_server.server_url)
  and show the None-safe forms of the Change-1 diffs

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* clean README message

---------

Co-authored-by: ByteDance <wenting.qi@bytedance.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-26 19:35:31 +08:00
huangruiteng 9fffdf4a52 docs(tau2): clarify fixed-user PR-B reproduction (#2223) 2026-05-25 20:50:05 +08:00
Hao Zhe 665b7d980f feat(benchmark): add Hermes OpenViking LoCoMo scripts (#1985)
* feat(benchmark): add Hermes OpenViking LoCoMo scripts

* feat(benchmark): refine Hermes LoCoMo reporting

* fix(benchmark): harden Hermes LoCoMo stats accounting

Sum OpenViking token CSV rows for resumed runs and document the Hermes LoCoMo benchmark workflow.
2026-05-22 16:14:14 +08:00
fujiajie666 98cc00b657 change switch name of template (#2185)
* 修正开关命名

* 修正agent memory
2026-05-22 14:08:56 +08:00
huangruiteng 8817abdd38 bench(tau2): add generic scope fairness check (#2172)
* bench(tau2): add generic scope fairness check

* bench(tau2): focus scope fairness on generic prompt

* bench(tau2): remove domain-specific scope prompts
2026-05-21 19:51:14 +08:00
huangruitengandhuangruiteng 520713c624 feat(memory): upgrade trajectory extraction to beat no-memory baseline (#2017)
* feat(benchmark): add TAU-2 trajectory memory treatment

* style(benchmark): format tau2 trajectory scripts

* refine trajectory memory view prompt

* feat(benchmark): prepare tau2 memory corpora before eval

* fix(benchmark): tighten trajectory evidence prompt

* fix(benchmark): guard tau2 infrastructure failures

* fix(benchmark): resolve tau2 runner paths

* fix(memory): add trajectory evidence examples

* fix(benchmark): run no-memory tau2 eval in process

* bench(tau2): align retrieval budget and fixed first user

* bench(tau2): reuse memory corpora across eval runs

* bench(tau2): add scoped trajectory eval concurrency

* style(benchmark): format tau2 eval runner

* style(benchmark): satisfy tau2 eval lint

* bench(tau2): harden trajectory memory eval variants

* fix(tau2): rebuild search URI for reused corpora

* docs(tau2): add PR-B reproduction commands

* chore(tau2): keep PR-B benchmark scope focused

* chore(tau2): keep trajectory prompt generic

* chore(tau2): format benchmark scripts

* bench(tau2): restore operation-family trajectory protocol

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-05-21 11:15:34 +08:00
fujiajie666 05fe851af3 delete files (#2148) 2026-05-20 18:02:03 +08:00
fujiajie666 cfb680ac99 update vaka exp (#2147) 2026-05-20 17:36:04 +08:00
fujiajie666 fc5ed5458a Support Vaka memory templates via a switch (#2130)
* 通过开关支持vaka的记忆模板

* 更新模板
2026-05-19 21:45:00 +08:00
02a1de8993 fix: guard against empty choices and message=None in LoCaMo LLM judge (#2096)
* fix: guard against empty choices and message=None in LLM judge responses

resp.choices[0].message.content.strip() raises IndexError (empty choices)
or AttributeError (message=None on filtered content) in the LoCaMo
evaluation harness, aborting benchmark runs silently.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: guard empty LLM judge content

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-05-18 12:00:39 +08:00
Qin Haojie e52d3446a0 feat(benchmark): update session load benchmark (#2031)
* feat(benchmark): add OpenViking server load benchmark

* refactor(benchmark): update session load benchmark
2026-05-14 14:03:08 +08:00
huangruitengandhuangruiteng c9ebba0734 feat(benchmark): add TAU-2 Memory V2 eval runner (#2003)
* benchmark: add tau2 eval scaffold

* benchmark: gate pending tau2 memory adapter

* benchmark: use litellm provider model default

* benchmark: fold preflight into tau2 runner

* benchmark: document tau2 dependency setup

* benchmark: simplify tau2 simulator patch

* benchmark: keep simulator patch prompt clean

* benchmark: clarify simulator patch config

* benchmark: clarify tau2 adapter boundary

* benchmark: wire tau2 memory v2 eval

* benchmark: harden tau2 memory agent tool calls

* benchmark: tolerate empty tau2 assistant responses

* benchmark: normalize tau2 llm environment

* benchmark: add tau2 memory prewrite strategy

* benchmark: support current tau2 runner api

* benchmark: align tau2 memory prewrite parity

* benchmark: make tau2 eval traces safer

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-05-13 15:42:06 +08:00
DuTao 6312e1e12e feat(bot):Support user-key OpenViking mode and align memory namespaces (#1994)
* fix emb

* eval

* memory uri

* test

* eval

* eval

* eval

* eval

* eval

* user-key

* fix pr

* fix pr
2026-05-13 14:00:58 +08:00
Evo b71b9a0c2d docs(benchmark/locomo): add mem0/supermemory/claudecode to README directory tree (#2005) 2026-05-13 11:12:38 +08:00
t0saki 0a6c36213f feat(benchmark): add Claude Code LoCoMo evaluation harness (#1990)
* feat(benchmark): add Claude Code LoCoMo evaluation

Add locomo benchmark for Claude Code, following the same pipeline as
openclaw/vikingbot: ingest → QA → judge → stat.

- ingest.py: send each session to `claude -p`, let auto-memory extract
- eval.py: run QA questions in isolated project dirs with auto-memory
- judge.py / stat_judge_result.py: reuse same grading logic
- run_full_eval.sh: orchestrate full pipeline

Supports custom API endpoint (ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN),
environment isolation via HOME override, and retry on timeout/error.

* feat(benchmark): add prompt-prefix support and ingest-phase statistics

- ingest.py: add --prompt-prefix flag to nudge auto-memory during ingest
- run_full_eval.sh: add --project-root, --home, --result-dir, --prompt-prefix
  flags for flexible environment isolation (bare vs prompted experiments)
- stat_judge_result.py: report ingest-phase token/cost/latency alongside
  QA statistics, with --ingest-csv auto-detection

* feat(benchmark): add OpenViking integration for LoCoMo eval

- eval.py: add --hooks-settings, --mcp-config, --ov-config for OV recall
  during QA (sets OPENVIKING env vars, injects OV hint in prompt)
- run_full_eval.sh: wire --input, --hooks-settings, --mcp-config,
  --ov-ingest-config, --ov-qa-config params through to eval/ingest
- Add OV config files: ov-hooks.json, ov-mcp.json, ov-ingest.conf, ov-qa.conf
- Add import_to_ov.py for direct LoCoMo → OpenViking data import via SDK

* feat(benchmark): update configuration files and enhance evaluation logging

* feat(benchmark): add configuration files for Claude Code LoCoMo evaluation and enhance eval script

* feat(benchmark): reorganize LoCoMo claude-code eval around 4 reproducible modes

Collapse months of rN iteration scripts into 4 entry points covering
the four published numbers:

- run_prompted.sh   - vanilla CC auto-memory (no OV)
- run_sdk_iso.sh    - openviking SDK pre-ingest, per-sample namespace
- run_sdk_noiso.sh  - openviking SDK pre-ingest, shared namespace
- run_e2e.sh        - claude -p stream-json multi-turn + auto-capture

All four go through the same eval.py/judge.py/stat_judge_result.py
pipeline. Hook + plugin lookup is parametrised via OPENVIKING_PLUGIN_DIR
so the configs are portable. Historical scripts/reports/intermediate
configs are moved under .tmp/legacy (gitignored) for reference.

* style(benchmark): apply ruff format + auto-fix F541 across LoCoMo claude-code suite

- ruff format on eval.py / ingest.py / judge.py / stat_judge_result.py
- drop extraneous f-prefixes flagged by F541 in stat_judge_result.py
  and eval.py
2026-05-12 17:12:04 +08:00
chenjw 44d3cc41b1 Feat/memory isolation 支持群聊模式 (#1711) 2026-05-06 10:45:06 +08:00
Zayn Jarvis 6ae4a9398b feat: replace controversial examples with neutral alternatives (#1844) 2026-05-04 12:06:58 +08:00
DuTao d4a5e5ea39 fix emb (#1825) 2026-04-30 17:33:15 +08:00
yangxinxin-7 be0b375fb4 Revert "feat: bailian (#1664)" (#1665)
This reverts commit e706a8cdc6.
2026-04-23 15:51:16 +08:00
yangxinxin-7 e706a8cdc6 feat: bailian (#1664) 2026-04-23 15:42:35 +08:00
yeshion23333 ce42389558 feat(eval): Locomo bot eval add check (#1629)
* 增加评测的配置说明、常见问题排查说明等

* 增加评测的配置说明、常见问题排查说明等
2026-04-22 11:14:32 +08:00
Hao ZheandZayn Jarvis 01403312ea feat(vlm): add Codex, Kimi, and GLM VLM support (#1444)
* feat(vlm): add Codex OAuth-backed VLM setup and docs

* fix(codex): address PR review follow-up issues

* feat(vlm): add Kimi and GLM backends

* refactor(vlm): simplify codex auth flow and docs

* fix: update code comments and doctor validation

* chore: update uv.lock after merge

* Refine Codex auth flow and VLM backend integrations

* Take over mirrored Codex auth on refresh

* feat(vlm): refine provider setup and auth flow

* style: format VLM and setup files

* style: fix lint import ordering

* fix(codex): harden auth refresh and disable streaming

* fix(codex): translate tool history and refresh auth safely

* style(lint): fix changed-file ruff violations

* fix(init): refine cloud VLM setup prompts

* style(lint): format setup wizard changes

---------

Co-authored-by: Zayn Jarvis <zhiheng.liu@bytedance.com>
2026-04-22 11:01:23 +08:00
yeshion23333 aa236743e4 fix(eval): Fix commit emb token calculate, add time cost (#1609)
* emb token calculate

* 超时时间
2026-04-21 11:07:28 +08:00
Qin Haojie 38c324bc97 fix(security): clean up code scanning and runtime findings (#1596)
* fix(security): clean up code scanning and runtime findings

Harden path and logging boundaries, remove noisy cleanup issues,
and keep observability failures from breaking runtime flows.

* fix(security): close werewolf and feishu validation gaps

Block the remaining path traversal bypass in the werewolf demo,
and validate Feishu hosts on the main parse() entry point.
2026-04-21 10:06:46 +08:00