DuTao
55a9d12cd6
1. 优化评测参数化; ( #3332 )
...
2. 优化评测显示;
3. 修复gpt-5.6 api返回 无 choices时bot兼容问题。
2026-07-17 17:46:57 +08:00
fujiajie666
4847ffa378
模板优化 ( #3242 )
2026-07-15 11:37:47 +08:00
chenjw and Claude
fd73dcf23a
Feat/自进化(经验记忆)框架重构 ( #2503 )
...
* Add trajectory experience learning redesign doc
* auto-commit before eval 20260607_043406
* auto-commit before eval 20260607_044129
* auto-commit before eval 20260607_123706
* auto-commit before eval 20260607_125514
* auto-commit before eval 20260607_133737
* auto-commit before eval 20260607_144649
* auto-commit before eval 20260607_154631
* Refine streaming memory train merge pipeline
* Refine session train policy optimization architecture
* Add VikingMem ARA paper analysis
* Force merge for mixed extraction memory patches
* auto-commit before eval 20260608_134426
* auto-commit before eval 20260608_142108
* auto-commit before eval 20260608_153909
* auto-commit before eval 20260608_154845
* auto-commit before eval 20260608_170143
* update
* auto-commit before eval 20260611_150946
* auto-commit before eval 20260611_153933
* auto-commit before eval 20260611_154251
* Fix tau2 reward wrapper call
* auto-commit before eval 20260611_193803
* auto-commit before eval 20260611_194939
* update
* auto-commit before eval 20260612_111029
* auto-commit before eval 20260612_112104
* auto-commit before eval 20260612_122603
* auto-commit before eval 20260612_123359
* auto-commit before eval 20260612_124303
* auto-commit before eval 20260612_130257
* Fallback peer routing to first conversation peer
* Route self memory through self peer sentinel
* Keep self sentinel out of peer memory paths
* auto-commit before eval 20260612_154051
* auto-commit before eval 20260612_154850
* auto-commit before eval 20260612_161633
* auto-commit before eval 20260612_184022
* auto-commit before eval 20260612_201845
* auto-commit before eval 20260612_202637
* auto-commit before eval 20260612_204040
* auto-commit before eval 20260612_224621
* Fix locomo progress column initialization
* Add memory field versioning
* auto-commit before eval 20260612_232318
* Simplify locomo progress display
* Remove locomo progress elapsed time
* Batch streaming memory merges by group
* Derive patch merge language from patches
* Detect patch merge language from updated files
* auto-commit before eval 20260613_004339
* auto-commit before eval 20260613_005835
* Persist memory update trace id
* auto-commit before eval 20260613_012722
* auto-commit before eval 20260613_013923
* auto-commit before eval 20260613_014708
* Enforce peer scope after memory merge
* auto-commit before eval 20260613_033402
* auto-commit before eval 20260613_151931
* auto-commit before eval 20260613_164217
* chore: raise vikingbot eval parallelism
* chore: tune vikingbot parallelism to 150
* auto-commit before eval 20260613_185807
* chore: restore vikingbot parallelism default
* feat(locomo): add import progress reporting
* chore(memory): restore profile and preference templates
* Fix tau2 reward JSON serialization
* Refactor tau2 batch memory training
* Stream batch train JSONL events
* Add fast path for batch training case specs
* Optimize streaming train gradient chunking
* Optimize patch merge prompt context
* fix tau2 memory training vectorization
* fix(memory): revert profile preference granularity rules
* bd init: initialize beads issue tracking
* update
* Log memory template fallback failures
* Record all rollout artifacts
* Fix OpenViking peer search forwarding
* Stop tracking Beads local state
* auto-commit before eval 20260616_002037
* Deprecate memory version selector
* Retry transient LoCoMo import HTTP failures
* Add memory schema stage and peer routing
* Organize LoCoMo benchmark outputs
* Restore VikingBot user memory auto recall
* Show elapsed time on LoCoMo progress bars
* Quiet transient import retries
* Shorten LoCoMo progress bars
* Route non-peer memories to self scope
* auto-commit before eval 20260616_124513
* Suppress memory read not found logs
* Limit LoCoMo import memory types
* Rename peer routing schema flag
* Rename peer schema flag to enable_peer
* Rename schema peer flag to peer_enabled
* auto-commit before eval 20260616_135946
* auto-commit before eval 20260616_140641
* auto-commit before eval 20260616_141753
* Show cached baseline eval at start of training
* Preserve remote policy contents
* Show failed work in progress bars
* Hide zero failed progress counts
* Disable tau2 service progress by default
* Reuse policy lock for policy deletes
* feat: add session skill extraction to Memory V3 streaming trainer
- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names
Co-authored-by: Claude <noreply@anthropic.com >
* Persist experience reminders in tau2 rollouts
* Enable tau2 epoch test eval by default
* Persist train rollout artifacts incrementally
* Ensure tau2 vikingbot user simulator deps
* Auto repair tau2 vikingbot simulator deps
* Avoid blocking tau2 vikingbot service loop
* Avoid tau2 gym reset when loading cases
* Clean tau2 rollout commit messages
* Clean tau2 tool trajectory serialization
* Retry vikingbot VLM rate limits
* Refine tau2 training case selection
* Promote vikingbot hook execution log level
* Improve VLM rate limit retry detection
* Update trajectory analysis prompt format
* Limit tau2 service logs to warnings
* Run tau2 vikingbot rollouts on service loop
* Lower vikingbot experience recall threshold
* Offload tau2 vikingbot blocking setup
* Retry tau2 LiteLLM rate limits
* Pin trajectory and experience outputs to Chinese
* Retry tau2 rate limits indefinitely
* Highlight tau2 training accuracy summaries
* Hide redundant avg reward console metrics
* Tighten memory extraction templates
* Reduce tau2 memory template noise
Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.
Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.
* Constrain tau2 memory extraction sources
Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.
Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.
* Preserve tau2 train non-run results
* Improve memory extraction guardrails
Run: result/tau2/train/run_airline_20260619_044051
tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.
* Support train split eval in tau2 batch runs
* Add slot support to tau2 vikingbot launcher
* Copy OpenViking configs for tau2 slots
* Tune tau2 case1 memory extraction
Run: result/tau2/train_1/run_airline_20260619_201546
Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.
* Advise tau2 train case1 best result
Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.
* Tune tau2 memory gate extraction
* Advise tau2 train case1 50pct result
* Guard failed write experience branches
* Advise tau2 train case1 100pct result
* Guard tau2 oracle training memories
* Recall trajectory diagnostics for tau2 rollouts
* Recall tau2 case specs for training rollouts
* Guard evaluated tau2 final states
* Inject compact tau2 oracle checklists
* Stabilize tau2 slot train multi-case runs
* Guard tau2 case10 oracle terminal state
* Use supported tau2 training memory types
* Match tau2 oracle writes by expected subset
* Autofill tau2 case10 oracle writes before done
* Enable tau2 case10 guard for train split
* Record slot1 S008 case10 guard best advice
* Generalize tau2 S008 oracle terminal guard
* Record slot1 S008 general guard best advice
* Remove tau2 benchmark oracle guard
* Prevent training ground truth memory recall
* Refine tau2 training memory extraction
* Fix epoch train rollout artifact stage
* Refine memory training rollout pipeline
* update
* auto-commit before eval 20260623_120317
* fix sdk read_raw for memory metadata
* use visible case links for experience recall
* auto-commit before eval 20260623_225354
* tau2/train: cap run_batch_train_eval rollout concurrency at 100
* update
* update
* update
* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2
- Port _same_memory_file filter to compressor_v3._build_memory_diff so
no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
(aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
extract_long_term_memories so session skill URIs written by the
streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
apply_result
- Remove four dead skill-related imports left from the unbuilt v3
execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
extract_execution_memories method exists
* fix(memory,v3): also filter unchanged experience updates in training memory diff
* train: finish rollout and memory refactor
* memory: refine runtime-visible extraction prompts
* train: constrain communication memory extraction
* auto-commit before eval 20260629_235623
* memory: address training review fixes
* update
* update
* message: reuse part deserializer
* train: snapshot memory prompt yaml
* prompts: restore memory yaml templates from main
* memory: scope streaming update results
* update
* update
* session: train canonical merged cases
---------
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-03 11:27:17 +08:00
fujiajie666
79cb571074
locomo数据导入优化 ( #2852 )
2026-06-26 16:00:24 +08:00
87329714dd
feat(grep): integrate VikingDB bm25 keyword search for grep engine ( #2144 )
...
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786 )
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes : #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com >
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com >
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com >
2026-06-24 18:46:02 +08:00
DuTao
07326bd827
feat(eval):Opt memory eval script ( #2563 )
...
* api_key
* 兼容最新的peer逻辑
* 调整 peer 逻辑
* 工具检索self + peer memory,以及对应的resource、SKILLS
* 调整评测,兼容peer逻辑
* 调整loop中的profile获取
* 1. 恢复agent workspace;
2. 评测使用ov_server.url;
3. 调整文档
* 1. 修复 root_api_key兼容;
2. 调整 Peer 别名
* 调整评测逻辑
* PR review
* test
* fix md
* fix md
* Eval 逻辑优化;vlm 增加token统计;
2026-06-11 21:23:09 +08:00
DuTao
a702d38a8b
feat(bot): Change bot api_key to user mode, support ov's peers, eval support peers ( #2527 )
...
* api_key
* 兼容最新的peer逻辑
* 调整 peer 逻辑
* 工具检索self + peer memory,以及对应的resource、SKILLS
* 调整评测,兼容peer逻辑
* 调整loop中的profile获取
* 1. 恢复agent workspace;
2. 评测使用ov_server.url;
3. 调整文档
* 1. 修复 root_api_key兼容;
2. 调整 Peer 别名
* 调整评测逻辑
* PR review
* test
* fix md
* fix md
2026-06-10 15:40:10 +08:00
Qin Haojie
a6fc0424bc
fix(session): apply memory type policy whitelist ( #2530 )
...
* fix(session): apply memory type policy whitelist
Restore top-level memory_types filtering for session memory extraction and validate it against enabled registry schemas. Ensure initialization and peer-aware smoke coverage honor the whitelist.
* fix(session): scope session skills to execution memory policy
* refactor(session): remove per-commit memory policy
2026-06-10 14:54:24 +08:00
chenjw
738cee7395
Fix/peer fix ( #2469 )
...
* auto-commit before eval 20260605_110036
(cherry picked from commit a4741cd60f0ea689b4e65156eb943b76a41cf2ba)
* auto-commit before eval 20260605_154023
(cherry picked from commit 3791a21c8cf88ae3fdabef12cfb99760c7bbbe5f)
* auto-commit before eval 20260605_174235
(cherry picked from commit 323c75b697369db736ff6ce0a071a14458ccc034)
* fix(vikingbot): preserve legacy memory search compatibility
* refactor(vikingbot): restore legacy memory parameter names
* fix(user-dirs): lazily create user subdirectories
* update
2026-06-08 11:03:13 +08:00
Qin Haojie
ff258768c2
feat(memory): 引入 User/Peer 记忆隔离模型 ( #2236 )
...
* feat(memory): introduce user and peer memory isolation
Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.
* feat(memory): align session identity around peer IDs
* feat(search): pass peer id through retrieval
* refactor(memory): remove agent identity from integrations
* fix(memory): isolate peer identity from self extraction
* fix(tau2): provision benchmark user configs
* fix(auth): allow admin keys to access data APIs
* fix(openclaw): enable peer memory policy for peer roles
* fix(openclaw): resolve sender for peer recall
* refactor(session): simplify memory extraction routing
* refactor(ov-cli): reduce formatting-only diff
* refactor(message): remove unused message helpers
* refactor(retrieval): simplify peer target resolution
* refactor(namespace): remove deprecated agent namespace policy
* fix(agent): propagate peer id through integrations
* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
DuTao
be1e7fc482
feat(eval)Opt vikingbot eval script ( #2305 )
...
* 优化评测逻辑
* 兼容 飞书的卡片消息
2026-05-29 20:01:36 +08:00
DuTao
6312e1e12e
feat(bot):Support user-key OpenViking mode and align memory namespaces ( #1994 )
...
* fix emb
* eval
* memory uri
* test
* eval
* eval
* eval
* eval
* eval
* user-key
* fix pr
* fix pr
2026-05-13 14:00:58 +08:00
chenjw
44d3cc41b1
Feat/memory isolation 支持群聊模式 ( #1711 )
2026-05-06 10:45:06 +08:00
DuTao
d4a5e5ea39
fix emb ( #1825 )
2026-04-30 17:33:15 +08:00
yeshion23333
ce42389558
feat(eval): Locomo bot eval add check ( #1629 )
...
* 增加评测的配置说明、常见问题排查说明等
* 增加评测的配置说明、常见问题排查说明等
2026-04-22 11:14:32 +08:00
yeshion23333
117d1e2e95
feat(eval): add openclaw eval sh ( #1287 )
...
* 增加完整的一键评测脚本
* 增加完整的一键评测脚本
2026-04-08 00:47:04 +08:00
chenjw
7f05828f53
Feature/memory opt ( #1159 )
2026-04-06 15:50:18 +08:00
yeshion23333
3d2037aaea
fix(eval) Fix import async ( #1203 )
...
* import async
* import async
2026-04-03 15:54:52 +08:00
yeshion23333
2f4b1480de
feat(eval): add locomo eval scripts for openclaw and readme ( #1152 )
...
* Locomo eval
* openclaw
2026-04-01 17:57:09 +08:00