Commit Graph
37 Commits
Author SHA1 Message Date
DuTao 55a9d12cd6 1. 优化评测参数化; (#3332)
2. 优化评测显示;
3. 修复gpt-5.6 api返回 无 choices时bot兼容问题。
2026-07-17 17:46:57 +08:00
fujiajie666 4847ffa378 模板优化 (#3242) 2026-07-15 11:37:47 +08:00
chenjwandClaude fd73dcf23a Feat/自进化(经验记忆)框架重构 (#2503)
* Add trajectory experience learning redesign doc

* auto-commit before eval 20260607_043406

* auto-commit before eval 20260607_044129

* auto-commit before eval 20260607_123706

* auto-commit before eval 20260607_125514

* auto-commit before eval 20260607_133737

* auto-commit before eval 20260607_144649

* auto-commit before eval 20260607_154631

* Refine streaming memory train merge pipeline

* Refine session train policy optimization architecture

* Add VikingMem ARA paper analysis

* Force merge for mixed extraction memory patches

* auto-commit before eval 20260608_134426

* auto-commit before eval 20260608_142108

* auto-commit before eval 20260608_153909

* auto-commit before eval 20260608_154845

* auto-commit before eval 20260608_170143

* update

* auto-commit before eval 20260611_150946

* auto-commit before eval 20260611_153933

* auto-commit before eval 20260611_154251

* Fix tau2 reward wrapper call

* auto-commit before eval 20260611_193803

* auto-commit before eval 20260611_194939

* update

* auto-commit before eval 20260612_111029

* auto-commit before eval 20260612_112104

* auto-commit before eval 20260612_122603

* auto-commit before eval 20260612_123359

* auto-commit before eval 20260612_124303

* auto-commit before eval 20260612_130257

* Fallback peer routing to first conversation peer

* Route self memory through self peer sentinel

* Keep self sentinel out of peer memory paths

* auto-commit before eval 20260612_154051

* auto-commit before eval 20260612_154850

* auto-commit before eval 20260612_161633

* auto-commit before eval 20260612_184022

* auto-commit before eval 20260612_201845

* auto-commit before eval 20260612_202637

* auto-commit before eval 20260612_204040

* auto-commit before eval 20260612_224621

* Fix locomo progress column initialization

* Add memory field versioning

* auto-commit before eval 20260612_232318

* Simplify locomo progress display

* Remove locomo progress elapsed time

* Batch streaming memory merges by group

* Derive patch merge language from patches

* Detect patch merge language from updated files

* auto-commit before eval 20260613_004339

* auto-commit before eval 20260613_005835

* Persist memory update trace id

* auto-commit before eval 20260613_012722

* auto-commit before eval 20260613_013923

* auto-commit before eval 20260613_014708

* Enforce peer scope after memory merge

* auto-commit before eval 20260613_033402

* auto-commit before eval 20260613_151931

* auto-commit before eval 20260613_164217

* chore: raise vikingbot eval parallelism

* chore: tune vikingbot parallelism to 150

* auto-commit before eval 20260613_185807

* chore: restore vikingbot parallelism default

* feat(locomo): add import progress reporting

* chore(memory): restore profile and preference templates

* Fix tau2 reward JSON serialization

* Refactor tau2 batch memory training

* Stream batch train JSONL events

* Add fast path for batch training case specs

* Optimize streaming train gradient chunking

* Optimize patch merge prompt context

* fix tau2 memory training vectorization

* fix(memory): revert profile preference granularity rules

* bd init: initialize beads issue tracking

* update

* Log memory template fallback failures

* Record all rollout artifacts

* Fix OpenViking peer search forwarding

* Stop tracking Beads local state

* auto-commit before eval 20260616_002037

* Deprecate memory version selector

* Retry transient LoCoMo import HTTP failures

* Add memory schema stage and peer routing

* Organize LoCoMo benchmark outputs

* Restore VikingBot user memory auto recall

* Show elapsed time on LoCoMo progress bars

* Quiet transient import retries

* Shorten LoCoMo progress bars

* Route non-peer memories to self scope

* auto-commit before eval 20260616_124513

* Suppress memory read not found logs

* Limit LoCoMo import memory types

* Rename peer routing schema flag

* Rename peer schema flag to enable_peer

* Rename schema peer flag to peer_enabled

* auto-commit before eval 20260616_135946

* auto-commit before eval 20260616_140641

* auto-commit before eval 20260616_141753

* Show cached baseline eval at start of training

* Preserve remote policy contents

* Show failed work in progress bars

* Hide zero failed progress counts

* Disable tau2 service progress by default

* Reuse policy lock for policy deletes

* feat: add session skill extraction to Memory V3 streaming trainer

- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names

Co-authored-by: Claude <noreply@anthropic.com>

* Persist experience reminders in tau2 rollouts

* Enable tau2 epoch test eval by default

* Persist train rollout artifacts incrementally

* Ensure tau2 vikingbot user simulator deps

* Auto repair tau2 vikingbot simulator deps

* Avoid blocking tau2 vikingbot service loop

* Avoid tau2 gym reset when loading cases

* Clean tau2 rollout commit messages

* Clean tau2 tool trajectory serialization

* Retry vikingbot VLM rate limits

* Refine tau2 training case selection

* Promote vikingbot hook execution log level

* Improve VLM rate limit retry detection

* Update trajectory analysis prompt format

* Limit tau2 service logs to warnings

* Run tau2 vikingbot rollouts on service loop

* Lower vikingbot experience recall threshold

* Offload tau2 vikingbot blocking setup

* Retry tau2 LiteLLM rate limits

* Pin trajectory and experience outputs to Chinese

* Retry tau2 rate limits indefinitely

* Highlight tau2 training accuracy summaries

* Hide redundant avg reward console metrics

* Tighten memory extraction templates

* Reduce tau2 memory template noise

Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.

Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.

* Constrain tau2 memory extraction sources

Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.

Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.

* Preserve tau2 train non-run results

* Improve memory extraction guardrails

Run: result/tau2/train/run_airline_20260619_044051

tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.

* Support train split eval in tau2 batch runs

* Add slot support to tau2 vikingbot launcher

* Copy OpenViking configs for tau2 slots

* Tune tau2 case1 memory extraction

Run: result/tau2/train_1/run_airline_20260619_201546

Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.

* Advise tau2 train case1 best result

Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.

* Tune tau2 memory gate extraction

* Advise tau2 train case1 50pct result

* Guard failed write experience branches

* Advise tau2 train case1 100pct result

* Guard tau2 oracle training memories

* Recall trajectory diagnostics for tau2 rollouts

* Recall tau2 case specs for training rollouts

* Guard evaluated tau2 final states

* Inject compact tau2 oracle checklists

* Stabilize tau2 slot train multi-case runs

* Guard tau2 case10 oracle terminal state

* Use supported tau2 training memory types

* Match tau2 oracle writes by expected subset

* Autofill tau2 case10 oracle writes before done

* Enable tau2 case10 guard for train split

* Record slot1 S008 case10 guard best advice

* Generalize tau2 S008 oracle terminal guard

* Record slot1 S008 general guard best advice

* Remove tau2 benchmark oracle guard

* Prevent training ground truth memory recall

* Refine tau2 training memory extraction

* Fix epoch train rollout artifact stage

* Refine memory training rollout pipeline

* update

* auto-commit before eval 20260623_120317

* fix sdk read_raw for memory metadata

* use visible case links for experience recall

* auto-commit before eval 20260623_225354

* tau2/train: cap run_batch_train_eval rollout concurrency at 100

* update

* update

* update

* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2

- Port _same_memory_file filter to compressor_v3._build_memory_diff so
  no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
  (aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
  extract_long_term_memories so session skill URIs written by the
  streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
  apply_result
- Remove four dead skill-related imports left from the unbuilt v3
  execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
  extract_execution_memories method exists

* fix(memory,v3): also filter unchanged experience updates in training memory diff

* train: finish rollout and memory refactor

* memory: refine runtime-visible extraction prompts

* train: constrain communication memory extraction

* auto-commit before eval 20260629_235623

* memory: address training review fixes

* update

* update

* message: reuse part deserializer

* train: snapshot memory prompt yaml

* prompts: restore memory yaml templates from main

* memory: scope streaming update results

* update

* update

* session: train canonical merged cases

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-03 11:27:17 +08:00
fujiajie666 79cb571074 locomo数据导入优化 (#2852) 2026-06-26 16:00:24 +08:00
87329714dd feat(grep): integrate VikingDB bm25 keyword search for grep engine (#2144)
* feat(grep): integrate VikingDB bm25 keyword search for grep engine

* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)

* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison

* fix(schema): upsert data to vikingdb lack of content

* chore: add benchmark for retrieval

* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs

* fix(benchmark): sub uri args; add report

* refactor: code format by ruff

* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf

* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search

* fix: adjust benchmark scripts

* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls

* refactor: new benchmark

* fix: step1 add resource by real code data

* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex

* optimize (benchmark): adjust keywords and ground truth for testing

* fix: truncate 64KB for content field

* optimize: effectiveness add resource plainly

* optimize: change param use of SearchByKeywords from "keywords" to "query"

* optimize(benchmark): refactor effectiveness scripts

* optimize: ensure raw data for content field

* optimize: fulltext analyzer's stop-words only use symbols

* fix: adapt to new ov cli for benchmark

* optimize: reuse file content to avoid re-read AGFS file

* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts

* optimize: benchmark client timeout

* update README

* fix: rm unused param

* fix: default values in docs

* optimize: increase truncate byte size to 1MB for content field for VikingDB

* fix(logger): harden queued stream logging (#2786)

* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock

When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.

During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.

Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.

Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.

Closes: #2752

* fix(logger): harden queued stream logging

---------

Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
2026-06-24 18:46:02 +08:00
agent a4aefac1f7 feat(session): Support image message extraction (#2578)
* Support image message extraction

* fix: fix image url

* fix: bug

* fix: image parts readme
2026-06-15 18:03:27 +08:00
DuTao 07326bd827 feat(eval):Opt memory eval script (#2563)
* api_key

* 兼容最新的peer逻辑

* 调整 peer 逻辑

* 工具检索self + peer memory,以及对应的resource、SKILLS

* 调整评测,兼容peer逻辑

* 调整loop中的profile获取

* 1. 恢复agent workspace;
2. 评测使用ov_server.url;
3. 调整文档

* 1. 修复 root_api_key兼容;
2. 调整 Peer 别名

* 调整评测逻辑

* PR review

* test

* fix md

* fix md

* Eval 逻辑优化;vlm 增加token统计;
2026-06-11 21:23:09 +08:00
agent 8a3f12d174 feat(exp): add LongMemEval and LoCoMo OpenViking benchmarks (#1937)
* feat: add longmemeval

* feat: longmemeval

* feat: openviking in longmemeval

* feat: run eval

* fix

* feat: add openviking in locomo and longmem eval

* feat: remove unless expr code

* fix: unless code

* feat: locomo and longmemeval

* feat: remove openviking

* feat: locomo

* fix: id

* feat: locomo

* fix: locomo

* fix: bug

* feat: prompt 同步

* fix: ov exp

* revert viking bot

* fix: template

* fix

* fix: remove file

* fix: eval

* feat: model

* feat: ov import and judge

* feat: exp readme

* fix: eval

* feat: long message split

* fix: agent id

* fix: doc

* fix: test case

* fix

* fix: readme

* fix

* chore: split memory chunking into separate branch

* chore: lint openviking benchmark scripts
2026-06-11 11:46:24 +08:00
DuTao a702d38a8b feat(bot): Change bot api_key to user mode, support ov's peers, eval support peers (#2527)
* api_key

* 兼容最新的peer逻辑

* 调整 peer 逻辑

* 工具检索self + peer memory,以及对应的resource、SKILLS

* 调整评测,兼容peer逻辑

* 调整loop中的profile获取

* 1. 恢复agent workspace;
2. 评测使用ov_server.url;
3. 调整文档

* 1. 修复 root_api_key兼容;
2. 调整 Peer 别名

* 调整评测逻辑

* PR review

* test

* fix md

* fix md
2026-06-10 15:40:10 +08:00
Qin Haojie a6fc0424bc fix(session): apply memory type policy whitelist (#2530)
* fix(session): apply memory type policy whitelist

Restore top-level memory_types filtering for session memory extraction and validate it against enabled registry schemas. Ensure initialization and peer-aware smoke coverage honor the whitelist.

* fix(session): scope session skills to execution memory policy

* refactor(session): remove per-commit memory policy
2026-06-10 14:54:24 +08:00
DuTao 1e8833b533 bad case doc (#2536) 2026-06-10 11:47:02 +08:00
chenjw 738cee7395 Fix/peer fix (#2469)
* auto-commit before eval 20260605_110036

(cherry picked from commit a4741cd60f0ea689b4e65156eb943b76a41cf2ba)

* auto-commit before eval 20260605_154023

(cherry picked from commit 3791a21c8cf88ae3fdabef12cfb99760c7bbbe5f)

* auto-commit before eval 20260605_174235

(cherry picked from commit 323c75b697369db736ff6ce0a071a14458ccc034)

* fix(vikingbot): preserve legacy memory search compatibility

* refactor(vikingbot): restore legacy memory parameter names

* fix(user-dirs): lazily create user subdirectories

* update
2026-06-08 11:03:13 +08:00
Qin Haojie ff258768c2 feat(memory): 引入 User/Peer 记忆隔离模型 (#2236)
* feat(memory): introduce user and peer memory isolation

Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.

* feat(memory): align session identity around peer IDs

* feat(search): pass peer id through retrieval

* refactor(memory): remove agent identity from integrations

* fix(memory): isolate peer identity from self extraction

* fix(tau2): provision benchmark user configs

* fix(auth): allow admin keys to access data APIs

* fix(openclaw): enable peer memory policy for peer roles

* fix(openclaw): resolve sender for peer recall

* refactor(session): simplify memory extraction routing

* refactor(ov-cli): reduce formatting-only diff

* refactor(message): remove unused message helpers

* refactor(retrieval): simplify peer target resolution

* refactor(namespace): remove deprecated agent namespace policy

* fix(agent): propagate peer id through integrations

* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
DuTao be1e7fc482 feat(eval)Opt vikingbot eval script (#2305)
* 优化评测逻辑

* 兼容 飞书的卡片消息
2026-05-29 20:01:36 +08:00
Hao Zhe 665b7d980f feat(benchmark): add Hermes OpenViking LoCoMo scripts (#1985)
* feat(benchmark): add Hermes OpenViking LoCoMo scripts

* feat(benchmark): refine Hermes LoCoMo reporting

* fix(benchmark): harden Hermes LoCoMo stats accounting

Sum OpenViking token CSV rows for resumed runs and document the Hermes LoCoMo benchmark workflow.
2026-05-22 16:14:14 +08:00
02a1de8993 fix: guard against empty choices and message=None in LoCaMo LLM judge (#2096)
* fix: guard against empty choices and message=None in LLM judge responses

resp.choices[0].message.content.strip() raises IndexError (empty choices)
or AttributeError (message=None on filtered content) in the LoCaMo
evaluation harness, aborting benchmark runs silently.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: guard empty LLM judge content

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-05-18 12:00:39 +08:00
DuTao 6312e1e12e feat(bot):Support user-key OpenViking mode and align memory namespaces (#1994)
* fix emb

* eval

* memory uri

* test

* eval

* eval

* eval

* eval

* eval

* user-key

* fix pr

* fix pr
2026-05-13 14:00:58 +08:00
Evo b71b9a0c2d docs(benchmark/locomo): add mem0/supermemory/claudecode to README directory tree (#2005) 2026-05-13 11:12:38 +08:00
t0saki 0a6c36213f feat(benchmark): add Claude Code LoCoMo evaluation harness (#1990)
* feat(benchmark): add Claude Code LoCoMo evaluation

Add locomo benchmark for Claude Code, following the same pipeline as
openclaw/vikingbot: ingest → QA → judge → stat.

- ingest.py: send each session to `claude -p`, let auto-memory extract
- eval.py: run QA questions in isolated project dirs with auto-memory
- judge.py / stat_judge_result.py: reuse same grading logic
- run_full_eval.sh: orchestrate full pipeline

Supports custom API endpoint (ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN),
environment isolation via HOME override, and retry on timeout/error.

* feat(benchmark): add prompt-prefix support and ingest-phase statistics

- ingest.py: add --prompt-prefix flag to nudge auto-memory during ingest
- run_full_eval.sh: add --project-root, --home, --result-dir, --prompt-prefix
  flags for flexible environment isolation (bare vs prompted experiments)
- stat_judge_result.py: report ingest-phase token/cost/latency alongside
  QA statistics, with --ingest-csv auto-detection

* feat(benchmark): add OpenViking integration for LoCoMo eval

- eval.py: add --hooks-settings, --mcp-config, --ov-config for OV recall
  during QA (sets OPENVIKING env vars, injects OV hint in prompt)
- run_full_eval.sh: wire --input, --hooks-settings, --mcp-config,
  --ov-ingest-config, --ov-qa-config params through to eval/ingest
- Add OV config files: ov-hooks.json, ov-mcp.json, ov-ingest.conf, ov-qa.conf
- Add import_to_ov.py for direct LoCoMo → OpenViking data import via SDK

* feat(benchmark): update configuration files and enhance evaluation logging

* feat(benchmark): add configuration files for Claude Code LoCoMo evaluation and enhance eval script

* feat(benchmark): reorganize LoCoMo claude-code eval around 4 reproducible modes

Collapse months of rN iteration scripts into 4 entry points covering
the four published numbers:

- run_prompted.sh   - vanilla CC auto-memory (no OV)
- run_sdk_iso.sh    - openviking SDK pre-ingest, per-sample namespace
- run_sdk_noiso.sh  - openviking SDK pre-ingest, shared namespace
- run_e2e.sh        - claude -p stream-json multi-turn + auto-capture

All four go through the same eval.py/judge.py/stat_judge_result.py
pipeline. Hook + plugin lookup is parametrised via OPENVIKING_PLUGIN_DIR
so the configs are portable. Historical scripts/reports/intermediate
configs are moved under .tmp/legacy (gitignored) for reference.

* style(benchmark): apply ruff format + auto-fix F541 across LoCoMo claude-code suite

- ruff format on eval.py / ingest.py / judge.py / stat_judge_result.py
- drop extraneous f-prefixes flagged by F541 in stat_judge_result.py
  and eval.py
2026-05-12 17:12:04 +08:00
chenjw 44d3cc41b1 Feat/memory isolation 支持群聊模式 (#1711) 2026-05-06 10:45:06 +08:00
DuTao d4a5e5ea39 fix emb (#1825) 2026-04-30 17:33:15 +08:00
yangxinxin-7 be0b375fb4 Revert "feat: bailian (#1664)" (#1665)
This reverts commit e706a8cdc6.
2026-04-23 15:51:16 +08:00
yangxinxin-7 e706a8cdc6 feat: bailian (#1664) 2026-04-23 15:42:35 +08:00
yeshion23333 ce42389558 feat(eval): Locomo bot eval add check (#1629)
* 增加评测的配置说明、常见问题排查说明等

* 增加评测的配置说明、常见问题排查说明等
2026-04-22 11:14:32 +08:00
Hao ZheandZayn Jarvis 01403312ea feat(vlm): add Codex, Kimi, and GLM VLM support (#1444)
* feat(vlm): add Codex OAuth-backed VLM setup and docs

* fix(codex): address PR review follow-up issues

* feat(vlm): add Kimi and GLM backends

* refactor(vlm): simplify codex auth flow and docs

* fix: update code comments and doctor validation

* chore: update uv.lock after merge

* Refine Codex auth flow and VLM backend integrations

* Take over mirrored Codex auth on refresh

* feat(vlm): refine provider setup and auth flow

* style: format VLM and setup files

* style: fix lint import ordering

* fix(codex): harden auth refresh and disable streaming

* fix(codex): translate tool history and refresh auth safely

* style(lint): fix changed-file ruff violations

* fix(init): refine cloud VLM setup prompts

* style(lint): format setup wizard changes

---------

Co-authored-by: Zayn Jarvis <zhiheng.liu@bytedance.com>
2026-04-22 11:01:23 +08:00
yeshion23333 aa236743e4 fix(eval): Fix commit emb token calculate, add time cost (#1609)
* emb token calculate

* 超时时间
2026-04-21 11:07:28 +08:00
yangxinxin-7andClaude Sonnet 4.6 26bbfd2c24 benchmark: add LoCoMo evaluation for Supermemory (#1401)
* benchmark: add LoCoMo evaluation scripts for supermemory

* benchmark(locomo): improve supermemory ingest and eval robustness

- ingest.py: parallelize session upload/poll with ThreadPoolExecutor,
  add sample-level concurrency, parse LoCoMo dates to ISO 8601,
  simplify session content format
- supermemory/eval.py: force explicit supermemory_search in prompt to
  work around first-turn autoRecall skip, pass question_time to gateway
- mem0/eval.py: increase gateway startup sleep from 3s to 5s

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(benchmark): remove dead code and fix potential IndexError in delete_container.py

- Remove unused variable `prefix_sanitized`
- Guard `k.split(":")[1]` access with length check to avoid IndexError
  on malformed ingest record keys

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-13 11:19:16 +08:00
yeshion23333 e3ac0ba575 feat(eval):Readme add qa (#1400)
* 增加关闭ov的配置

* 增加常见QA
2026-04-13 10:58:41 +08:00
yeshion23333 c6e8de9a5a 增加关闭ov的配置 (#1352) 2026-04-10 18:59:13 +08:00
yeshion23333 aaa7e2a859 fix(bot):Response language, Multi user memory commit (#1329)
* 修复语种问题

* 多user提交ov
2026-04-09 15:29:31 +08:00
yeshion23333 b05ef1d025 fix(eval): OpenClaw eval, import to ov use default user (#1305)
* 增加完整的一键评测脚本

* 增加完整的一键评测脚本

* 完善评测脚本

* 完善评测脚本

* 完善评测脚本
2026-04-08 21:05:12 +08:00
yangxinxin-7 b5da895a64 benchmark: add LoCoMo evaluation scripts for mem0 (#1290)
Implements a two-phase benchmark pipeline for evaluating mem0 on the
LoCoMo long-term conversation dataset (10 samples, 1540 non-adversarial QA pairs).

- ingest.py: imports LoCoMo conversation sessions into mem0, using
  sample_id as the userId namespace. All messages use "user" role with
  [SpeakerName]: prefix to preserve two-person dialogue structure.
  Temporal context is added via a [System] prefix on each session.

- eval.py: sends QA questions to an OpenClaw agent backed by the
  openclaw-mem0 plugin. Restarts the gateway per sample to switch the
  active userId, verifies the correct user is loaded before running
  questions, then parallelizes questions within each sample using unique
  session keys. Parses session jsonl to collect accurate per-turn token
  usage. Optionally judges answers with a Volcengine ARK LLM.

- delete_user.py: utility to clear mem0 memories for given user_ids.

- README.md: documents prerequisites, ingest/eval parameters, output
  format, and per-sample run commands.
2026-04-08 14:05:20 +08:00
yeshion23333 117d1e2e95 feat(eval): add openclaw eval sh (#1287)
* 增加完整的一键评测脚本

* 增加完整的一键评测脚本
2026-04-08 00:47:04 +08:00
chenjw 7f05828f53 Feature/memory opt (#1159) 2026-04-06 15:50:18 +08:00
yeshion23333 3d2037aaea fix(eval) Fix import async (#1203)
* import async

* import async
2026-04-03 15:54:52 +08:00
yeshion23333 020cc17b3c feat(bot):Single Channel (BotChannel) Integration, Werewolf demo (#1196)
* add bot api channel

* revert index.ts

* werewolf

* werewolf

* werewolf

* werewolf

* werewolf

* revert

* api
2026-04-03 11:59:31 +08:00
yeshion23333 2f4b1480de feat(eval): add locomo eval scripts for openclaw and readme (#1152)
* Locomo eval

* openclaw
2026-04-01 17:57:09 +08:00