Commit Graph
10 Commits
Author SHA1 Message Date
chenjwandClaude fd73dcf23a Feat/自进化(经验记忆)框架重构 (#2503)
* Add trajectory experience learning redesign doc

* auto-commit before eval 20260607_043406

* auto-commit before eval 20260607_044129

* auto-commit before eval 20260607_123706

* auto-commit before eval 20260607_125514

* auto-commit before eval 20260607_133737

* auto-commit before eval 20260607_144649

* auto-commit before eval 20260607_154631

* Refine streaming memory train merge pipeline

* Refine session train policy optimization architecture

* Add VikingMem ARA paper analysis

* Force merge for mixed extraction memory patches

* auto-commit before eval 20260608_134426

* auto-commit before eval 20260608_142108

* auto-commit before eval 20260608_153909

* auto-commit before eval 20260608_154845

* auto-commit before eval 20260608_170143

* update

* auto-commit before eval 20260611_150946

* auto-commit before eval 20260611_153933

* auto-commit before eval 20260611_154251

* Fix tau2 reward wrapper call

* auto-commit before eval 20260611_193803

* auto-commit before eval 20260611_194939

* update

* auto-commit before eval 20260612_111029

* auto-commit before eval 20260612_112104

* auto-commit before eval 20260612_122603

* auto-commit before eval 20260612_123359

* auto-commit before eval 20260612_124303

* auto-commit before eval 20260612_130257

* Fallback peer routing to first conversation peer

* Route self memory through self peer sentinel

* Keep self sentinel out of peer memory paths

* auto-commit before eval 20260612_154051

* auto-commit before eval 20260612_154850

* auto-commit before eval 20260612_161633

* auto-commit before eval 20260612_184022

* auto-commit before eval 20260612_201845

* auto-commit before eval 20260612_202637

* auto-commit before eval 20260612_204040

* auto-commit before eval 20260612_224621

* Fix locomo progress column initialization

* Add memory field versioning

* auto-commit before eval 20260612_232318

* Simplify locomo progress display

* Remove locomo progress elapsed time

* Batch streaming memory merges by group

* Derive patch merge language from patches

* Detect patch merge language from updated files

* auto-commit before eval 20260613_004339

* auto-commit before eval 20260613_005835

* Persist memory update trace id

* auto-commit before eval 20260613_012722

* auto-commit before eval 20260613_013923

* auto-commit before eval 20260613_014708

* Enforce peer scope after memory merge

* auto-commit before eval 20260613_033402

* auto-commit before eval 20260613_151931

* auto-commit before eval 20260613_164217

* chore: raise vikingbot eval parallelism

* chore: tune vikingbot parallelism to 150

* auto-commit before eval 20260613_185807

* chore: restore vikingbot parallelism default

* feat(locomo): add import progress reporting

* chore(memory): restore profile and preference templates

* Fix tau2 reward JSON serialization

* Refactor tau2 batch memory training

* Stream batch train JSONL events

* Add fast path for batch training case specs

* Optimize streaming train gradient chunking

* Optimize patch merge prompt context

* fix tau2 memory training vectorization

* fix(memory): revert profile preference granularity rules

* bd init: initialize beads issue tracking

* update

* Log memory template fallback failures

* Record all rollout artifacts

* Fix OpenViking peer search forwarding

* Stop tracking Beads local state

* auto-commit before eval 20260616_002037

* Deprecate memory version selector

* Retry transient LoCoMo import HTTP failures

* Add memory schema stage and peer routing

* Organize LoCoMo benchmark outputs

* Restore VikingBot user memory auto recall

* Show elapsed time on LoCoMo progress bars

* Quiet transient import retries

* Shorten LoCoMo progress bars

* Route non-peer memories to self scope

* auto-commit before eval 20260616_124513

* Suppress memory read not found logs

* Limit LoCoMo import memory types

* Rename peer routing schema flag

* Rename peer schema flag to enable_peer

* Rename schema peer flag to peer_enabled

* auto-commit before eval 20260616_135946

* auto-commit before eval 20260616_140641

* auto-commit before eval 20260616_141753

* Show cached baseline eval at start of training

* Preserve remote policy contents

* Show failed work in progress bars

* Hide zero failed progress counts

* Disable tau2 service progress by default

* Reuse policy lock for policy deletes

* feat: add session skill extraction to Memory V3 streaming trainer

- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names

Co-authored-by: Claude <noreply@anthropic.com>

* Persist experience reminders in tau2 rollouts

* Enable tau2 epoch test eval by default

* Persist train rollout artifacts incrementally

* Ensure tau2 vikingbot user simulator deps

* Auto repair tau2 vikingbot simulator deps

* Avoid blocking tau2 vikingbot service loop

* Avoid tau2 gym reset when loading cases

* Clean tau2 rollout commit messages

* Clean tau2 tool trajectory serialization

* Retry vikingbot VLM rate limits

* Refine tau2 training case selection

* Promote vikingbot hook execution log level

* Improve VLM rate limit retry detection

* Update trajectory analysis prompt format

* Limit tau2 service logs to warnings

* Run tau2 vikingbot rollouts on service loop

* Lower vikingbot experience recall threshold

* Offload tau2 vikingbot blocking setup

* Retry tau2 LiteLLM rate limits

* Pin trajectory and experience outputs to Chinese

* Retry tau2 rate limits indefinitely

* Highlight tau2 training accuracy summaries

* Hide redundant avg reward console metrics

* Tighten memory extraction templates

* Reduce tau2 memory template noise

Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.

Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.

* Constrain tau2 memory extraction sources

Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.

Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.

* Preserve tau2 train non-run results

* Improve memory extraction guardrails

Run: result/tau2/train/run_airline_20260619_044051

tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.

* Support train split eval in tau2 batch runs

* Add slot support to tau2 vikingbot launcher

* Copy OpenViking configs for tau2 slots

* Tune tau2 case1 memory extraction

Run: result/tau2/train_1/run_airline_20260619_201546

Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.

* Advise tau2 train case1 best result

Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.

* Tune tau2 memory gate extraction

* Advise tau2 train case1 50pct result

* Guard failed write experience branches

* Advise tau2 train case1 100pct result

* Guard tau2 oracle training memories

* Recall trajectory diagnostics for tau2 rollouts

* Recall tau2 case specs for training rollouts

* Guard evaluated tau2 final states

* Inject compact tau2 oracle checklists

* Stabilize tau2 slot train multi-case runs

* Guard tau2 case10 oracle terminal state

* Use supported tau2 training memory types

* Match tau2 oracle writes by expected subset

* Autofill tau2 case10 oracle writes before done

* Enable tau2 case10 guard for train split

* Record slot1 S008 case10 guard best advice

* Generalize tau2 S008 oracle terminal guard

* Record slot1 S008 general guard best advice

* Remove tau2 benchmark oracle guard

* Prevent training ground truth memory recall

* Refine tau2 training memory extraction

* Fix epoch train rollout artifact stage

* Refine memory training rollout pipeline

* update

* auto-commit before eval 20260623_120317

* fix sdk read_raw for memory metadata

* use visible case links for experience recall

* auto-commit before eval 20260623_225354

* tau2/train: cap run_batch_train_eval rollout concurrency at 100

* update

* update

* update

* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2

- Port _same_memory_file filter to compressor_v3._build_memory_diff so
  no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
  (aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
  extract_long_term_memories so session skill URIs written by the
  streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
  apply_result
- Remove four dead skill-related imports left from the unbuilt v3
  execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
  extract_execution_memories method exists

* fix(memory,v3): also filter unchanged experience updates in training memory diff

* train: finish rollout and memory refactor

* memory: refine runtime-visible extraction prompts

* train: constrain communication memory extraction

* auto-commit before eval 20260629_235623

* memory: address training review fixes

* update

* update

* message: reuse part deserializer

* train: snapshot memory prompt yaml

* prompts: restore memory yaml templates from main

* memory: scope streaming update results

* update

* update

* session: train canonical merged cases

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-03 11:27:17 +08:00
Qin Haojie ff258768c2 feat(memory): 引入 User/Peer 记忆隔离模型 (#2236)
* feat(memory): introduce user and peer memory isolation

Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.

* feat(memory): align session identity around peer IDs

* feat(search): pass peer id through retrieval

* refactor(memory): remove agent identity from integrations

* fix(memory): isolate peer identity from self extraction

* fix(tau2): provision benchmark user configs

* fix(auth): allow admin keys to access data APIs

* fix(openclaw): enable peer memory policy for peer roles

* fix(openclaw): resolve sender for peer recall

* refactor(session): simplify memory extraction routing

* refactor(ov-cli): reduce formatting-only diff

* refactor(message): remove unused message helpers

* refactor(retrieval): simplify peer target resolution

* refactor(namespace): remove deprecated agent namespace policy

* fix(agent): propagate peer id through integrations

* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
yangxinxin-7andClaude Sonnet 4.6 936624c24b feat(tau2/vikingbot): config-driven experience recall + per-domain isolation (#2380)
* feat(tau2/vikingbot): config-driven experience recall + per-domain isolation

Switch tau2 self-improvement behaviour from core-code patches to three
ov.conf flags (recall_exp_first_round_only, exp_recall_limit,
exp_recall_max_chars), so the VikingBot core is unchanged for non-tau2
users.

- context.py: when recall_exp_first_round_only=true, skip per-turn
  user+agent memory retrieval and inject experience once on the first
  user-turn; accepts explicit agent_id to scope retrieval per domain
- memory.py: exp_recall_limit and exp_recall_max_chars read from config
  instead of hardcoded values
- schema.py: add three new OpenVikingConfig fields (all default to
  existing behaviour so existing deployments are unaffected)
- ov_server.py: extract _is_session_key() helper to unify the two
  places that distinguish session keys from per-domain agent ids;
  local mode now respects agent_id for namespace isolation (remote mode
  already supported this)
- tau2 runner: pass agent_id= instead of memory_users= to build_messages
- README: document Python >=3.12 prerequisite, correct pip install
  extra, clarify that isolation works in both local and remote modes

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ov_server): remove redundant underscore check in search_experiences

The "_" in self.agent_id guard was a leftover heuristic to distinguish
domain ids from session keys. Now that _is_session_key() handles that
check via "__", the extra "_" condition is unnecessary and actively
breaks agent ids without underscores (e.g. "airline", "retail").

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-02 15:04:40 +08:00
huangruiteng 646ff735db docs: simplify tau2 benchmark reproduction (#2267) 2026-05-27 20:18:49 +08:00
huangruiteng 76c8f559e1 feat(memory): add trajectory retrieval anchor (#2255) 2026-05-27 14:37:40 +08:00
e0ce670f5c feat(tau2/vikingbot): benchmark updates (#2244)
* feat(benchmark/tau2): add VikingBot agent runner for tau2-bench

Adds benchmark/tau2/vikingbot/, an end-to-end harness that runs the full
VikingBot AgentLoop on tau2-bench tasks and commits trajectories back into
OpenViking memory for epoch-based self-improvement. This complements the
existing memory-retrieval harness in benchmark/tau2/ (which is retrieval-only).

Contents:
- scripts/vikingbot_tau2_runner.py: run one tau2 task through the agent loop
  (tau2 tool registry swap, simulated-time patch, advisory memory scope guard).
- scripts/run_tau2_domain.sh / run_eval_reward.sh: run a domain split with
  bounded concurrency and score average reward.
- scripts/commit_trajectory_to_memory.py: commit train trajectories to memory.
- scripts/stat_trajectory.py, check_openviking_tool_calls.py: analysis helpers.
- tau2_env/: tau2 environment + tool-provider integration.
- run_full_test.sh and run_{airline,retail}_*epochs.sh: full / multi-epoch runs.
- setup_env.sh, README.md, .gitignore.

tau2-bench is referenced as an external dependency (cloned + installed by the
user); no OpenViking core changes are required. The runner is API-compatible
with bot/vikingbot on current main.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(benchmark/tau2): split into llm/ and vikingbot/ subfolders

Mirror the two evaluation approaches as sibling subfolders under benchmark/tau2/:

- llm/: the existing OpenViking Memory V2 retrieval harness, moved from
  benchmark/tau2/. All internal benchmark/tau2/... path references and the
  REPO_ROOT depth computations (run_full_eval.sh, tau2_common.py,
  run_memory_v2_eval.py) are updated for the extra directory level.
- vikingbot/: the VikingBot agent runner (added in the previous commit).

vikingbot/ cleanup:
- make memory-block extraction time-independent: anchor on the stable session
  header and trailing reply instruction instead of a fixed simulated timestamp
  (the sim-time patch was removed, so the current time is now system-generated).
- drop the now-removed sim-time / scope-guard notes from the README.
- remove the unused stat_trajectory.py and check_openviking_tool_calls.py helpers.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(benchmark/tau2/vikingbot): train-once/test-8x eval, drop smolagents, doc updates

- run_full_test.sh: run train once per epoch (experience extraction) and test
  N times in parallel (--test-repeats, default 8), reporting the averaged test
  accuracy; keep --commit/--no-commit.
- tau2_environment.py: remove the unused smolagents Tool path (CommunicateWithUser /
  create_tool_from_json_schema / self.tools); communicate_with_user is handled
  directly in tool_call. tau2-bench has no smolagents dependency, so it is dropped.
- README: reorder install (tau2-bench first so setup_env can derive TAU2_DATA_ROOT),
  explain train-once/test-8x methodology and train-only memory extraction, document
  the required bot/vikingbot core changes (agent_id isolation + agent-experience
  memory), fix sibling links to ../llm/.
- Remove run_retail_3epochs.sh.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(tau2/vikingbot): one-step setup_env.sh + communicate_with_user refactor

setup_env.sh now does full environment setup in a single `source`: creates a
fresh repo-root .venv, clones tau2-bench (external dep), installs openviking +
vikingbot (pip install -e ., runs the Cargo build) + tau2-bench + smolagents,
then activates and exports the runtime env vars. Idempotent via a marker file;
supports --reinstall. README updated to document the one-step flow and the
overridable env vars.

Also move the communicate_with_user tool into a CommunicateWithUser class in
tau2_environment.py (owns both schema and execution) and drop the duplicated
inline schema from tau2_tool_provider.py.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(tau2/vikingbot): sync setup_env.sh fixes + README port/diff clarifications

Backport the environment-setup fixes and README clarifications discovered while
running the harness end-to-end (the core bot/vikingbot code changes live on the
test/tau2-vikingbot-core-changes branch, not here):

- setup_env.sh: install the [bot] extra (prompt_toolkit/gradio/mcp/...), build +
  bundle ragfs_python via maturin when the editable install skips it under pip
  build isolation, and install tau2-bench with the [gym] extra (gymnasium)
- README.md: explain the server port (default 1933 vs bot.ov_server.server_url)
  and show the None-safe forms of the Change-1 diffs

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* clean README message

---------

Co-authored-by: ByteDance <wenting.qi@bytedance.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-26 19:35:31 +08:00
huangruiteng 9fffdf4a52 docs(tau2): clarify fixed-user PR-B reproduction (#2223) 2026-05-25 20:50:05 +08:00
huangruiteng 8817abdd38 bench(tau2): add generic scope fairness check (#2172)
* bench(tau2): add generic scope fairness check

* bench(tau2): focus scope fairness on generic prompt

* bench(tau2): remove domain-specific scope prompts
2026-05-21 19:51:14 +08:00
huangruitengandhuangruiteng 520713c624 feat(memory): upgrade trajectory extraction to beat no-memory baseline (#2017)
* feat(benchmark): add TAU-2 trajectory memory treatment

* style(benchmark): format tau2 trajectory scripts

* refine trajectory memory view prompt

* feat(benchmark): prepare tau2 memory corpora before eval

* fix(benchmark): tighten trajectory evidence prompt

* fix(benchmark): guard tau2 infrastructure failures

* fix(benchmark): resolve tau2 runner paths

* fix(memory): add trajectory evidence examples

* fix(benchmark): run no-memory tau2 eval in process

* bench(tau2): align retrieval budget and fixed first user

* bench(tau2): reuse memory corpora across eval runs

* bench(tau2): add scoped trajectory eval concurrency

* style(benchmark): format tau2 eval runner

* style(benchmark): satisfy tau2 eval lint

* bench(tau2): harden trajectory memory eval variants

* fix(tau2): rebuild search URI for reused corpora

* docs(tau2): add PR-B reproduction commands

* chore(tau2): keep PR-B benchmark scope focused

* chore(tau2): keep trajectory prompt generic

* chore(tau2): format benchmark scripts

* bench(tau2): restore operation-family trajectory protocol

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-05-21 11:15:34 +08:00
huangruitengandhuangruiteng c9ebba0734 feat(benchmark): add TAU-2 Memory V2 eval runner (#2003)
* benchmark: add tau2 eval scaffold

* benchmark: gate pending tau2 memory adapter

* benchmark: use litellm provider model default

* benchmark: fold preflight into tau2 runner

* benchmark: document tau2 dependency setup

* benchmark: simplify tau2 simulator patch

* benchmark: keep simulator patch prompt clean

* benchmark: clarify simulator patch config

* benchmark: clarify tau2 adapter boundary

* benchmark: wire tau2 memory v2 eval

* benchmark: harden tau2 memory agent tool calls

* benchmark: tolerate empty tau2 assistant responses

* benchmark: normalize tau2 llm environment

* benchmark: add tau2 memory prewrite strategy

* benchmark: support current tau2 runner api

* benchmark: align tau2 memory prewrite parity

* benchmark: make tau2 eval traces safer

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-05-13 15:42:06 +08:00