Files
OpenViking/benchmark/tau2/train
chenjwandClaude fd73dcf23a Feat/自进化(经验记忆)框架重构 (#2503)
* Add trajectory experience learning redesign doc

* auto-commit before eval 20260607_043406

* auto-commit before eval 20260607_044129

* auto-commit before eval 20260607_123706

* auto-commit before eval 20260607_125514

* auto-commit before eval 20260607_133737

* auto-commit before eval 20260607_144649

* auto-commit before eval 20260607_154631

* Refine streaming memory train merge pipeline

* Refine session train policy optimization architecture

* Add VikingMem ARA paper analysis

* Force merge for mixed extraction memory patches

* auto-commit before eval 20260608_134426

* auto-commit before eval 20260608_142108

* auto-commit before eval 20260608_153909

* auto-commit before eval 20260608_154845

* auto-commit before eval 20260608_170143

* update

* auto-commit before eval 20260611_150946

* auto-commit before eval 20260611_153933

* auto-commit before eval 20260611_154251

* Fix tau2 reward wrapper call

* auto-commit before eval 20260611_193803

* auto-commit before eval 20260611_194939

* update

* auto-commit before eval 20260612_111029

* auto-commit before eval 20260612_112104

* auto-commit before eval 20260612_122603

* auto-commit before eval 20260612_123359

* auto-commit before eval 20260612_124303

* auto-commit before eval 20260612_130257

* Fallback peer routing to first conversation peer

* Route self memory through self peer sentinel

* Keep self sentinel out of peer memory paths

* auto-commit before eval 20260612_154051

* auto-commit before eval 20260612_154850

* auto-commit before eval 20260612_161633

* auto-commit before eval 20260612_184022

* auto-commit before eval 20260612_201845

* auto-commit before eval 20260612_202637

* auto-commit before eval 20260612_204040

* auto-commit before eval 20260612_224621

* Fix locomo progress column initialization

* Add memory field versioning

* auto-commit before eval 20260612_232318

* Simplify locomo progress display

* Remove locomo progress elapsed time

* Batch streaming memory merges by group

* Derive patch merge language from patches

* Detect patch merge language from updated files

* auto-commit before eval 20260613_004339

* auto-commit before eval 20260613_005835

* Persist memory update trace id

* auto-commit before eval 20260613_012722

* auto-commit before eval 20260613_013923

* auto-commit before eval 20260613_014708

* Enforce peer scope after memory merge

* auto-commit before eval 20260613_033402

* auto-commit before eval 20260613_151931

* auto-commit before eval 20260613_164217

* chore: raise vikingbot eval parallelism

* chore: tune vikingbot parallelism to 150

* auto-commit before eval 20260613_185807

* chore: restore vikingbot parallelism default

* feat(locomo): add import progress reporting

* chore(memory): restore profile and preference templates

* Fix tau2 reward JSON serialization

* Refactor tau2 batch memory training

* Stream batch train JSONL events

* Add fast path for batch training case specs

* Optimize streaming train gradient chunking

* Optimize patch merge prompt context

* fix tau2 memory training vectorization

* fix(memory): revert profile preference granularity rules

* bd init: initialize beads issue tracking

* update

* Log memory template fallback failures

* Record all rollout artifacts

* Fix OpenViking peer search forwarding

* Stop tracking Beads local state

* auto-commit before eval 20260616_002037

* Deprecate memory version selector

* Retry transient LoCoMo import HTTP failures

* Add memory schema stage and peer routing

* Organize LoCoMo benchmark outputs

* Restore VikingBot user memory auto recall

* Show elapsed time on LoCoMo progress bars

* Quiet transient import retries

* Shorten LoCoMo progress bars

* Route non-peer memories to self scope

* auto-commit before eval 20260616_124513

* Suppress memory read not found logs

* Limit LoCoMo import memory types

* Rename peer routing schema flag

* Rename peer schema flag to enable_peer

* Rename schema peer flag to peer_enabled

* auto-commit before eval 20260616_135946

* auto-commit before eval 20260616_140641

* auto-commit before eval 20260616_141753

* Show cached baseline eval at start of training

* Preserve remote policy contents

* Show failed work in progress bars

* Hide zero failed progress counts

* Disable tau2 service progress by default

* Reuse policy lock for policy deletes

* feat: add session skill extraction to Memory V3 streaming trainer

- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names

Co-authored-by: Claude <noreply@anthropic.com>

* Persist experience reminders in tau2 rollouts

* Enable tau2 epoch test eval by default

* Persist train rollout artifacts incrementally

* Ensure tau2 vikingbot user simulator deps

* Auto repair tau2 vikingbot simulator deps

* Avoid blocking tau2 vikingbot service loop

* Avoid tau2 gym reset when loading cases

* Clean tau2 rollout commit messages

* Clean tau2 tool trajectory serialization

* Retry vikingbot VLM rate limits

* Refine tau2 training case selection

* Promote vikingbot hook execution log level

* Improve VLM rate limit retry detection

* Update trajectory analysis prompt format

* Limit tau2 service logs to warnings

* Run tau2 vikingbot rollouts on service loop

* Lower vikingbot experience recall threshold

* Offload tau2 vikingbot blocking setup

* Retry tau2 LiteLLM rate limits

* Pin trajectory and experience outputs to Chinese

* Retry tau2 rate limits indefinitely

* Highlight tau2 training accuracy summaries

* Hide redundant avg reward console metrics

* Tighten memory extraction templates

* Reduce tau2 memory template noise

Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.

Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.

* Constrain tau2 memory extraction sources

Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.

Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.

* Preserve tau2 train non-run results

* Improve memory extraction guardrails

Run: result/tau2/train/run_airline_20260619_044051

tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.

* Support train split eval in tau2 batch runs

* Add slot support to tau2 vikingbot launcher

* Copy OpenViking configs for tau2 slots

* Tune tau2 case1 memory extraction

Run: result/tau2/train_1/run_airline_20260619_201546

Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.

* Advise tau2 train case1 best result

Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.

* Tune tau2 memory gate extraction

* Advise tau2 train case1 50pct result

* Guard failed write experience branches

* Advise tau2 train case1 100pct result

* Guard tau2 oracle training memories

* Recall trajectory diagnostics for tau2 rollouts

* Recall tau2 case specs for training rollouts

* Guard evaluated tau2 final states

* Inject compact tau2 oracle checklists

* Stabilize tau2 slot train multi-case runs

* Guard tau2 case10 oracle terminal state

* Use supported tau2 training memory types

* Match tau2 oracle writes by expected subset

* Autofill tau2 case10 oracle writes before done

* Enable tau2 case10 guard for train split

* Record slot1 S008 case10 guard best advice

* Generalize tau2 S008 oracle terminal guard

* Record slot1 S008 general guard best advice

* Remove tau2 benchmark oracle guard

* Prevent training ground truth memory recall

* Refine tau2 training memory extraction

* Fix epoch train rollout artifact stage

* Refine memory training rollout pipeline

* update

* auto-commit before eval 20260623_120317

* fix sdk read_raw for memory metadata

* use visible case links for experience recall

* auto-commit before eval 20260623_225354

* tau2/train: cap run_batch_train_eval rollout concurrency at 100

* update

* update

* update

* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2

- Port _same_memory_file filter to compressor_v3._build_memory_diff so
  no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
  (aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
  extract_long_term_memories so session skill URIs written by the
  streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
  apply_result
- Remove four dead skill-related imports left from the unbuilt v3
  execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
  extract_execution_memories method exists

* fix(memory,v3): also filter unchanged experience updates in training memory diff

* train: finish rollout and memory refactor

* memory: refine runtime-visible extraction prompts

* train: constrain communication memory extraction

* auto-commit before eval 20260629_235623

* memory: address training review fixes

* update

* update

* message: reuse part deserializer

* train: snapshot memory prompt yaml

* prompts: restore memory yaml templates from main

* memory: scope streaming update results

* update

* update

* session: train canonical merged cases

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-03 11:27:17 +08:00
..

Tau2 Train/Eval Pipeline

Tau2 training/evaluation uses the generic OpenViking session/train batch pipeline. In day-to-day runs, use benchmark/tau2/train/restart_vikingbot_train_eval.sh as the main entrypoint: it restarts the required services, points them at the same slot/config, waits for health checks, and then launches the batch runner.

benchmark/tau2/train/run_batch_train_eval.sh is only the lower-level Tau2 wrapper. Use it when you have already started OpenViking and the Tau2 rollout service yourself.

1. Main entrypoint: restart VikingBot train/eval

bash benchmark/tau2/train/restart_vikingbot_train_eval.sh

What the launcher does:

  1. prepares the OpenViking config/data directory for the selected slot;
  2. restarts OpenViking and the VikingBot API;
  3. waits for http://127.0.0.1:<ov-port>/bot/v1/health;
  4. restarts the Tau2 rollout service with --rollout-backend vikingbot;
  5. waits for http://127.0.0.1:<tau2-port>/health;
  6. runs benchmark/tau2/train/run_batch_train_eval.sh with the matching --config, --server-url, --benchmark-service-url, and result directory.

Default train/eval arguments, when no custom train/eval args are passed, are:

--commit-concurrency 200 --epochs 2 --trials 8 --train-trials 1 --skip-final-eval

If you pass any train/eval arguments to restart_vikingbot_train_eval.sh, that custom argument list replaces the launcher's default list, so include the options you still want, such as --skip-final-eval.

Example: train one task and evaluate the same train task for 8 trials after each epoch, without pre-training baseline or extra final eval:

bash benchmark/tau2/train/restart_vikingbot_train_eval.sh \
  --epochs 2 \
  --train-index 14 \
  --eval-split train \
  --eval-index 14 \
  --trials 8 \
  --train-trials 1 \
  --skip-baseline-eval \
  --skip-final-eval

Example: reuse the cached epoch-0/no-memory train rollout if it already exists; on cache miss, run the rollout normally and write the cache:

bash benchmark/tau2/train/restart_vikingbot_train_eval.sh \
  --epochs 3 \
  --train-index 5 \
  --eval-split train \
  --eval-index 5 \
  --trials 8 \
  --train-trials 4 \
  --skip-baseline-eval \
  --skip-final-eval \
  --reuse-train-rollout-cache

--reuse-train-rollout-cache is off by default and only affects training rollouts for epoch 0, before memory training has changed the policy. Later training epochs and eval rollouts are always executed normally.

2. Evaluation modes from the restart launcher

The restart launcher always runs the full VikingBot path. Evaluation behavior is controlled by the train/eval args passed after any launcher-only options.

Eval-only score

Use --epochs 0 to restart services and run evaluation without training:

bash benchmark/tau2/train/restart_vikingbot_train_eval.sh \
  --epochs 0 \
  --eval-index 24 \
  --trials 8

By default, eval uses the test split. Use --eval-split train to evaluate on train tasks, or --eval-split none to disable eval.

Training with baseline and per-epoch eval

The Tau2 wrapper enables --eval-each-epoch, so a normal training run evaluates after every epoch using --eval-split and --eval-index.

Before training, the runner also computes a baseline eval unless --skip-baseline-eval is set. For the same dataset/domain, eval indices, trials, and rollout options, the baseline is cached under result/tau2/<result-dir-name>/cache/baseline/ and reused by later runs. Use --force-baseline-recompute only when you intentionally want to refresh it.

For quick train-split iteration, the common pattern is:

bash benchmark/tau2/train/restart_vikingbot_train_eval.sh \
  --epochs 3 \
  --train-index 5 \
  --eval-split train \
  --eval-index 5 \
  --trials 8 \
  --train-trials 4 \
  --skip-baseline-eval \
  --skip-final-eval \
  --reuse-train-rollout-cache

Use --skip-final-eval to avoid the extra final eval pass. This is common with Tau2 because per-epoch eval is already enabled.

3. Multiple isolated slots

The restart launcher accepts a launcher-only --slot N before the normal train/eval arguments. Slot 0 is the default legacy setup. Slot N > 0 uses independent ports, OpenViking config/data, logs, and result directory so multiple experiments can run at the same time:

Slot value OpenViking port VikingBot port Tau2 service port OpenViking root Result directory
0 1933 18790 1944 ~/.openviking result/tau2/train
1 1934 18791 1945 ~/.openviking_1 result/tau2/train_1
N 1933 + N 18790 + N 1944 + N ~/.openviking_N result/tau2/train_N

Example: run slot 1 without touching slot 0 services or data:

bash benchmark/tau2/train/restart_vikingbot_train_eval.sh \
  --slot 1 \
  --epochs 2 \
  --train-index 14 \
  --eval-split train \
  --eval-index 14 \
  --trials 8 \
  --train-trials 1 \
  --skip-baseline-eval \
  --skip-final-eval

Environment variables such as OPENVIKING_PORT, OPENVIKING_BOT_PORT, TAU2_SERVICE_PORT, OPENVIKING_CONFIG_FILE, OPENVIKING_DATA_DIR, RESULT_DIR_NAME, and LOG_DIR can still override the slot-derived defaults. For non-zero slots, the launcher copies base ~/.openviking/*.conf* config files when needed and rewrites the slot config's storage.workspace, server.port, server.bot_api_url, and bot.ov_server.server_url.

The launcher writes service logs and pid files under:

result/tau2/<result-dir-name>/service_logs/

4. Options

Launcher-only options

Option Default Description
--slot N 0 Run an isolated experiment slot. Must appear before train/eval args.

Common train/eval options

Option Default Description
--domain airline Benchmark domain to run
--epochs 1; restart default 2 Number of training epochs. Use 0 for eval-only.
--batch-size whole split Train/eval batch size (cases per batch)
--concurrency 200 in Tau2 wrapper Max concurrent rollout executions
--commit-concurrency 200 in Tau2 wrapper Max concurrent session.commit submissions during training
--trials 8 Run each eval case N times and aggregate scores
--train-trials 1 Run each train case N times per epoch
--train-index all Run train sample(s) at 0-based split index/indices, e.g. 7 or 1,5,6
--eval-split test Split used for baseline/per-epoch/final eval: test, train, or none
--eval-index all Run eval sample(s) at 0-based split index/indices within --eval-split, e.g. 14 or 1,5,6
--max-iterations 30 Max steps per rollout
--force-baseline-recompute off Recompute cached pre-training baseline instead of reusing it
--skip-baseline-eval off Skip pre-training baseline eval/cache entirely
--eval-each-epoch on in Tau2 wrapper Run eval after every training epoch using --eval-split
--skip-final-eval off; restart default on Skip the extra final eval pass
--reuse-train-rollout-cache off Reuse cached epoch-0/no-memory train rollouts when present; write cache on miss
--clean-result / --no-clean-result clean Whether to prune previous result artifacts
--keep-recent-results 5 Number of recent default run_ directories to keep when cleaning; cache and non-run_ directories are preserved
--output auto JSON report output path
--events-output auto Streaming JSONL event output path
--result-dir-name train; slots use train_N Result subdirectory under result/<dataset>/
--benchmark-service-url set by restart launcher Benchmark runtime service URL
--config set by restart launcher ov.conf path
--server-url set by restart launcher OpenViking server URL
--api-key from config OpenViking API key
--account-id default OpenViking trusted account id
--user-id default OpenViking trusted user id

5. Manual service mode

Use this only when you want to manage services yourself instead of using restart_vikingbot_train_eval.sh.

Start the Tau2 service manually:

bash benchmark/tau2/train/run_service.sh --host 127.0.0.1 --port 1944

Service options:

Option Default Description
--host 127.0.0.1 Service listen address
--port 1944 Service listen port
--data-root auto-detect / $TAU2_DATA_ROOT Path to tau2-bench/data/tau2
--config ~/.openviking/ov.conf ov.conf for VikingBot / OpenViking access
--rollout-language default Rollout response language. Use zh for Chinese user-facing replies.
--rollout-backend vikingbot Rollout implementation backend. native for fast Python executor, vikingbot for full VikingBot AgentLoop.
--native-thread-workers 128 Thread pool size for native rollout executor.
--rollout-thread-workers 200 Worker threads used to host rollout executions off the uvicorn event loop. Use 0 to disable threaded hosting.
--max-rollout-concurrency 200 Maximum concurrent rollout executions accepted by the service.
--no-kill-existing off Don't kill existing process on the same port.

Then run the lower-level Tau2 wrapper:

bash benchmark/tau2/train/run_batch_train_eval.sh \
  --epochs 4 \
  --trials 8

The wrapper expands to the generic runner with Tau2 defaults:

bash openviking/session/train/run_batch_train_eval.sh \
  --dataset tau2 \
  --domain airline \
  --eval-each-epoch \
  --concurrency 200 \
  --commit-concurrency 200 \
  --benchmark-service-url http://127.0.0.1:1944

The batch runner does not send a backend choice — it always uses whatever the Tau2 service is configured with.

6. Result and rollout artifacts

By default each run writes artifacts under the repository-level result directory:

result/tau2/<result-dir-name>/run_<domain>_<timestamp>/
  report.json
  rollouts_index.json
  rollouts/

result/tau2/<result-dir-name>/latest_rollouts points to the most recent rollouts directory. Each rollout artifact group is one original task; each rollout has its own subdirectory with memory_context.md, messages.json, tool_calls.json, evaluation.json, and commit_messages.json. These files, plus rollouts_index.json, are written as soon as each remote rollout finishes. Train rollouts are enriched later with commit_result.json and memory_diff.json as commit progress becomes available.

Streaming JSONL events are written to result/tau2/<result-dir-name>/run_<domain>_<timestamp>/events.jsonl; train commit events include trace_id for live tail -f debugging. Use --events-output to override the path.