Files
OpenViking/benchmark/tau2/llm
chenjwandClaude fd73dcf23a Feat/自进化(经验记忆)框架重构 (#2503)
* Add trajectory experience learning redesign doc

* auto-commit before eval 20260607_043406

* auto-commit before eval 20260607_044129

* auto-commit before eval 20260607_123706

* auto-commit before eval 20260607_125514

* auto-commit before eval 20260607_133737

* auto-commit before eval 20260607_144649

* auto-commit before eval 20260607_154631

* Refine streaming memory train merge pipeline

* Refine session train policy optimization architecture

* Add VikingMem ARA paper analysis

* Force merge for mixed extraction memory patches

* auto-commit before eval 20260608_134426

* auto-commit before eval 20260608_142108

* auto-commit before eval 20260608_153909

* auto-commit before eval 20260608_154845

* auto-commit before eval 20260608_170143

* update

* auto-commit before eval 20260611_150946

* auto-commit before eval 20260611_153933

* auto-commit before eval 20260611_154251

* Fix tau2 reward wrapper call

* auto-commit before eval 20260611_193803

* auto-commit before eval 20260611_194939

* update

* auto-commit before eval 20260612_111029

* auto-commit before eval 20260612_112104

* auto-commit before eval 20260612_122603

* auto-commit before eval 20260612_123359

* auto-commit before eval 20260612_124303

* auto-commit before eval 20260612_130257

* Fallback peer routing to first conversation peer

* Route self memory through self peer sentinel

* Keep self sentinel out of peer memory paths

* auto-commit before eval 20260612_154051

* auto-commit before eval 20260612_154850

* auto-commit before eval 20260612_161633

* auto-commit before eval 20260612_184022

* auto-commit before eval 20260612_201845

* auto-commit before eval 20260612_202637

* auto-commit before eval 20260612_204040

* auto-commit before eval 20260612_224621

* Fix locomo progress column initialization

* Add memory field versioning

* auto-commit before eval 20260612_232318

* Simplify locomo progress display

* Remove locomo progress elapsed time

* Batch streaming memory merges by group

* Derive patch merge language from patches

* Detect patch merge language from updated files

* auto-commit before eval 20260613_004339

* auto-commit before eval 20260613_005835

* Persist memory update trace id

* auto-commit before eval 20260613_012722

* auto-commit before eval 20260613_013923

* auto-commit before eval 20260613_014708

* Enforce peer scope after memory merge

* auto-commit before eval 20260613_033402

* auto-commit before eval 20260613_151931

* auto-commit before eval 20260613_164217

* chore: raise vikingbot eval parallelism

* chore: tune vikingbot parallelism to 150

* auto-commit before eval 20260613_185807

* chore: restore vikingbot parallelism default

* feat(locomo): add import progress reporting

* chore(memory): restore profile and preference templates

* Fix tau2 reward JSON serialization

* Refactor tau2 batch memory training

* Stream batch train JSONL events

* Add fast path for batch training case specs

* Optimize streaming train gradient chunking

* Optimize patch merge prompt context

* fix tau2 memory training vectorization

* fix(memory): revert profile preference granularity rules

* bd init: initialize beads issue tracking

* update

* Log memory template fallback failures

* Record all rollout artifacts

* Fix OpenViking peer search forwarding

* Stop tracking Beads local state

* auto-commit before eval 20260616_002037

* Deprecate memory version selector

* Retry transient LoCoMo import HTTP failures

* Add memory schema stage and peer routing

* Organize LoCoMo benchmark outputs

* Restore VikingBot user memory auto recall

* Show elapsed time on LoCoMo progress bars

* Quiet transient import retries

* Shorten LoCoMo progress bars

* Route non-peer memories to self scope

* auto-commit before eval 20260616_124513

* Suppress memory read not found logs

* Limit LoCoMo import memory types

* Rename peer routing schema flag

* Rename peer schema flag to enable_peer

* Rename schema peer flag to peer_enabled

* auto-commit before eval 20260616_135946

* auto-commit before eval 20260616_140641

* auto-commit before eval 20260616_141753

* Show cached baseline eval at start of training

* Preserve remote policy contents

* Show failed work in progress bars

* Hide zero failed progress counts

* Disable tau2 service progress by default

* Reuse policy lock for policy deletes

* feat: add session skill extraction to Memory V3 streaming trainer

- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names

Co-authored-by: Claude <noreply@anthropic.com>

* Persist experience reminders in tau2 rollouts

* Enable tau2 epoch test eval by default

* Persist train rollout artifacts incrementally

* Ensure tau2 vikingbot user simulator deps

* Auto repair tau2 vikingbot simulator deps

* Avoid blocking tau2 vikingbot service loop

* Avoid tau2 gym reset when loading cases

* Clean tau2 rollout commit messages

* Clean tau2 tool trajectory serialization

* Retry vikingbot VLM rate limits

* Refine tau2 training case selection

* Promote vikingbot hook execution log level

* Improve VLM rate limit retry detection

* Update trajectory analysis prompt format

* Limit tau2 service logs to warnings

* Run tau2 vikingbot rollouts on service loop

* Lower vikingbot experience recall threshold

* Offload tau2 vikingbot blocking setup

* Retry tau2 LiteLLM rate limits

* Pin trajectory and experience outputs to Chinese

* Retry tau2 rate limits indefinitely

* Highlight tau2 training accuracy summaries

* Hide redundant avg reward console metrics

* Tighten memory extraction templates

* Reduce tau2 memory template noise

Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.

Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.

* Constrain tau2 memory extraction sources

Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.

Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.

* Preserve tau2 train non-run results

* Improve memory extraction guardrails

Run: result/tau2/train/run_airline_20260619_044051

tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.

* Support train split eval in tau2 batch runs

* Add slot support to tau2 vikingbot launcher

* Copy OpenViking configs for tau2 slots

* Tune tau2 case1 memory extraction

Run: result/tau2/train_1/run_airline_20260619_201546

Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.

* Advise tau2 train case1 best result

Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.

* Tune tau2 memory gate extraction

* Advise tau2 train case1 50pct result

* Guard failed write experience branches

* Advise tau2 train case1 100pct result

* Guard tau2 oracle training memories

* Recall trajectory diagnostics for tau2 rollouts

* Recall tau2 case specs for training rollouts

* Guard evaluated tau2 final states

* Inject compact tau2 oracle checklists

* Stabilize tau2 slot train multi-case runs

* Guard tau2 case10 oracle terminal state

* Use supported tau2 training memory types

* Match tau2 oracle writes by expected subset

* Autofill tau2 case10 oracle writes before done

* Enable tau2 case10 guard for train split

* Record slot1 S008 case10 guard best advice

* Generalize tau2 S008 oracle terminal guard

* Record slot1 S008 general guard best advice

* Remove tau2 benchmark oracle guard

* Prevent training ground truth memory recall

* Refine tau2 training memory extraction

* Fix epoch train rollout artifact stage

* Refine memory training rollout pipeline

* update

* auto-commit before eval 20260623_120317

* fix sdk read_raw for memory metadata

* use visible case links for experience recall

* auto-commit before eval 20260623_225354

* tau2/train: cap run_batch_train_eval rollout concurrency at 100

* update

* update

* update

* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2

- Port _same_memory_file filter to compressor_v3._build_memory_diff so
  no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
  (aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
  extract_long_term_memories so session skill URIs written by the
  streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
  apply_result
- Remove four dead skill-related imports left from the unbuilt v3
  execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
  extract_execution_memories method exists

* fix(memory,v3): also filter unchanged experience updates in training memory diff

* train: finish rollout and memory refactor

* memory: refine runtime-visible extraction prompts

* train: constrain communication memory extraction

* auto-commit before eval 20260629_235623

* memory: address training review fixes

* update

* update

* message: reuse part deserializer

* train: snapshot memory prompt yaml

* prompts: restore memory yaml templates from main

* memory: scope streaming update results

* update

* update

* session: train canonical merged cases

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-03 11:27:17 +08:00
..

TAU-2 Benchmark

This directory contains the OpenViking TAU-2 LLM benchmark entry point. The reproduction surface is intentionally narrow:

  • no_memory: same-seed TAU-2 baseline without OpenViking memory injection;
  • template_indexed_trajectory_top4_prewrite_top2: the current best template-indexed trajectory memory treatment.

The template-indexed trajectory treatment trains OpenViking Memory V2 from TAU-2 train conversations, retrieves generated trajectories, and uses the trajectory embedding template {{ trajectory_name }}\n\n{{ retrieval_anchor }} instead of broad procedure bodies for retrieval. It injects trajectory top4 at the first user turn and top2 before write-like tool calls, with the generic memory scope prompt enabled.

Category rerank, experience-memory routes, fixed-count-only ablations, character-budget ablations, and official-user parity controls are intentionally left out of this README and config set so reproduction agents do not mistake diagnostic routes for current evidence.

Layout

benchmark/tau2/llm/
├── config/
│   ├── baseline.yaml
│   ├── fixed_first_user_bootstrap.yaml
│   ├── no_memory.yaml
│   ├── scope_prompts/
│   │   └── generic_memory_scope.md
│   └── template_indexed_trajectory.yaml
├── scripts/
│   ├── build_fixed_first_user_fixture.py
│   ├── run_eval.py
│   ├── setup_tau2_repo.sh
│   └── tau2_common.py
└── run_full_eval.sh

baseline.yaml is a shared protocol/defaults file, not a runnable evidence cell by itself. Use no_memory.yaml for the baseline-only run and template_indexed_trajectory.yaml for the paired no-memory + trajectory run.

Generated eval artifacts are written to benchmark/tau2/llm/result/<run_id>/. Memory corpus artifacts are cached outside the run id at benchmark/tau2/llm/result/memory_corpora/ by default.

Setup

This benchmark delegates task simulation and scoring to an external TAU-2 checkout. Point the runner at that checkout and CLI explicitly when they are not on the default path:

export TAU2_REPO=/path/to/tau2-bench
export TAU2_CLI=/path/to/tau2

For a local one-command setup, clone and install TAU-2 into ignored benchmark directories:

benchmark/tau2/llm/scripts/setup_tau2_repo.sh
source benchmark/tau2/llm/.env.tau2

The default OpenViking TAU-2 memory evidence protocol is fixed_first_user_full8: retail + airline, 8 repeats, same seeds, confirmation-aware user simulator, and fixed first-user fixtures for both domains. Later user simulator turns remain live.

The confirmation-aware simulator behavior is available from sierra-research/tau2-bench#297. Pin the local TAU-2 checkout to a ref that includes that behavior when reproducing these numbers:

benchmark/tau2/llm/scripts/setup_tau2_repo.sh \
  --ref refs/pull/297/head
source benchmark/tau2/llm/.env.tau2

When using Doubao through an OpenAI-compatible endpoint, set OPENAI_API_KEY and OPENAI_API_BASE for LiteLLM before running upstream TAU-2.

Fixed-First-User Fixtures

Strict reproduction requires fixed first-user fixtures:

export TAU2_RETAIL_FIXED_FIRST_USER_FILE=/path/to/retail/fixed_first_user_fixture.json
export TAU2_AIRLINE_FIXED_FIRST_USER_FILE=/path/to/airline/fixed_first_user_fixture.json

--strict-preflight fails when eval.require_fixed_first_user=true and either fixture is missing.

For a fresh checkout, run one live-user bootstrap pass per domain:

benchmark/tau2/llm/run_full_eval.sh \
  --config benchmark/tau2/llm/config/fixed_first_user_bootstrap.yaml \
  --domain retail \
  --run-id fixed_first_user_bootstrap_retail \
  --strict-preflight \
  --execute

benchmark/tau2/llm/run_full_eval.sh \
  --config benchmark/tau2/llm/config/fixed_first_user_bootstrap.yaml \
  --domain airline \
  --run-id fixed_first_user_bootstrap_airline \
  --strict-preflight \
  --execute

Then convert each bootstrap results.json into a fixture:

RETAIL_RESULTS=benchmark/tau2/llm/result/fixed_first_user_bootstrap_retail/memory_cells/fixed_first_user_bootstrap_retail_retail_no_memory_r1/fixed_first_user_bootstrap_retail_retail_no_memory_r1.json
AIRLINE_RESULTS=benchmark/tau2/llm/result/fixed_first_user_bootstrap_airline/memory_cells/fixed_first_user_bootstrap_airline_airline_no_memory_r1/fixed_first_user_bootstrap_airline_airline_no_memory_r1.json

python benchmark/tau2/llm/scripts/build_fixed_first_user_fixture.py \
  --repo "$TAU2_REPO" \
  --results-json "$RETAIL_RESULTS" \
  --domain retail \
  --task-split-name test \
  --output benchmark/tau2/llm/result/fixed_first_user_fixtures/retail/fixed_first_user_fixture.json \
  --require-full-split

python benchmark/tau2/llm/scripts/build_fixed_first_user_fixture.py \
  --repo "$TAU2_REPO" \
  --results-json "$AIRLINE_RESULTS" \
  --domain airline \
  --task-split-name test \
  --output benchmark/tau2/llm/result/fixed_first_user_fixtures/airline/fixed_first_user_fixture.json \
  --require-full-split

Export the generated fixture paths for subsequent strict runs:

export TAU2_RETAIL_FIXED_FIRST_USER_FILE="$PWD/benchmark/tau2/llm/result/fixed_first_user_fixtures/retail/fixed_first_user_fixture.json"
export TAU2_AIRLINE_FIXED_FIRST_USER_FILE="$PWD/benchmark/tau2/llm/result/fixed_first_user_fixtures/airline/fixed_first_user_fixture.json"

Run Plans And Smoke Checks

Plan the no-memory baseline without running TAU-2:

python benchmark/tau2/llm/scripts/run_eval.py \
  --config benchmark/tau2/llm/config/no_memory.yaml \
  --plan-only

Plan the paired current-evidence config without running TAU-2:

python benchmark/tau2/llm/scripts/run_eval.py \
  --config benchmark/tau2/llm/config/template_indexed_trajectory.yaml \
  --plan-only

Run a tiny no-memory smoke:

benchmark/tau2/llm/run_full_eval.sh \
  --config benchmark/tau2/llm/config/no_memory.yaml \
  --domain retail \
  --strategy-id no_memory \
  --num-tasks 1 \
  --repeat-count 1 \
  --strict-preflight \
  --execute

Run a tiny template-indexed trajectory smoke against a clean local OpenViking service:

benchmark/tau2/llm/run_full_eval.sh \
  --config benchmark/tau2/llm/config/template_indexed_trajectory.yaml \
  --domain retail \
  --strategy-id template_indexed_trajectory_top4_prewrite_top2 \
  --num-tasks 1 \
  --train-num-tasks 1 \
  --repeat-count 1 \
  --strict-preflight \
  --execute

Start the OpenViking service before executing memory cells, and verify it with ov status. For trajectory memory evidence, start the service from this branch and inspect generated trajectory files; changing search_uri alone does not prove the template-indexed trajectory prompt was used.

Full Reproduction

Run the no-memory full8 baseline:

benchmark/tau2/llm/run_full_eval.sh \
  --config benchmark/tau2/llm/config/no_memory.yaml \
  --run-id no_memory_full8 \
  --strict-preflight \
  --execute

Run the paired no-memory + current trajectory evidence config:

benchmark/tau2/llm/run_full_eval.sh \
  --config benchmark/tau2/llm/config/template_indexed_trajectory.yaml \
  --run-id template_indexed_trajectory_full8 \
  --strict-preflight \
  --execute

The main result is written to benchmark/tau2/llm/result/template_indexed_trajectory_full8/scoreboard.json. Per-cell execution records live under cell_results/, raw TAU-2 result JSON lives under memory_cells/, and corpus identity / generated memory checks live under memory_corpora/.

Memory Adapter

Memory cells run through a small TAU-2 agent adapter in this directory:

  • train by writing TAU-2 training conversations into OpenViking sessions;
  • retrieve OpenViking memory at the first user turn;
  • for pre-write recall, retrieve again before write-like tool calls and regenerate that step with the matched memories;
  • optionally apply a generic scope prompt that keeps retrieved memories advisory and asks the agent to preserve the current task scope before write-like tool calls;
  • emit artifact metadata identifying the OpenViking account, agent, corpus, retrieval mode, search memory type, and simulator policy used by each cell.

The current trajectory config uses:

  • train_memory_mode: experience_only, which selects the Memory V2 session-commit path that writes generated memory artifacts;
  • train_transcript_format: role_tool_blocks, which preserves role-prefixed messages plus tool-call/tool-response blocks during training;
  • train_include_system_prompt: true, which includes the domain policy in the training session;
  • train_skip_failed_sessions: true, which avoids learning from failed train sessions;
  • search_memory_type: trajectories, which retrieves generated trajectory memory during eval.

The runner prepares each distinct domain + corpus_id once and reuses it across eval run ids when the cached corpus_manifest.json is present. Different corpora may be prepared in parallel with benchmark.corpus_prepare_concurrency; session commits inside one corpus remain serial to preserve OpenViking write semantics.

By default, trajectory extraction is transcript-only: the runner replays TAU-2 messages into an OpenViking session and does not expose held-out reward or assertion results to the extractor.

Eval cells run in parallel with benchmark.strategy_concurrency by default and can be overridden with --strategy-concurrency. This only parallelizes read-only TAU-2 eval cells; corpus writes inside one corpus are still serialized by the prepare step.

For exploratory gates, prefer a bounded run with --cell-timeout-seconds. Timed-out cells are recorded with return code 124, timed_out=true, and are excluded from scoreboard metrics, which keeps smoke runs from silently becoming long-running evidence jobs.

User Simulator Policy

The runner default is the official TAU-2 user simulator if eval.user_simulator_policy is omitted. The bundled OpenViking memory benchmark configs set confirmation_aware, because a memory benchmark should not treat user confirmation as task completion before the backend write has happened.

confirmation_aware applies a small idempotent prompt patch to the configured TAU-2 checkout before planning or running. The patch appends only the behavioral confirmation boundary to the TAU-2 user simulator guidelines; metadata such as the upstream PR link is kept in run artifacts, not in the simulator prompt.

Optional fixed-first-user fixtures keep the first simulated user turn stable while preserving live simulator behavior after that turn.

Evidence Boundary

Only completed retail + airline runs with the same config, same seeds/repeats, and non-empty artifacts should be read as benchmark evidence. Partial runs, single-task probes, or missing OpenViking corpus identity are diagnostics. Executed runs write per-cell JSON under cell_results/ and a strategy/domain aggregate under scoreboard.json. Memory training artifacts are shared by domain and strategy under memory_corpora/, so repeated eval cells reuse the same fresh corpus instead of rewriting it.