* Add trajectory experience learning redesign doc * auto-commit before eval 20260607_043406 * auto-commit before eval 20260607_044129 * auto-commit before eval 20260607_123706 * auto-commit before eval 20260607_125514 * auto-commit before eval 20260607_133737 * auto-commit before eval 20260607_144649 * auto-commit before eval 20260607_154631 * Refine streaming memory train merge pipeline * Refine session train policy optimization architecture * Add VikingMem ARA paper analysis * Force merge for mixed extraction memory patches * auto-commit before eval 20260608_134426 * auto-commit before eval 20260608_142108 * auto-commit before eval 20260608_153909 * auto-commit before eval 20260608_154845 * auto-commit before eval 20260608_170143 * update * auto-commit before eval 20260611_150946 * auto-commit before eval 20260611_153933 * auto-commit before eval 20260611_154251 * Fix tau2 reward wrapper call * auto-commit before eval 20260611_193803 * auto-commit before eval 20260611_194939 * update * auto-commit before eval 20260612_111029 * auto-commit before eval 20260612_112104 * auto-commit before eval 20260612_122603 * auto-commit before eval 20260612_123359 * auto-commit before eval 20260612_124303 * auto-commit before eval 20260612_130257 * Fallback peer routing to first conversation peer * Route self memory through self peer sentinel * Keep self sentinel out of peer memory paths * auto-commit before eval 20260612_154051 * auto-commit before eval 20260612_154850 * auto-commit before eval 20260612_161633 * auto-commit before eval 20260612_184022 * auto-commit before eval 20260612_201845 * auto-commit before eval 20260612_202637 * auto-commit before eval 20260612_204040 * auto-commit before eval 20260612_224621 * Fix locomo progress column initialization * Add memory field versioning * auto-commit before eval 20260612_232318 * Simplify locomo progress display * Remove locomo progress elapsed time * Batch streaming memory merges by group * Derive patch merge language from patches * Detect patch merge language from updated files * auto-commit before eval 20260613_004339 * auto-commit before eval 20260613_005835 * Persist memory update trace id * auto-commit before eval 20260613_012722 * auto-commit before eval 20260613_013923 * auto-commit before eval 20260613_014708 * Enforce peer scope after memory merge * auto-commit before eval 20260613_033402 * auto-commit before eval 20260613_151931 * auto-commit before eval 20260613_164217 * chore: raise vikingbot eval parallelism * chore: tune vikingbot parallelism to 150 * auto-commit before eval 20260613_185807 * chore: restore vikingbot parallelism default * feat(locomo): add import progress reporting * chore(memory): restore profile and preference templates * Fix tau2 reward JSON serialization * Refactor tau2 batch memory training * Stream batch train JSONL events * Add fast path for batch training case specs * Optimize streaming train gradient chunking * Optimize patch merge prompt context * fix tau2 memory training vectorization * fix(memory): revert profile preference granularity rules * bd init: initialize beads issue tracking * update * Log memory template fallback failures * Record all rollout artifacts * Fix OpenViking peer search forwarding * Stop tracking Beads local state * auto-commit before eval 20260616_002037 * Deprecate memory version selector * Retry transient LoCoMo import HTTP failures * Add memory schema stage and peer routing * Organize LoCoMo benchmark outputs * Restore VikingBot user memory auto recall * Show elapsed time on LoCoMo progress bars * Quiet transient import retries * Shorten LoCoMo progress bars * Route non-peer memories to self scope * auto-commit before eval 20260616_124513 * Suppress memory read not found logs * Limit LoCoMo import memory types * Rename peer routing schema flag * Rename peer schema flag to enable_peer * Rename schema peer flag to peer_enabled * auto-commit before eval 20260616_135946 * auto-commit before eval 20260616_140641 * auto-commit before eval 20260616_141753 * Show cached baseline eval at start of training * Preserve remote policy contents * Show failed work in progress bars * Hide zero failed progress counts * Disable tau2 service progress by default * Reuse policy lock for policy deletes * feat: add session skill extraction to Memory V3 streaming trainer - Generalize domain types: Experience → Policy, ExperienceSet → PolicySet - Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type - Generalize PatchSemanticGradient target names - Add SkillSetLoader (reads skills/ dir into PolicySet) - Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater) - Add RolloutAnalysis.gradients for co-extracted policy patches - Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients - Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission - Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases() - Generalize PatchMergePolicyOptimizer for any memory_type - Update tests to use new field/kind names Co-authored-by: Claude <noreply@anthropic.com> * Persist experience reminders in tau2 rollouts * Enable tau2 epoch test eval by default * Persist train rollout artifacts incrementally * Ensure tau2 vikingbot user simulator deps * Auto repair tau2 vikingbot simulator deps * Avoid blocking tau2 vikingbot service loop * Avoid tau2 gym reset when loading cases * Clean tau2 rollout commit messages * Clean tau2 tool trajectory serialization * Retry vikingbot VLM rate limits * Refine tau2 training case selection * Promote vikingbot hook execution log level * Improve VLM rate limit retry detection * Update trajectory analysis prompt format * Limit tau2 service logs to warnings * Run tau2 vikingbot rollouts on service loop * Lower vikingbot experience recall threshold * Offload tau2 vikingbot blocking setup * Retry tau2 LiteLLM rate limits * Pin trajectory and experience outputs to Chinese * Retry tau2 rate limits indefinitely * Highlight tau2 training accuracy summaries * Hide redundant avg reward console metrics * Tighten memory extraction templates * Reduce tau2 memory template noise Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service. Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%. * Constrain tau2 memory extraction sources Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories. Evaluation: - Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval - Result dir: result/tau2/train/airline_20260619_000757 - Baseline test: 55.00% (88/160) - Epoch0 train: 66.67% (20/30) - Epoch0 test: 56.25% (90/160) - Epoch1 train: 60.00% (18/30) - Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%. * Preserve tau2 train non-run results * Improve memory extraction guardrails Run: result/tau2/train/run_airline_20260619_044051 tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp. * Support train split eval in tau2 batch runs * Add slot support to tau2 vikingbot launcher * Copy OpenViking configs for tau2 slots * Tune tau2 case1 memory extraction Run: result/tau2/train_1/run_airline_20260619_201546 Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp. * Advise tau2 train case1 best result Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%. * Tune tau2 memory gate extraction * Advise tau2 train case1 50pct result * Guard failed write experience branches * Advise tau2 train case1 100pct result * Guard tau2 oracle training memories * Recall trajectory diagnostics for tau2 rollouts * Recall tau2 case specs for training rollouts * Guard evaluated tau2 final states * Inject compact tau2 oracle checklists * Stabilize tau2 slot train multi-case runs * Guard tau2 case10 oracle terminal state * Use supported tau2 training memory types * Match tau2 oracle writes by expected subset * Autofill tau2 case10 oracle writes before done * Enable tau2 case10 guard for train split * Record slot1 S008 case10 guard best advice * Generalize tau2 S008 oracle terminal guard * Record slot1 S008 general guard best advice * Remove tau2 benchmark oracle guard * Prevent training ground truth memory recall * Refine tau2 training memory extraction * Fix epoch train rollout artifact stage * Refine memory training rollout pipeline * update * auto-commit before eval 20260623_120317 * fix sdk read_raw for memory metadata * use visible case links for experience recall * auto-commit before eval 20260623_225354 * tau2/train: cap run_batch_train_eval rollout concurrency at 100 * update * update * update * fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2 - Port _same_memory_file filter to compressor_v3._build_memory_diff so no-op merges/patches don't inflate memory_diff.json update counts - Write memory_diff.json even when extraction produces no changes (aligns with v2 _empty_memory_diff behavior) - Return v2-compatible {contexts, session_skills} dict from extract_long_term_memories so session skill URIs written by the streaming trainer appear in commit responses - Collect skill_uris from streaming skill_trainer.submit_gradients apply_result - Remove four dead skill-related imports left from the unbuilt v3 execution-memory path - Fix lock_manager caller to handle both list and dict return shapes - Fix test_session_commit assertions that assumed v2-only extract_execution_memories method exists * fix(memory,v3): also filter unchanged experience updates in training memory diff * train: finish rollout and memory refactor * memory: refine runtime-visible extraction prompts * train: constrain communication memory extraction * auto-commit before eval 20260629_235623 * memory: address training review fixes * update * update * message: reuse part deserializer * train: snapshot memory prompt yaml * prompts: restore memory yaml templates from main * memory: scope streaming update results * update * update * session: train canonical merged cases --------- Co-authored-by: Claude <noreply@anthropic.com>
TAU-2 Benchmark
This directory contains the OpenViking TAU-2 LLM benchmark entry point. The reproduction surface is intentionally narrow:
no_memory: same-seed TAU-2 baseline without OpenViking memory injection;template_indexed_trajectory_top4_prewrite_top2: the current best template-indexed trajectory memory treatment.
The template-indexed trajectory treatment trains OpenViking Memory V2 from
TAU-2 train conversations, retrieves generated trajectories, and uses the
trajectory embedding template {{ trajectory_name }}\n\n{{ retrieval_anchor }}
instead of broad procedure bodies for retrieval. It injects trajectory top4 at
the first user turn and top2 before write-like tool calls, with the generic
memory scope prompt enabled.
Category rerank, experience-memory routes, fixed-count-only ablations, character-budget ablations, and official-user parity controls are intentionally left out of this README and config set so reproduction agents do not mistake diagnostic routes for current evidence.
Layout
benchmark/tau2/llm/
├── config/
│ ├── baseline.yaml
│ ├── fixed_first_user_bootstrap.yaml
│ ├── no_memory.yaml
│ ├── scope_prompts/
│ │ └── generic_memory_scope.md
│ └── template_indexed_trajectory.yaml
├── scripts/
│ ├── build_fixed_first_user_fixture.py
│ ├── run_eval.py
│ ├── setup_tau2_repo.sh
│ └── tau2_common.py
└── run_full_eval.sh
baseline.yaml is a shared protocol/defaults file, not a runnable evidence
cell by itself. Use no_memory.yaml for the baseline-only run and
template_indexed_trajectory.yaml for the paired no-memory + trajectory run.
Generated eval artifacts are written to benchmark/tau2/llm/result/<run_id>/.
Memory corpus artifacts are cached outside the run id at
benchmark/tau2/llm/result/memory_corpora/ by default.
Setup
This benchmark delegates task simulation and scoring to an external TAU-2 checkout. Point the runner at that checkout and CLI explicitly when they are not on the default path:
export TAU2_REPO=/path/to/tau2-bench
export TAU2_CLI=/path/to/tau2
For a local one-command setup, clone and install TAU-2 into ignored benchmark directories:
benchmark/tau2/llm/scripts/setup_tau2_repo.sh
source benchmark/tau2/llm/.env.tau2
The default OpenViking TAU-2 memory evidence protocol is
fixed_first_user_full8: retail + airline, 8 repeats, same seeds,
confirmation-aware user simulator, and fixed first-user fixtures for both
domains. Later user simulator turns remain live.
The confirmation-aware simulator behavior is available from sierra-research/tau2-bench#297. Pin the local TAU-2 checkout to a ref that includes that behavior when reproducing these numbers:
benchmark/tau2/llm/scripts/setup_tau2_repo.sh \
--ref refs/pull/297/head
source benchmark/tau2/llm/.env.tau2
When using Doubao through an OpenAI-compatible endpoint, set OPENAI_API_KEY
and OPENAI_API_BASE for LiteLLM before running upstream TAU-2.
Fixed-First-User Fixtures
Strict reproduction requires fixed first-user fixtures:
export TAU2_RETAIL_FIXED_FIRST_USER_FILE=/path/to/retail/fixed_first_user_fixture.json
export TAU2_AIRLINE_FIXED_FIRST_USER_FILE=/path/to/airline/fixed_first_user_fixture.json
--strict-preflight fails when eval.require_fixed_first_user=true and either
fixture is missing.
For a fresh checkout, run one live-user bootstrap pass per domain:
benchmark/tau2/llm/run_full_eval.sh \
--config benchmark/tau2/llm/config/fixed_first_user_bootstrap.yaml \
--domain retail \
--run-id fixed_first_user_bootstrap_retail \
--strict-preflight \
--execute
benchmark/tau2/llm/run_full_eval.sh \
--config benchmark/tau2/llm/config/fixed_first_user_bootstrap.yaml \
--domain airline \
--run-id fixed_first_user_bootstrap_airline \
--strict-preflight \
--execute
Then convert each bootstrap results.json into a fixture:
RETAIL_RESULTS=benchmark/tau2/llm/result/fixed_first_user_bootstrap_retail/memory_cells/fixed_first_user_bootstrap_retail_retail_no_memory_r1/fixed_first_user_bootstrap_retail_retail_no_memory_r1.json
AIRLINE_RESULTS=benchmark/tau2/llm/result/fixed_first_user_bootstrap_airline/memory_cells/fixed_first_user_bootstrap_airline_airline_no_memory_r1/fixed_first_user_bootstrap_airline_airline_no_memory_r1.json
python benchmark/tau2/llm/scripts/build_fixed_first_user_fixture.py \
--repo "$TAU2_REPO" \
--results-json "$RETAIL_RESULTS" \
--domain retail \
--task-split-name test \
--output benchmark/tau2/llm/result/fixed_first_user_fixtures/retail/fixed_first_user_fixture.json \
--require-full-split
python benchmark/tau2/llm/scripts/build_fixed_first_user_fixture.py \
--repo "$TAU2_REPO" \
--results-json "$AIRLINE_RESULTS" \
--domain airline \
--task-split-name test \
--output benchmark/tau2/llm/result/fixed_first_user_fixtures/airline/fixed_first_user_fixture.json \
--require-full-split
Export the generated fixture paths for subsequent strict runs:
export TAU2_RETAIL_FIXED_FIRST_USER_FILE="$PWD/benchmark/tau2/llm/result/fixed_first_user_fixtures/retail/fixed_first_user_fixture.json"
export TAU2_AIRLINE_FIXED_FIRST_USER_FILE="$PWD/benchmark/tau2/llm/result/fixed_first_user_fixtures/airline/fixed_first_user_fixture.json"
Run Plans And Smoke Checks
Plan the no-memory baseline without running TAU-2:
python benchmark/tau2/llm/scripts/run_eval.py \
--config benchmark/tau2/llm/config/no_memory.yaml \
--plan-only
Plan the paired current-evidence config without running TAU-2:
python benchmark/tau2/llm/scripts/run_eval.py \
--config benchmark/tau2/llm/config/template_indexed_trajectory.yaml \
--plan-only
Run a tiny no-memory smoke:
benchmark/tau2/llm/run_full_eval.sh \
--config benchmark/tau2/llm/config/no_memory.yaml \
--domain retail \
--strategy-id no_memory \
--num-tasks 1 \
--repeat-count 1 \
--strict-preflight \
--execute
Run a tiny template-indexed trajectory smoke against a clean local OpenViking service:
benchmark/tau2/llm/run_full_eval.sh \
--config benchmark/tau2/llm/config/template_indexed_trajectory.yaml \
--domain retail \
--strategy-id template_indexed_trajectory_top4_prewrite_top2 \
--num-tasks 1 \
--train-num-tasks 1 \
--repeat-count 1 \
--strict-preflight \
--execute
Start the OpenViking service before executing memory cells, and verify it with
ov status. For trajectory memory evidence, start the service from this branch
and inspect generated trajectory files; changing search_uri alone does not
prove the template-indexed trajectory prompt was used.
Full Reproduction
Run the no-memory full8 baseline:
benchmark/tau2/llm/run_full_eval.sh \
--config benchmark/tau2/llm/config/no_memory.yaml \
--run-id no_memory_full8 \
--strict-preflight \
--execute
Run the paired no-memory + current trajectory evidence config:
benchmark/tau2/llm/run_full_eval.sh \
--config benchmark/tau2/llm/config/template_indexed_trajectory.yaml \
--run-id template_indexed_trajectory_full8 \
--strict-preflight \
--execute
The main result is written to
benchmark/tau2/llm/result/template_indexed_trajectory_full8/scoreboard.json.
Per-cell execution records live under cell_results/, raw TAU-2 result JSON
lives under memory_cells/, and corpus identity / generated memory checks live
under memory_corpora/.
Memory Adapter
Memory cells run through a small TAU-2 agent adapter in this directory:
- train by writing TAU-2 training conversations into OpenViking sessions;
- retrieve OpenViking memory at the first user turn;
- for pre-write recall, retrieve again before write-like tool calls and regenerate that step with the matched memories;
- optionally apply a generic scope prompt that keeps retrieved memories advisory and asks the agent to preserve the current task scope before write-like tool calls;
- emit artifact metadata identifying the OpenViking account, agent, corpus, retrieval mode, search memory type, and simulator policy used by each cell.
The current trajectory config uses:
train_memory_mode: experience_only, which selects the Memory V2 session-commit path that writes generated memory artifacts;train_transcript_format: role_tool_blocks, which preserves role-prefixed messages plus tool-call/tool-response blocks during training;train_include_system_prompt: true, which includes the domain policy in the training session;train_skip_failed_sessions: true, which avoids learning from failed train sessions;search_memory_type: trajectories, which retrieves generated trajectory memory during eval.
The runner prepares each distinct domain + corpus_id once and reuses it across
eval run ids when the cached corpus_manifest.json is present. Different
corpora may be prepared in parallel with benchmark.corpus_prepare_concurrency;
session commits inside one corpus remain serial to preserve OpenViking write
semantics.
By default, trajectory extraction is transcript-only: the runner replays TAU-2 messages into an OpenViking session and does not expose held-out reward or assertion results to the extractor.
Eval cells run in parallel with benchmark.strategy_concurrency by default and
can be overridden with --strategy-concurrency. This only parallelizes read-only
TAU-2 eval cells; corpus writes inside one corpus are still serialized by the
prepare step.
For exploratory gates, prefer a bounded run with --cell-timeout-seconds.
Timed-out cells are recorded with return code 124, timed_out=true, and are
excluded from scoreboard metrics, which keeps smoke runs from silently becoming
long-running evidence jobs.
User Simulator Policy
The runner default is the official TAU-2 user simulator if
eval.user_simulator_policy is omitted. The bundled OpenViking memory benchmark
configs set confirmation_aware, because a memory benchmark should not treat
user confirmation as task completion before the backend write has happened.
confirmation_aware applies a small idempotent prompt patch to the configured
TAU-2 checkout before planning or running. The patch appends only the behavioral
confirmation boundary to the TAU-2 user simulator guidelines; metadata such as
the upstream PR link is kept in run artifacts, not in the simulator prompt.
Optional fixed-first-user fixtures keep the first simulated user turn stable while preserving live simulator behavior after that turn.
Evidence Boundary
Only completed retail + airline runs with the same config, same seeds/repeats,
and non-empty artifacts should be read as benchmark evidence. Partial runs,
single-task probes, or missing OpenViking corpus identity are diagnostics.
Executed runs write per-cell JSON under cell_results/ and a strategy/domain
aggregate under scoreboard.json. Memory training artifacts are shared by
domain and strategy under memory_corpora/, so repeated eval cells reuse the
same fresh corpus instead of rewriting it.