* session: add restricted Python DSL extraction protocol and make it the default
Introduce a restricted Python memory SDK output protocol as an alternative to
the JSON extraction protocol, and switch the default to python. Both protocols
share the same ResolvedOperations post-processing, schema rules, and patch-repair
path via a new ExtractionOutputProtocol abstraction.
- Add extraction_output_protocol/{base,json,python}.py; python compiles a
restricted AST into the same operations model as json.
- Default memory.extraction_output_format flips json -> python.
- Surface the offending source line on Python syntax errors and add targeted
triple-quote retry guidance for string-literal breaks.
- Preserve every distinct fact on canonical merges; remove hardcoded memory
type names from prompts so custom memory_types render dynamically.
- Downgrade benign batch-delete link-inheritance read failures to WARNING.
- Add memory_organization A/B benchmark and message_format pretty-printer.
Tests: extraction protocol, config loader, memory react suites pass.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: drop redundant entity split hint and duplicate abstractmethod
- entities.yaml: remove the size-triggered split hint; when to split/compact is
decided at read time by memory_maintenance_notice, so the static schema
description only keeps the identity semantics and fact-preservation rule.
- vlm/base.py: remove a duplicated @abstractmethod on get_completion_async.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* vlm: drop redundant *.vlm.call span decorators for trace parity
volcengine already dropped its @tracer("volcengine.vlm.call") wrapper to avoid
duplicate spans now that the request is logged via tracer.info(llm_input_messages=...).
Remove the symmetric litellm/openai decorators so all three backends behave the same.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* benchmark/locomo: fix commit_session kwarg for CLI AsyncHTTPClient
ov.AsyncHTTPClient resolves to openviking_cli.client._http_compat.AsyncHTTPClient,
whose commit_session takes a flat telemetry= kwarg and has no options= parameter.
Passing options={...} (the SDK-client shape) raised TypeError during import.
Use telemetry=True to match the CLI client, consistent with the other locomo
import scripts.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: don't parse error-target sentinels as URIs in extraction telemetry
The by-type extraction telemetry treated result.errors[].uri as a valid viking
URI and fell back to MemoryUpdater.memory_type_from_uri(), but that field is an
error *target* — it can be a sentinel like "unknown" or "events(page_id=100)".
VikingURI() then raised 'URI must start with viking://', turning a single
recorded extraction error into a crash of the whole long_term extraction step.
Count failed errors by the known uri->type map only, defaulting to "unknown".
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* bot: pass context_type via FindOptions after SDK find/search sync
The SDK find/search sync moved context_type from a top-level find() kwarg into
FindOptions. VikingBot still called client.find(context_type='memory') for peer
recall, so every per-turn type-quota recall raised 'unexpected keyword argument
context_type' and silently returned no memories. The answer agent then fell back
to manual multi-round search (iteration ~1.3 -> ~3.9) and accuracy dropped from
~83% to ~72-77%. Pass it via options={'context_type': 'memory'} instead.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* bot: adapt VikingClient.find to SDK FindOptions for context_type/filter
The SDK find/search sync moved context_type and filter out of top-level find()
kwargs into FindOptions. VikingClient.find still forwarded them as top-level
kwargs to the SDK client, so peer memory recall raised 'unexpected keyword
argument context_type' (and after the prior partial fix, 'options') and returned
no memories — the answer agent fell back to manual multi-round search, spiking
iteration ~1.3 -> ~4 and dropping accuracy ~83% -> ~76%.
Do the SDK adaptation once in VikingClient.find (pack context_type/filter into
options={...}); callers keep the stable VikingClient.find(context_type=...)
interface, so memory.py reverts to passing context_type= directly. Verified via
a single-question smoke: type_quota recall returns 13 memories, injection is
non-empty, iteration=1, answer correct.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: route event resolution repair through the output protocol
The event resolution-repair instruction was hardcoded to demand a JSON object,
but under the default Python protocol the repaired response is parsed by the
Python SDK compiler. When a first-pass event had out-of-bounds ranges, an
assistant-only span, or an ambiguous peer, the repair round returned JSON, the
compiler rejected it as an invalid program, retries were exhausted, and the
recoverable event memory was never written.
Add ExtractionOutputProtocol.render_resolution_repair(); JSON keeps the existing
JSON-object wording, Python asks for corrected sdk.create_events(...) calls.
_build_resolution_repair_instruction now delegates to the active protocol, like
patch-repair and the final instruction already do.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session,bot,benchmark: address PR review findings
- extract_loop: document 32768 as the extraction output floor (tuned for Doubao;
lower-max models override via vlm.max_tokens) and extract
_resolve_effective_max_output_tokens; ov.conf.example notes the override.
- python_protocol: alias non-identifier memory_type/field names on the Python DSL
surface only (real names kept in URIs/storage/JSON); map aliases back when
compiling, instead of hard-rejecting kebab-case custom schemas.
- run_full_eval.sh: move auto-commit + GIT_COMMIT_ID capture AFTER arg parsing so
--auto-commit is honored and run metadata records the committed HEAD.
- litellm_vlm: strip Gemini cache_control from the already-sanitized messages so
empty assistant turns are not reintroduced; sanitize_openai_messages passes
through non-dict entries.
- Tests for each fix.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: reject Python DSL alias collisions instead of silently overwriting
_identifier_alias() is not one-to-one: memory_type 'project-notes' and
'project_notes' (or fields 'note-body'/'note_body') fold to the same DSL alias.
The alias->real dict comprehensions would silently drop one, making a schema/
field unreachable and routing writes to the wrong target. Add
_validate_alias_uniqueness(), invoked in render_contract() and the compiler
__init__ (so parse() paths without render are also guarded), which fails loudly
with a rename hint. Distinct-identifier names never collide, so real configs are
unaffected. Tests cover type, field, and no-render parse collisions.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: bound str.replace result size before allocation in Python DSL
The _MAX_EXPRESSION_SIZE guard covered * (repeat) and + (concat) but not the
whitelisted string methods: only join() had a projected-size check, so
('x'*1000).replace('x','y'*10000) could still allocate a >1MB result and bypass
the limit. Add _check_replace_size() that bounds source + occurrences*(len(new)
-len(old)) BEFORE calling str.replace (which builds the whole result in C), so
the oversized string is never allocated. replace is the only whitelisted method
that can materially inflate output (join already guarded; upper/lower/strip/
split/startswith/endswith do not grow). Tests cover an oversized replace being
rejected and a normal replace passing.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: close f-string width and str %-format inflation in Python DSL
The replace guard alone was insufficient: f-string format specs (f"{'x':>1000001}")
and str %-formatting ("%1000001s" % "x") also turn a small integer literal into an
arbitrarily large string with no repeat operator, bypassing _MAX_EXPRESSION_SIZE.
Neither has a legitimate use in memory content, so disallow them outright rather
than bounding width inflation: reject any non-empty f-string format spec and reject
str/bytes %-formatting (numeric % still allowed). Combined with the existing
*/+/join/replace pre-allocation checks, all small-input->large-output amplifiers
are now closed. Tests cover f-string width rejection, plain f-string, and str %.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat(sdk): sync go/ts/python SDKs with server find/search, recall, and admin changes
Server-side changes recently landed that the language SDKs had drifted from:
- find/search results now return `tags` and no longer return
`category`/`match_reason`/`relations`/`overview` (#3730). Go's strict
struct was the only one broken; update MatchedContext accordingly.
- new admin endpoints for agent-evolution and per-account settings (#3695).
- public `search/recall` endpoint was missing from all SDKs.
Changes:
- python: add `level`/`since`/`until`/`time_field` to find/search; add an
`extra` escape hatch to find/search/add_resource/write/batch_write so new
server fields can be passed without an SDK bump (only forwarded when set,
preserving `level=0`); add `recall` and the four admin methods.
- go: fix MatchedContext (add Tags, drop removed fields), add Recall and the
four admin methods.
- typescript: type MatchedContext/FindResult, add RecallOptions, add `recall`
and the four admin methods.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(sdk): unify options APIs and sync latest server interfaces
- migrate complex Python SDK calls to typed options dictionaries
- add dedicated context search and consistent extra-field handling
- align Go and TypeScript options with omission-aware serialization
- support session config, event tags, Agent Evolution date filters,
OpenViking Assets, batch write, downloads, and create_parent
- refresh SDK tests and examples across all three languages
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): address options API review findings
- fix Go session extra merging and Python message precedence
- adapt LangChain calls to the Python options API
- migrate repository examples, tests, and documentation
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): complete options migration and message parity
- migrate remaining Python SDK benchmarks to options dictionaries
- normalize empty parts consistently for single and batch messages
- add regression guards for repository SDK call sites
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): align reindex options after main rebase
- preserve reindex tags in Python typed options
- add reindex extra support for Go and TypeScript
- reject official fields passed through extra across SDKs
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(sdk): support legacy keyword options
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
docs(sdk): use explicit Python SDK arguments
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): support set tags extra options
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): expose Go add resource options
Expose AddType and ProcessingMode through Go AddResourceOptions and serialize them to the resources API. Add a regression test covering the resulting request payload.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(sdk): flatten core Python client options
Co-authored-by: TRAE CLI <traecli@bytedance.com>
docs(sdk): align Python examples with flattened options
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): preserve core API compatibility
Co-authored-by: TRAE CLI <traecli@bytedance.com>
refactor(python-sdk): move resource hints to options
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): align resource option callers
Co-authored-by: TRAE CLI <traecli@bytedance.com>
test(sdk): cover recursive reindex forwarding
Co-authored-by: TRAE CLI <traecli@bytedance.com>
fix(sdk): preserve Go options compatibility
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(python-sdk): expose message peer id
Co-authored-by: TRAE CLI <traecli@bytedance.com>
test(python-sdk): consolidate options coverage
Co-authored-by: TRAE CLI <traecli@bytedance.com>
feat(python-sdk): add parts and flatten image search
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* docs(sdk): align Python call examples
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
* feat(uri)!: reject uid-less current-user shorthand in favor of viking://~
viking://user/<segment> (memories/resources/skills/peers/privacy/sessions
without a user id) was ambiguous with a user literally named after the
segment, and a user actually named e.g. "memories" was unreachable for
USER/ADMIN callers. Now that the viking://~ home alias (#4167) covers the
same need unambiguously, the shorthand fails closed at the request
boundary instead of expanding:
- resolve_current_user_uri raises NamespaceShapeError with a corrective
hint naming both viking://~/<rest> and the explicit-uid form. Silently
parsing the reserved segment as a peer user id would misdirect reads
and writes, so rejection is the only safe removal.
- Bare viking://user falls through to the canonical parser and keeps
container semantics (a user key listing it sees only its own space).
- The self-id escape stays: a caller whose user_id equals a reserved
name keeps viking://user/<own-id> as their canonical root. ROOT-role
literal parsing and the legacy viking://session alias are unchanged.
- AddTargetsConfig normalizes stored legacy config spellings
(viking://user/resources|skills) to the viking://~ form at validation
so existing ov.conf/user_config deployments keep working; the accepted
per-user spelling is now viking://~/resources and viking://~/skills.
- usage_reporter keeps canonicalizing the historical shorthand found in
old transcripts and additionally recognizes viking://~/memories/.
BREAKING CHANGE: requests using the uid-less viking://user/<segment>
spelling now fail with 400; use viking://~/<segment> or an explicit
viking://user/{user_id}/<segment> URI.
* refactor(clients): migrate first-party emitters to the viking://~ home alias
Every in-repo client that emitted the removed uid-less current-user
shorthand now sends viking://~/... instead: vikingbot fallbacks and
default sentinels, the LangChain store/tools defaults, the shared
recall-core.mjs (all synced plugin copies), the codex/claude-code/
openclaw/openwebui/dsh/zcode/pi plugin emitters, quick-app examples,
Go SDK example, tau2 benchmark targets, and the eval golden dataset.
Compat kept where legacy strings live in stored user configs: bot and
ov_dream sentinels accept both spellings while emitting only ~, and
recall-core still rewrites legacy viking://user/<reserved> config values
client-side. langchain_openviking._uri now classifies viking://~ with
the explicit-user shape so canonicalized server responses keep matching
a ~ root. Plugin READMEs note the server requirement for the alias.
* docs: replace current-user shorthand guidance with the viking://~ home alias
Rewrite every EN/ZH doc and model-facing prompt that advertised the
uid-less viking://user/<segment> spelling: URI concept catalogue,
context-types/storage/extraction/retrieval/session/privacy concepts,
configuration guide (with the legacy add_targets auto-normalization
note), resources/skills/sessions/retrieval/admin API references, FAQ,
capability reference, and the openviking-memory / ov-experience-memory /
openclaw / ov-resources skills. The stale MCP viking://user/<path>
dialect passage in the MCP guide is replaced by ~ guidance, and bare
viking://user is documented as the container of user spaces.
* test(api): migrate live API session-used tests off the removed shorthand
tests/api_test/sessions sent uid-less viking://user/skills/... URIs to
record_used, which the request boundary now rejects with 400 (caught by
the API & CLI Integration Tests CI job; these tests need a live server
and are not part of the local suites). The api_test client authenticates
as an admin-role user key, so the viking://~ home alias expands for it.
tests/api_test/common/test_edge_cases.py is left as is: it asserts a 400
for a non-resource add target, which still holds.
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest
Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.
- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit
Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.
* fix(vectordb): reject stale index replacements
* fix(vectordb): harden bulk rebuild lifecycle
---------
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
* feat(storage): optimize glob func
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* feat(rgafs): implement paged glob traversal without full tree materialization
* fix(localfs): offload blocking fs operations to spawn_blocking
* feat(glob): cap glob api default node_limit at 256
* feat(sdk): add node_limit options for glob in python and go SDKs
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* feat(tau2/vikingbot): config-driven experience recall + per-domain isolation
Switch tau2 self-improvement behaviour from core-code patches to three
ov.conf flags (recall_exp_first_round_only, exp_recall_limit,
exp_recall_max_chars), so the VikingBot core is unchanged for non-tau2
users.
- context.py: when recall_exp_first_round_only=true, skip per-turn
user+agent memory retrieval and inject experience once on the first
user-turn; accepts explicit agent_id to scope retrieval per domain
- memory.py: exp_recall_limit and exp_recall_max_chars read from config
instead of hardcoded values
- schema.py: add three new OpenVikingConfig fields (all default to
existing behaviour so existing deployments are unaffected)
- ov_server.py: extract _is_session_key() helper to unify the two
places that distinguish session keys from per-domain agent ids;
local mode now respects agent_id for namespace isolation (remote mode
already supported this)
- tau2 runner: pass agent_id= instead of memory_users= to build_messages
- README: document Python >=3.12 prerequisite, correct pip install
extra, clarify that isolation works in both local and remote modes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ov_server): remove redundant underscore check in search_experiences
The "_" in self.agent_id guard was a leftover heuristic to distinguish
domain ids from session keys. Now that _is_session_key() handles that
check via "__", the extra "_" condition is unnecessary and actively
breaks agent ids without underscores (e.g. "airline", "retail").
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(benchmark/tau2): add VikingBot agent runner for tau2-bench
Adds benchmark/tau2/vikingbot/, an end-to-end harness that runs the full
VikingBot AgentLoop on tau2-bench tasks and commits trajectories back into
OpenViking memory for epoch-based self-improvement. This complements the
existing memory-retrieval harness in benchmark/tau2/ (which is retrieval-only).
Contents:
- scripts/vikingbot_tau2_runner.py: run one tau2 task through the agent loop
(tau2 tool registry swap, simulated-time patch, advisory memory scope guard).
- scripts/run_tau2_domain.sh / run_eval_reward.sh: run a domain split with
bounded concurrency and score average reward.
- scripts/commit_trajectory_to_memory.py: commit train trajectories to memory.
- scripts/stat_trajectory.py, check_openviking_tool_calls.py: analysis helpers.
- tau2_env/: tau2 environment + tool-provider integration.
- run_full_test.sh and run_{airline,retail}_*epochs.sh: full / multi-epoch runs.
- setup_env.sh, README.md, .gitignore.
tau2-bench is referenced as an external dependency (cloned + installed by the
user); no OpenViking core changes are required. The runner is API-compatible
with bot/vikingbot on current main.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* refactor(benchmark/tau2): split into llm/ and vikingbot/ subfolders
Mirror the two evaluation approaches as sibling subfolders under benchmark/tau2/:
- llm/: the existing OpenViking Memory V2 retrieval harness, moved from
benchmark/tau2/. All internal benchmark/tau2/... path references and the
REPO_ROOT depth computations (run_full_eval.sh, tau2_common.py,
run_memory_v2_eval.py) are updated for the extra directory level.
- vikingbot/: the VikingBot agent runner (added in the previous commit).
vikingbot/ cleanup:
- make memory-block extraction time-independent: anchor on the stable session
header and trailing reply instruction instead of a fixed simulated timestamp
(the sim-time patch was removed, so the current time is now system-generated).
- drop the now-removed sim-time / scope-guard notes from the README.
- remove the unused stat_trajectory.py and check_openviking_tool_calls.py helpers.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* refactor(benchmark/tau2/vikingbot): train-once/test-8x eval, drop smolagents, doc updates
- run_full_test.sh: run train once per epoch (experience extraction) and test
N times in parallel (--test-repeats, default 8), reporting the averaged test
accuracy; keep --commit/--no-commit.
- tau2_environment.py: remove the unused smolagents Tool path (CommunicateWithUser /
create_tool_from_json_schema / self.tools); communicate_with_user is handled
directly in tool_call. tau2-bench has no smolagents dependency, so it is dropped.
- README: reorder install (tau2-bench first so setup_env can derive TAU2_DATA_ROOT),
explain train-once/test-8x methodology and train-only memory extraction, document
the required bot/vikingbot core changes (agent_id isolation + agent-experience
memory), fix sibling links to ../llm/.
- Remove run_retail_3epochs.sh.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(tau2/vikingbot): one-step setup_env.sh + communicate_with_user refactor
setup_env.sh now does full environment setup in a single `source`: creates a
fresh repo-root .venv, clones tau2-bench (external dep), installs openviking +
vikingbot (pip install -e ., runs the Cargo build) + tau2-bench + smolagents,
then activates and exports the runtime env vars. Idempotent via a marker file;
supports --reinstall. README updated to document the one-step flow and the
overridable env vars.
Also move the communicate_with_user tool into a CommunicateWithUser class in
tau2_environment.py (owns both schema and execution) and drop the duplicated
inline schema from tau2_tool_provider.py.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs(tau2/vikingbot): sync setup_env.sh fixes + README port/diff clarifications
Backport the environment-setup fixes and README clarifications discovered while
running the harness end-to-end (the core bot/vikingbot code changes live on the
test/tau2-vikingbot-core-changes branch, not here):
- setup_env.sh: install the [bot] extra (prompt_toolkit/gradio/mcp/...), build +
bundle ragfs_python via maturin when the editable install skips it under pip
build isolation, and install tau2-bench with the [gym] extra (gymnasium)
- README.md: explain the server port (default 1933 vs bot.ov_server.server_url)
and show the None-safe forms of the Change-1 diffs
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* clean README message
---------
Co-authored-by: ByteDance <wenting.qi@bytedance.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* feat(benchmark): add Claude Code LoCoMo evaluation
Add locomo benchmark for Claude Code, following the same pipeline as
openclaw/vikingbot: ingest → QA → judge → stat.
- ingest.py: send each session to `claude -p`, let auto-memory extract
- eval.py: run QA questions in isolated project dirs with auto-memory
- judge.py / stat_judge_result.py: reuse same grading logic
- run_full_eval.sh: orchestrate full pipeline
Supports custom API endpoint (ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN),
environment isolation via HOME override, and retry on timeout/error.
* feat(benchmark): add prompt-prefix support and ingest-phase statistics
- ingest.py: add --prompt-prefix flag to nudge auto-memory during ingest
- run_full_eval.sh: add --project-root, --home, --result-dir, --prompt-prefix
flags for flexible environment isolation (bare vs prompted experiments)
- stat_judge_result.py: report ingest-phase token/cost/latency alongside
QA statistics, with --ingest-csv auto-detection
* feat(benchmark): add OpenViking integration for LoCoMo eval
- eval.py: add --hooks-settings, --mcp-config, --ov-config for OV recall
during QA (sets OPENVIKING env vars, injects OV hint in prompt)
- run_full_eval.sh: wire --input, --hooks-settings, --mcp-config,
--ov-ingest-config, --ov-qa-config params through to eval/ingest
- Add OV config files: ov-hooks.json, ov-mcp.json, ov-ingest.conf, ov-qa.conf
- Add import_to_ov.py for direct LoCoMo → OpenViking data import via SDK
* feat(benchmark): update configuration files and enhance evaluation logging
* feat(benchmark): add configuration files for Claude Code LoCoMo evaluation and enhance eval script
* feat(benchmark): reorganize LoCoMo claude-code eval around 4 reproducible modes
Collapse months of rN iteration scripts into 4 entry points covering
the four published numbers:
- run_prompted.sh - vanilla CC auto-memory (no OV)
- run_sdk_iso.sh - openviking SDK pre-ingest, per-sample namespace
- run_sdk_noiso.sh - openviking SDK pre-ingest, shared namespace
- run_e2e.sh - claude -p stream-json multi-turn + auto-capture
All four go through the same eval.py/judge.py/stat_judge_result.py
pipeline. Hook + plugin lookup is parametrised via OPENVIKING_PLUGIN_DIR
so the configs are portable. Historical scripts/reports/intermediate
configs are moved under .tmp/legacy (gitignored) for reference.
* style(benchmark): apply ruff format + auto-fix F541 across LoCoMo claude-code suite
- ruff format on eval.py / ingest.py / judge.py / stat_judge_result.py
- drop extraneous f-prefixes flagged by F541 in stat_judge_result.py
and eval.py