Commit Graph
72 Commits
Author SHA1 Message Date
chenjwandTRAE CLI a843ab6bf2 session: add restricted Python DSL extraction protocol and make it the default (#4581)
* session: add restricted Python DSL extraction protocol and make it the default

Introduce a restricted Python memory SDK output protocol as an alternative to
the JSON extraction protocol, and switch the default to python. Both protocols
share the same ResolvedOperations post-processing, schema rules, and patch-repair
path via a new ExtractionOutputProtocol abstraction.

- Add extraction_output_protocol/{base,json,python}.py; python compiles a
  restricted AST into the same operations model as json.
- Default memory.extraction_output_format flips json -> python.
- Surface the offending source line on Python syntax errors and add targeted
  triple-quote retry guidance for string-literal breaks.
- Preserve every distinct fact on canonical merges; remove hardcoded memory
  type names from prompts so custom memory_types render dynamically.
- Downgrade benign batch-delete link-inheritance read failures to WARNING.
- Add memory_organization A/B benchmark and message_format pretty-printer.

Tests: extraction protocol, config loader, memory react suites pass.
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: drop redundant entity split hint and duplicate abstractmethod

- entities.yaml: remove the size-triggered split hint; when to split/compact is
  decided at read time by memory_maintenance_notice, so the static schema
  description only keeps the identity semantics and fact-preservation rule.
- vlm/base.py: remove a duplicated @abstractmethod on get_completion_async.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* vlm: drop redundant *.vlm.call span decorators for trace parity

volcengine already dropped its @tracer("volcengine.vlm.call") wrapper to avoid
duplicate spans now that the request is logged via tracer.info(llm_input_messages=...).
Remove the symmetric litellm/openai decorators so all three backends behave the same.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* benchmark/locomo: fix commit_session kwarg for CLI AsyncHTTPClient

ov.AsyncHTTPClient resolves to openviking_cli.client._http_compat.AsyncHTTPClient,
whose commit_session takes a flat telemetry= kwarg and has no options= parameter.
Passing options={...} (the SDK-client shape) raised TypeError during import.
Use telemetry=True to match the CLI client, consistent with the other locomo
import scripts.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: don't parse error-target sentinels as URIs in extraction telemetry

The by-type extraction telemetry treated result.errors[].uri as a valid viking
URI and fell back to MemoryUpdater.memory_type_from_uri(), but that field is an
error *target* — it can be a sentinel like "unknown" or "events(page_id=100)".
VikingURI() then raised 'URI must start with viking://', turning a single
recorded extraction error into a crash of the whole long_term extraction step.
Count failed errors by the known uri->type map only, defaulting to "unknown".

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* bot: pass context_type via FindOptions after SDK find/search sync

The SDK find/search sync moved context_type from a top-level find() kwarg into
FindOptions. VikingBot still called client.find(context_type='memory') for peer
recall, so every per-turn type-quota recall raised 'unexpected keyword argument
context_type' and silently returned no memories. The answer agent then fell back
to manual multi-round search (iteration ~1.3 -> ~3.9) and accuracy dropped from
~83% to ~72-77%. Pass it via options={'context_type': 'memory'} instead.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* bot: adapt VikingClient.find to SDK FindOptions for context_type/filter

The SDK find/search sync moved context_type and filter out of top-level find()
kwargs into FindOptions. VikingClient.find still forwarded them as top-level
kwargs to the SDK client, so peer memory recall raised 'unexpected keyword
argument context_type' (and after the prior partial fix, 'options') and returned
no memories — the answer agent fell back to manual multi-round search, spiking
iteration ~1.3 -> ~4 and dropping accuracy ~83% -> ~76%.

Do the SDK adaptation once in VikingClient.find (pack context_type/filter into
options={...}); callers keep the stable VikingClient.find(context_type=...)
interface, so memory.py reverts to passing context_type= directly. Verified via
a single-question smoke: type_quota recall returns 13 memories, injection is
non-empty, iteration=1, answer correct.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: route event resolution repair through the output protocol

The event resolution-repair instruction was hardcoded to demand a JSON object,
but under the default Python protocol the repaired response is parsed by the
Python SDK compiler. When a first-pass event had out-of-bounds ranges, an
assistant-only span, or an ambiguous peer, the repair round returned JSON, the
compiler rejected it as an invalid program, retries were exhausted, and the
recoverable event memory was never written.

Add ExtractionOutputProtocol.render_resolution_repair(); JSON keeps the existing
JSON-object wording, Python asks for corrected sdk.create_events(...) calls.
_build_resolution_repair_instruction now delegates to the active protocol, like
patch-repair and the final instruction already do.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session,bot,benchmark: address PR review findings

- extract_loop: document 32768 as the extraction output floor (tuned for Doubao;
  lower-max models override via vlm.max_tokens) and extract
  _resolve_effective_max_output_tokens; ov.conf.example notes the override.
- python_protocol: alias non-identifier memory_type/field names on the Python DSL
  surface only (real names kept in URIs/storage/JSON); map aliases back when
  compiling, instead of hard-rejecting kebab-case custom schemas.
- run_full_eval.sh: move auto-commit + GIT_COMMIT_ID capture AFTER arg parsing so
  --auto-commit is honored and run metadata records the committed HEAD.
- litellm_vlm: strip Gemini cache_control from the already-sanitized messages so
  empty assistant turns are not reintroduced; sanitize_openai_messages passes
  through non-dict entries.
- Tests for each fix.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: reject Python DSL alias collisions instead of silently overwriting

_identifier_alias() is not one-to-one: memory_type 'project-notes' and
'project_notes' (or fields 'note-body'/'note_body') fold to the same DSL alias.
The alias->real dict comprehensions would silently drop one, making a schema/
field unreachable and routing writes to the wrong target. Add
_validate_alias_uniqueness(), invoked in render_contract() and the compiler
__init__ (so parse() paths without render are also guarded), which fails loudly
with a rename hint. Distinct-identifier names never collide, so real configs are
unaffected. Tests cover type, field, and no-render parse collisions.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: bound str.replace result size before allocation in Python DSL

The _MAX_EXPRESSION_SIZE guard covered * (repeat) and + (concat) but not the
whitelisted string methods: only join() had a projected-size check, so
('x'*1000).replace('x','y'*10000) could still allocate a >1MB result and bypass
the limit. Add _check_replace_size() that bounds source + occurrences*(len(new)
-len(old)) BEFORE calling str.replace (which builds the whole result in C), so
the oversized string is never allocated. replace is the only whitelisted method
that can materially inflate output (join already guarded; upper/lower/strip/
split/startswith/endswith do not grow). Tests cover an oversized replace being
rejected and a normal replace passing.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: close f-string width and str %-format inflation in Python DSL

The replace guard alone was insufficient: f-string format specs (f"{'x':>1000001}")
and str %-formatting ("%1000001s" % "x") also turn a small integer literal into an
arbitrarily large string with no repeat operator, bypassing _MAX_EXPRESSION_SIZE.
Neither has a legitimate use in memory content, so disallow them outright rather
than bounding width inflation: reject any non-empty f-string format spec and reject
str/bytes %-formatting (numeric % still allowed). Combined with the existing
*/+/join/replace pre-allocation checks, all small-input->large-output amplifiers
are now closed. Tests cover f-string width rejection, plain f-string, and str %.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-07 20:49:02 +08:00
DuTao 0e754112ad 删除无用的bot逻辑 (#4769) 2026-09-07 17:35:42 +08:00
9eac8a6d3d feat(sdk): sync go/ts/python SDKs with server find/search, recall, an… (#3737)
* feat(sdk): sync go/ts/python SDKs with server find/search, recall, and admin changes

Server-side changes recently landed that the language SDKs had drifted from:

- find/search results now return `tags` and no longer return
  `category`/`match_reason`/`relations`/`overview` (#3730). Go's strict
  struct was the only one broken; update MatchedContext accordingly.
- new admin endpoints for agent-evolution and per-account settings (#3695).
- public `search/recall` endpoint was missing from all SDKs.

Changes:
- python: add `level`/`since`/`until`/`time_field` to find/search; add an
  `extra` escape hatch to find/search/add_resource/write/batch_write so new
  server fields can be passed without an SDK bump (only forwarded when set,
  preserving `level=0`); add `recall` and the four admin methods.
- go: fix MatchedContext (add Tags, drop removed fields), add Recall and the
  four admin methods.
- typescript: type MatchedContext/FindResult, add RecallOptions, add `recall`
  and the four admin methods.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): unify options APIs and sync latest server interfaces

- migrate complex Python SDK calls to typed options dictionaries
- add dedicated context search and consistent extra-field handling
- align Go and TypeScript options with omission-aware serialization
- support session config, event tags, Agent Evolution date filters,
  OpenViking Assets, batch write, downloads, and create_parent
- refresh SDK tests and examples across all three languages

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): address options API review findings

- fix Go session extra merging and Python message precedence
- adapt LangChain calls to the Python options API
- migrate repository examples, tests, and documentation

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): complete options migration and message parity

- migrate remaining Python SDK benchmarks to options dictionaries
- normalize empty parts consistently for single and batch messages
- add regression guards for repository SDK call sites

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): align reindex options after main rebase

- preserve reindex tags in Python typed options
- add reindex extra support for Go and TypeScript
- reject official fields passed through extra across SDKs

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): support legacy keyword options

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

docs(sdk): use explicit Python SDK arguments

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): support set tags extra options

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): expose Go add resource options

Expose AddType and ProcessingMode through Go AddResourceOptions and serialize them to the resources API. Add a regression test covering the resulting request payload.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): flatten core Python client options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

docs(sdk): align Python examples with flattened options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): preserve core API compatibility

Co-authored-by: TRAE CLI <traecli@bytedance.com>

refactor(python-sdk): move resource hints to options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): align resource option callers

Co-authored-by: TRAE CLI <traecli@bytedance.com>

test(sdk): cover recursive reindex forwarding

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): preserve Go options compatibility

Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(python-sdk): expose message peer id

Co-authored-by: TRAE CLI <traecli@bytedance.com>

test(python-sdk): consolidate options coverage

Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(python-sdk): add parts and flatten image search

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* docs(sdk): align Python call examples

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-24 14:09:11 +08:00
t0saki a83b81715b feat(uri)!: remove uid-less current-user shorthand in favor of viking://~ (#4196)
* feat(uri)!: reject uid-less current-user shorthand in favor of viking://~

viking://user/<segment> (memories/resources/skills/peers/privacy/sessions
without a user id) was ambiguous with a user literally named after the
segment, and a user actually named e.g. "memories" was unreachable for
USER/ADMIN callers. Now that the viking://~ home alias (#4167) covers the
same need unambiguously, the shorthand fails closed at the request
boundary instead of expanding:

- resolve_current_user_uri raises NamespaceShapeError with a corrective
  hint naming both viking://~/<rest> and the explicit-uid form. Silently
  parsing the reserved segment as a peer user id would misdirect reads
  and writes, so rejection is the only safe removal.
- Bare viking://user falls through to the canonical parser and keeps
  container semantics (a user key listing it sees only its own space).
- The self-id escape stays: a caller whose user_id equals a reserved
  name keeps viking://user/<own-id> as their canonical root. ROOT-role
  literal parsing and the legacy viking://session alias are unchanged.
- AddTargetsConfig normalizes stored legacy config spellings
  (viking://user/resources|skills) to the viking://~ form at validation
  so existing ov.conf/user_config deployments keep working; the accepted
  per-user spelling is now viking://~/resources and viking://~/skills.
- usage_reporter keeps canonicalizing the historical shorthand found in
  old transcripts and additionally recognizes viking://~/memories/.

BREAKING CHANGE: requests using the uid-less viking://user/<segment>
spelling now fail with 400; use viking://~/<segment> or an explicit
viking://user/{user_id}/<segment> URI.

* refactor(clients): migrate first-party emitters to the viking://~ home alias

Every in-repo client that emitted the removed uid-less current-user
shorthand now sends viking://~/... instead: vikingbot fallbacks and
default sentinels, the LangChain store/tools defaults, the shared
recall-core.mjs (all synced plugin copies), the codex/claude-code/
openclaw/openwebui/dsh/zcode/pi plugin emitters, quick-app examples,
Go SDK example, tau2 benchmark targets, and the eval golden dataset.

Compat kept where legacy strings live in stored user configs: bot and
ov_dream sentinels accept both spellings while emitting only ~, and
recall-core still rewrites legacy viking://user/<reserved> config values
client-side. langchain_openviking._uri now classifies viking://~ with
the explicit-user shape so canonicalized server responses keep matching
a ~ root. Plugin READMEs note the server requirement for the alias.

* docs: replace current-user shorthand guidance with the viking://~ home alias

Rewrite every EN/ZH doc and model-facing prompt that advertised the
uid-less viking://user/<segment> spelling: URI concept catalogue,
context-types/storage/extraction/retrieval/session/privacy concepts,
configuration guide (with the legacy add_targets auto-normalization
note), resources/skills/sessions/retrieval/admin API references, FAQ,
capability reference, and the openviking-memory / ov-experience-memory /
openclaw / ov-resources skills. The stale MCP viking://user/<path>
dialect passage in the MCP guide is replaced by ~ guidance, and bare
viking://user is documented as the container of user spaces.

* test(api): migrate live API session-used tests off the removed shorthand

tests/api_test/sessions sent uid-less viking://user/skills/... URIs to
record_used, which the request boundary now rejects with 400 (caught by
the API & CLI Integration Tests CI job; these tests need a live server
and are not part of the local suites). The api_test client authenticates
as an admin-role user key, so the viking://~ home alias expands for it.
tests/api_test/common/test_edge_cases.py is left as is: it asserts a 400
for a non-resource add target, which still holds.
2026-08-21 19:00:19 +08:00
Qin Haojie 7abd6ab249 refactor(client): remove Python embedded mode (#3712)
* refactor(client): remove Python embedded mode

Consolidate Python consumers on the HTTP SDK while keeping shared server and storage capabilities unchanged.

* refactor(client): remove obsolete embedded leftovers
2026-08-10 18:00:00 +08:00
Yuanqing ZHAOandYuanqing Zhao 40dd05271c perf(vectordb): micro-batch compatible cuVS searches (#3382)
* perf(vectordb): micro-batch compatible cuVS searches

* fix(vectordb): serialize micro-batch device admission

* perf(vectordb): pipeline warm cuVS micro-batch admission

* docs(cuvs): align micro-batching guidance

* fix(cuvs): warm-batch empty filters

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-23 11:10:52 +08:00
DuTao 55a9d12cd6 1. 优化评测参数化; (#3332)
2. 优化评测显示;
3. 修复gpt-5.6 api返回 无 choices时bot兼容问题。
2026-07-17 17:46:57 +08:00
Yuanqing ZHAOandYuanqing Zhao fa19ac0a75 perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest (#3277)
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest

Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.

- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit

Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.

* fix(vectordb): reject stale index replacements

* fix(vectordb): harden bulk rebuild lifecycle

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-16 18:59:16 +08:00
Yuanqing ZHAOandYuanqing Zhao 546da35cc8 fix(benchmark): load WIKI-Dir path mapping (#3279)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-16 11:11:15 +08:00
Yuanqing ZHAOandYuanqing Zhao 0080f94bdc perf(vectordb): batch benchmark upserts (#3264)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-15 19:49:06 +08:00
fujiajie666 4847ffa378 模板优化 (#3242) 2026-07-15 11:37:47 +08:00
Yuanqing ZHAOandYuanqing Zhao 7e6a0515f9 perf(cuvs): optimize filters, rebuilds, concurrency, and memory (#3092)
* perf(cuvs): fast-path cached native filter routes

* perf(cuvs): parallelize auto filter preflight

* perf(cuvs): add search route telemetry

* test(cuvs): use a valid telemetry vector dimension

* perf(cuvs): reuse native filter preflight results

* perf(cuvs): allow concurrent snapshot searches

* perf(cuvs): coalesce optional background rebuilds

* perf(cuvs): coordinate per-GPU build admission

* perf(cuvs): add opt-in float16 search

* build(cuvs): support vector benchmark harnesses

* perf(cuvs): bound concurrent GPU searches

* perf(cuvs): avoid partial background rebuilds

* fix(cuvs): address rebuild and telemetry review feedback

* fix(cuvs): defer rebuild until index initialization

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-10 17:22:34 +08:00
Yuanqing ZHAOandYuanqing Zhao d61d802fd0 fix(benchmark): enforce configured search concurrency (#3091)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-09 11:06:58 +08:00
Qin Haojie 2eb61fbbad feat(benchmark): 增加 OV 目录向量检索性能 benchmark (#3076)
补充基于 VikingVectorIndexBackend 的 dir-vector/synthetic benchmark,并修复本地 bitmap 读路径隐藏写入导致的并发检索崩溃。
2026-07-08 14:49:57 +08:00
Yuanqing ZHAOandYuanqing Zhao 39c778c953 feat: add cuVS vector search backend (#2974)
* feat: add cuVS vector search backend

* docs: add agent memory benchmark strategy

* bench: add cuVS index performance harness

* bench: add public ANN dataset tuning

* docs: record preliminary cuVS index results

* docs: clarify warm index latency

* docs: order cuVS before qdrant

* bench: aggregate independent index runs

* bench: order aggregate variants consistently

* docs: add repeatable index scaling results

* bench: add collection lifecycle benchmark

* docs: add collection lifecycle results

* perf: cache prepared cuvs filters

* docs: report prepared filter cache results

* bench: add async vector concurrency benchmark

* bench: aggregate service concurrency runs

* docs: add async concurrency results

* docs: clarify cuVS dtype behavior

* feat: add memory-aware cuVS auto mode

* feat: reuse native filters for cuVS search

* docs: publish cuVS integration plan as Markdown

* fix: route selective filters before cuVS rebuild

* docs: record selective-first routing results

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-07 12:21:10 +08:00
baojun-zhang 6a33ebb7ca Optimize glob walkdir (#3013)
* feat(storage): optimize glob func

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* feat(rgafs): implement paged glob traversal without full tree materialization

* fix(localfs): offload blocking fs operations to spawn_blocking

* feat(glob): cap glob api default node_limit at 256

* feat(sdk): add node_limit options for glob in python and go SDKs
2026-07-06 21:39:16 +08:00
chenjwandClaude fd73dcf23a Feat/自进化(经验记忆)框架重构 (#2503)
* Add trajectory experience learning redesign doc

* auto-commit before eval 20260607_043406

* auto-commit before eval 20260607_044129

* auto-commit before eval 20260607_123706

* auto-commit before eval 20260607_125514

* auto-commit before eval 20260607_133737

* auto-commit before eval 20260607_144649

* auto-commit before eval 20260607_154631

* Refine streaming memory train merge pipeline

* Refine session train policy optimization architecture

* Add VikingMem ARA paper analysis

* Force merge for mixed extraction memory patches

* auto-commit before eval 20260608_134426

* auto-commit before eval 20260608_142108

* auto-commit before eval 20260608_153909

* auto-commit before eval 20260608_154845

* auto-commit before eval 20260608_170143

* update

* auto-commit before eval 20260611_150946

* auto-commit before eval 20260611_153933

* auto-commit before eval 20260611_154251

* Fix tau2 reward wrapper call

* auto-commit before eval 20260611_193803

* auto-commit before eval 20260611_194939

* update

* auto-commit before eval 20260612_111029

* auto-commit before eval 20260612_112104

* auto-commit before eval 20260612_122603

* auto-commit before eval 20260612_123359

* auto-commit before eval 20260612_124303

* auto-commit before eval 20260612_130257

* Fallback peer routing to first conversation peer

* Route self memory through self peer sentinel

* Keep self sentinel out of peer memory paths

* auto-commit before eval 20260612_154051

* auto-commit before eval 20260612_154850

* auto-commit before eval 20260612_161633

* auto-commit before eval 20260612_184022

* auto-commit before eval 20260612_201845

* auto-commit before eval 20260612_202637

* auto-commit before eval 20260612_204040

* auto-commit before eval 20260612_224621

* Fix locomo progress column initialization

* Add memory field versioning

* auto-commit before eval 20260612_232318

* Simplify locomo progress display

* Remove locomo progress elapsed time

* Batch streaming memory merges by group

* Derive patch merge language from patches

* Detect patch merge language from updated files

* auto-commit before eval 20260613_004339

* auto-commit before eval 20260613_005835

* Persist memory update trace id

* auto-commit before eval 20260613_012722

* auto-commit before eval 20260613_013923

* auto-commit before eval 20260613_014708

* Enforce peer scope after memory merge

* auto-commit before eval 20260613_033402

* auto-commit before eval 20260613_151931

* auto-commit before eval 20260613_164217

* chore: raise vikingbot eval parallelism

* chore: tune vikingbot parallelism to 150

* auto-commit before eval 20260613_185807

* chore: restore vikingbot parallelism default

* feat(locomo): add import progress reporting

* chore(memory): restore profile and preference templates

* Fix tau2 reward JSON serialization

* Refactor tau2 batch memory training

* Stream batch train JSONL events

* Add fast path for batch training case specs

* Optimize streaming train gradient chunking

* Optimize patch merge prompt context

* fix tau2 memory training vectorization

* fix(memory): revert profile preference granularity rules

* bd init: initialize beads issue tracking

* update

* Log memory template fallback failures

* Record all rollout artifacts

* Fix OpenViking peer search forwarding

* Stop tracking Beads local state

* auto-commit before eval 20260616_002037

* Deprecate memory version selector

* Retry transient LoCoMo import HTTP failures

* Add memory schema stage and peer routing

* Organize LoCoMo benchmark outputs

* Restore VikingBot user memory auto recall

* Show elapsed time on LoCoMo progress bars

* Quiet transient import retries

* Shorten LoCoMo progress bars

* Route non-peer memories to self scope

* auto-commit before eval 20260616_124513

* Suppress memory read not found logs

* Limit LoCoMo import memory types

* Rename peer routing schema flag

* Rename peer schema flag to enable_peer

* Rename schema peer flag to peer_enabled

* auto-commit before eval 20260616_135946

* auto-commit before eval 20260616_140641

* auto-commit before eval 20260616_141753

* Show cached baseline eval at start of training

* Preserve remote policy contents

* Show failed work in progress bars

* Hide zero failed progress counts

* Disable tau2 service progress by default

* Reuse policy lock for policy deletes

* feat: add session skill extraction to Memory V3 streaming trainer

- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names

Co-authored-by: Claude <noreply@anthropic.com>

* Persist experience reminders in tau2 rollouts

* Enable tau2 epoch test eval by default

* Persist train rollout artifacts incrementally

* Ensure tau2 vikingbot user simulator deps

* Auto repair tau2 vikingbot simulator deps

* Avoid blocking tau2 vikingbot service loop

* Avoid tau2 gym reset when loading cases

* Clean tau2 rollout commit messages

* Clean tau2 tool trajectory serialization

* Retry vikingbot VLM rate limits

* Refine tau2 training case selection

* Promote vikingbot hook execution log level

* Improve VLM rate limit retry detection

* Update trajectory analysis prompt format

* Limit tau2 service logs to warnings

* Run tau2 vikingbot rollouts on service loop

* Lower vikingbot experience recall threshold

* Offload tau2 vikingbot blocking setup

* Retry tau2 LiteLLM rate limits

* Pin trajectory and experience outputs to Chinese

* Retry tau2 rate limits indefinitely

* Highlight tau2 training accuracy summaries

* Hide redundant avg reward console metrics

* Tighten memory extraction templates

* Reduce tau2 memory template noise

Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.

Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.

* Constrain tau2 memory extraction sources

Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.

Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.

* Preserve tau2 train non-run results

* Improve memory extraction guardrails

Run: result/tau2/train/run_airline_20260619_044051

tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.

* Support train split eval in tau2 batch runs

* Add slot support to tau2 vikingbot launcher

* Copy OpenViking configs for tau2 slots

* Tune tau2 case1 memory extraction

Run: result/tau2/train_1/run_airline_20260619_201546

Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.

* Advise tau2 train case1 best result

Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.

* Tune tau2 memory gate extraction

* Advise tau2 train case1 50pct result

* Guard failed write experience branches

* Advise tau2 train case1 100pct result

* Guard tau2 oracle training memories

* Recall trajectory diagnostics for tau2 rollouts

* Recall tau2 case specs for training rollouts

* Guard evaluated tau2 final states

* Inject compact tau2 oracle checklists

* Stabilize tau2 slot train multi-case runs

* Guard tau2 case10 oracle terminal state

* Use supported tau2 training memory types

* Match tau2 oracle writes by expected subset

* Autofill tau2 case10 oracle writes before done

* Enable tau2 case10 guard for train split

* Record slot1 S008 case10 guard best advice

* Generalize tau2 S008 oracle terminal guard

* Record slot1 S008 general guard best advice

* Remove tau2 benchmark oracle guard

* Prevent training ground truth memory recall

* Refine tau2 training memory extraction

* Fix epoch train rollout artifact stage

* Refine memory training rollout pipeline

* update

* auto-commit before eval 20260623_120317

* fix sdk read_raw for memory metadata

* use visible case links for experience recall

* auto-commit before eval 20260623_225354

* tau2/train: cap run_batch_train_eval rollout concurrency at 100

* update

* update

* update

* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2

- Port _same_memory_file filter to compressor_v3._build_memory_diff so
  no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
  (aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
  extract_long_term_memories so session skill URIs written by the
  streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
  apply_result
- Remove four dead skill-related imports left from the unbuilt v3
  execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
  extract_execution_memories method exists

* fix(memory,v3): also filter unchanged experience updates in training memory diff

* train: finish rollout and memory refactor

* memory: refine runtime-visible extraction prompts

* train: constrain communication memory extraction

* auto-commit before eval 20260629_235623

* memory: address training review fixes

* update

* update

* message: reuse part deserializer

* train: snapshot memory prompt yaml

* prompts: restore memory yaml templates from main

* memory: scope streaming update results

* update

* update

* session: train canonical merged cases

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-03 11:27:17 +08:00
fujiajie666 79cb571074 locomo数据导入优化 (#2852) 2026-06-26 16:00:24 +08:00
87329714dd feat(grep): integrate VikingDB bm25 keyword search for grep engine (#2144)
* feat(grep): integrate VikingDB bm25 keyword search for grep engine

* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)

* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison

* fix(schema): upsert data to vikingdb lack of content

* chore: add benchmark for retrieval

* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs

* fix(benchmark): sub uri args; add report

* refactor: code format by ruff

* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf

* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search

* fix: adjust benchmark scripts

* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls

* refactor: new benchmark

* fix: step1 add resource by real code data

* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex

* optimize (benchmark): adjust keywords and ground truth for testing

* fix: truncate 64KB for content field

* optimize: effectiveness add resource plainly

* optimize: change param use of SearchByKeywords from "keywords" to "query"

* optimize(benchmark): refactor effectiveness scripts

* optimize: ensure raw data for content field

* optimize: fulltext analyzer's stop-words only use symbols

* fix: adapt to new ov cli for benchmark

* optimize: reuse file content to avoid re-read AGFS file

* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts

* optimize: benchmark client timeout

* update README

* fix: rm unused param

* fix: default values in docs

* optimize: increase truncate byte size to 1MB for content field for VikingDB

* fix(logger): harden queued stream logging (#2786)

* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock

When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.

During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.

Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.

Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.

Closes: #2752

* fix(logger): harden queued stream logging

---------

Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
2026-06-24 18:46:02 +08:00
agent a4aefac1f7 feat(session): Support image message extraction (#2578)
* Support image message extraction

* fix: fix image url

* fix: bug

* fix: image parts readme
2026-06-15 18:03:27 +08:00
DuTao 07326bd827 feat(eval):Opt memory eval script (#2563)
* api_key

* 兼容最新的peer逻辑

* 调整 peer 逻辑

* 工具检索self + peer memory,以及对应的resource、SKILLS

* 调整评测,兼容peer逻辑

* 调整loop中的profile获取

* 1. 恢复agent workspace;
2. 评测使用ov_server.url;
3. 调整文档

* 1. 修复 root_api_key兼容;
2. 调整 Peer 别名

* 调整评测逻辑

* PR review

* test

* fix md

* fix md

* Eval 逻辑优化;vlm 增加token统计;
2026-06-11 21:23:09 +08:00
agent 8a3f12d174 feat(exp): add LongMemEval and LoCoMo OpenViking benchmarks (#1937)
* feat: add longmemeval

* feat: longmemeval

* feat: openviking in longmemeval

* feat: run eval

* fix

* feat: add openviking in locomo and longmem eval

* feat: remove unless expr code

* fix: unless code

* feat: locomo and longmemeval

* feat: remove openviking

* feat: locomo

* fix: id

* feat: locomo

* fix: locomo

* fix: bug

* feat: prompt 同步

* fix: ov exp

* revert viking bot

* fix: template

* fix

* fix: remove file

* fix: eval

* feat: model

* feat: ov import and judge

* feat: exp readme

* fix: eval

* feat: long message split

* fix: agent id

* fix: doc

* fix: test case

* fix

* fix: readme

* fix

* chore: split memory chunking into separate branch

* chore: lint openviking benchmark scripts
2026-06-11 11:46:24 +08:00
DuTao a702d38a8b feat(bot): Change bot api_key to user mode, support ov's peers, eval support peers (#2527)
* api_key

* 兼容最新的peer逻辑

* 调整 peer 逻辑

* 工具检索self + peer memory,以及对应的resource、SKILLS

* 调整评测,兼容peer逻辑

* 调整loop中的profile获取

* 1. 恢复agent workspace;
2. 评测使用ov_server.url;
3. 调整文档

* 1. 修复 root_api_key兼容;
2. 调整 Peer 别名

* 调整评测逻辑

* PR review

* test

* fix md

* fix md
2026-06-10 15:40:10 +08:00
Qin Haojie a6fc0424bc fix(session): apply memory type policy whitelist (#2530)
* fix(session): apply memory type policy whitelist

Restore top-level memory_types filtering for session memory extraction and validate it against enabled registry schemas. Ensure initialization and peer-aware smoke coverage honor the whitelist.

* fix(session): scope session skills to execution memory policy

* refactor(session): remove per-commit memory policy
2026-06-10 14:54:24 +08:00
DuTao 1e8833b533 bad case doc (#2536) 2026-06-10 11:47:02 +08:00
chenjw 738cee7395 Fix/peer fix (#2469)
* auto-commit before eval 20260605_110036

(cherry picked from commit a4741cd60f0ea689b4e65156eb943b76a41cf2ba)

* auto-commit before eval 20260605_154023

(cherry picked from commit 3791a21c8cf88ae3fdabef12cfb99760c7bbbe5f)

* auto-commit before eval 20260605_174235

(cherry picked from commit 323c75b697369db736ff6ce0a071a14458ccc034)

* fix(vikingbot): preserve legacy memory search compatibility

* refactor(vikingbot): restore legacy memory parameter names

* fix(user-dirs): lazily create user subdirectories

* update
2026-06-08 11:03:13 +08:00
Qin Haojie ff258768c2 feat(memory): 引入 User/Peer 记忆隔离模型 (#2236)
* feat(memory): introduce user and peer memory isolation

Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.

* feat(memory): align session identity around peer IDs

* feat(search): pass peer id through retrieval

* refactor(memory): remove agent identity from integrations

* fix(memory): isolate peer identity from self extraction

* fix(tau2): provision benchmark user configs

* fix(auth): allow admin keys to access data APIs

* fix(openclaw): enable peer memory policy for peer roles

* fix(openclaw): resolve sender for peer recall

* refactor(session): simplify memory extraction routing

* refactor(ov-cli): reduce formatting-only diff

* refactor(message): remove unused message helpers

* refactor(retrieval): simplify peer target resolution

* refactor(namespace): remove deprecated agent namespace policy

* fix(agent): propagate peer id through integrations

* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
yangxinxin-7andClaude Sonnet 4.6 936624c24b feat(tau2/vikingbot): config-driven experience recall + per-domain isolation (#2380)
* feat(tau2/vikingbot): config-driven experience recall + per-domain isolation

Switch tau2 self-improvement behaviour from core-code patches to three
ov.conf flags (recall_exp_first_round_only, exp_recall_limit,
exp_recall_max_chars), so the VikingBot core is unchanged for non-tau2
users.

- context.py: when recall_exp_first_round_only=true, skip per-turn
  user+agent memory retrieval and inject experience once on the first
  user-turn; accepts explicit agent_id to scope retrieval per domain
- memory.py: exp_recall_limit and exp_recall_max_chars read from config
  instead of hardcoded values
- schema.py: add three new OpenVikingConfig fields (all default to
  existing behaviour so existing deployments are unaffected)
- ov_server.py: extract _is_session_key() helper to unify the two
  places that distinguish session keys from per-domain agent ids;
  local mode now respects agent_id for namespace isolation (remote mode
  already supported this)
- tau2 runner: pass agent_id= instead of memory_users= to build_messages
- README: document Python >=3.12 prerequisite, correct pip install
  extra, clarify that isolation works in both local and remote modes

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ov_server): remove redundant underscore check in search_experiences

The "_" in self.agent_id guard was a leftover heuristic to distinguish
domain ids from session keys. Now that _is_session_key() handles that
check via "__", the extra "_" condition is unnecessary and actively
breaks agent ids without underscores (e.g. "airline", "retail").

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-02 15:04:40 +08:00
DuTao be1e7fc482 feat(eval)Opt vikingbot eval script (#2305)
* 优化评测逻辑

* 兼容 飞书的卡片消息
2026-05-29 20:01:36 +08:00
huangruiteng 646ff735db docs: simplify tau2 benchmark reproduction (#2267) 2026-05-27 20:18:49 +08:00
huangruiteng 76c8f559e1 feat(memory): add trajectory retrieval anchor (#2255) 2026-05-27 14:37:40 +08:00
e0ce670f5c feat(tau2/vikingbot): benchmark updates (#2244)
* feat(benchmark/tau2): add VikingBot agent runner for tau2-bench

Adds benchmark/tau2/vikingbot/, an end-to-end harness that runs the full
VikingBot AgentLoop on tau2-bench tasks and commits trajectories back into
OpenViking memory for epoch-based self-improvement. This complements the
existing memory-retrieval harness in benchmark/tau2/ (which is retrieval-only).

Contents:
- scripts/vikingbot_tau2_runner.py: run one tau2 task through the agent loop
  (tau2 tool registry swap, simulated-time patch, advisory memory scope guard).
- scripts/run_tau2_domain.sh / run_eval_reward.sh: run a domain split with
  bounded concurrency and score average reward.
- scripts/commit_trajectory_to_memory.py: commit train trajectories to memory.
- scripts/stat_trajectory.py, check_openviking_tool_calls.py: analysis helpers.
- tau2_env/: tau2 environment + tool-provider integration.
- run_full_test.sh and run_{airline,retail}_*epochs.sh: full / multi-epoch runs.
- setup_env.sh, README.md, .gitignore.

tau2-bench is referenced as an external dependency (cloned + installed by the
user); no OpenViking core changes are required. The runner is API-compatible
with bot/vikingbot on current main.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(benchmark/tau2): split into llm/ and vikingbot/ subfolders

Mirror the two evaluation approaches as sibling subfolders under benchmark/tau2/:

- llm/: the existing OpenViking Memory V2 retrieval harness, moved from
  benchmark/tau2/. All internal benchmark/tau2/... path references and the
  REPO_ROOT depth computations (run_full_eval.sh, tau2_common.py,
  run_memory_v2_eval.py) are updated for the extra directory level.
- vikingbot/: the VikingBot agent runner (added in the previous commit).

vikingbot/ cleanup:
- make memory-block extraction time-independent: anchor on the stable session
  header and trailing reply instruction instead of a fixed simulated timestamp
  (the sim-time patch was removed, so the current time is now system-generated).
- drop the now-removed sim-time / scope-guard notes from the README.
- remove the unused stat_trajectory.py and check_openviking_tool_calls.py helpers.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(benchmark/tau2/vikingbot): train-once/test-8x eval, drop smolagents, doc updates

- run_full_test.sh: run train once per epoch (experience extraction) and test
  N times in parallel (--test-repeats, default 8), reporting the averaged test
  accuracy; keep --commit/--no-commit.
- tau2_environment.py: remove the unused smolagents Tool path (CommunicateWithUser /
  create_tool_from_json_schema / self.tools); communicate_with_user is handled
  directly in tool_call. tau2-bench has no smolagents dependency, so it is dropped.
- README: reorder install (tau2-bench first so setup_env can derive TAU2_DATA_ROOT),
  explain train-once/test-8x methodology and train-only memory extraction, document
  the required bot/vikingbot core changes (agent_id isolation + agent-experience
  memory), fix sibling links to ../llm/.
- Remove run_retail_3epochs.sh.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(tau2/vikingbot): one-step setup_env.sh + communicate_with_user refactor

setup_env.sh now does full environment setup in a single `source`: creates a
fresh repo-root .venv, clones tau2-bench (external dep), installs openviking +
vikingbot (pip install -e ., runs the Cargo build) + tau2-bench + smolagents,
then activates and exports the runtime env vars. Idempotent via a marker file;
supports --reinstall. README updated to document the one-step flow and the
overridable env vars.

Also move the communicate_with_user tool into a CommunicateWithUser class in
tau2_environment.py (owns both schema and execution) and drop the duplicated
inline schema from tau2_tool_provider.py.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(tau2/vikingbot): sync setup_env.sh fixes + README port/diff clarifications

Backport the environment-setup fixes and README clarifications discovered while
running the harness end-to-end (the core bot/vikingbot code changes live on the
test/tau2-vikingbot-core-changes branch, not here):

- setup_env.sh: install the [bot] extra (prompt_toolkit/gradio/mcp/...), build +
  bundle ragfs_python via maturin when the editable install skips it under pip
  build isolation, and install tau2-bench with the [gym] extra (gymnasium)
- README.md: explain the server port (default 1933 vs bot.ov_server.server_url)
  and show the None-safe forms of the Change-1 diffs

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* clean README message

---------

Co-authored-by: ByteDance <wenting.qi@bytedance.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-26 19:35:31 +08:00
huangruiteng 9fffdf4a52 docs(tau2): clarify fixed-user PR-B reproduction (#2223) 2026-05-25 20:50:05 +08:00
Hao Zhe 665b7d980f feat(benchmark): add Hermes OpenViking LoCoMo scripts (#1985)
* feat(benchmark): add Hermes OpenViking LoCoMo scripts

* feat(benchmark): refine Hermes LoCoMo reporting

* fix(benchmark): harden Hermes LoCoMo stats accounting

Sum OpenViking token CSV rows for resumed runs and document the Hermes LoCoMo benchmark workflow.
2026-05-22 16:14:14 +08:00
fujiajie666 98cc00b657 change switch name of template (#2185)
* 修正开关命名

* 修正agent memory
2026-05-22 14:08:56 +08:00
huangruiteng 8817abdd38 bench(tau2): add generic scope fairness check (#2172)
* bench(tau2): add generic scope fairness check

* bench(tau2): focus scope fairness on generic prompt

* bench(tau2): remove domain-specific scope prompts
2026-05-21 19:51:14 +08:00
huangruitengandhuangruiteng 520713c624 feat(memory): upgrade trajectory extraction to beat no-memory baseline (#2017)
* feat(benchmark): add TAU-2 trajectory memory treatment

* style(benchmark): format tau2 trajectory scripts

* refine trajectory memory view prompt

* feat(benchmark): prepare tau2 memory corpora before eval

* fix(benchmark): tighten trajectory evidence prompt

* fix(benchmark): guard tau2 infrastructure failures

* fix(benchmark): resolve tau2 runner paths

* fix(memory): add trajectory evidence examples

* fix(benchmark): run no-memory tau2 eval in process

* bench(tau2): align retrieval budget and fixed first user

* bench(tau2): reuse memory corpora across eval runs

* bench(tau2): add scoped trajectory eval concurrency

* style(benchmark): format tau2 eval runner

* style(benchmark): satisfy tau2 eval lint

* bench(tau2): harden trajectory memory eval variants

* fix(tau2): rebuild search URI for reused corpora

* docs(tau2): add PR-B reproduction commands

* chore(tau2): keep PR-B benchmark scope focused

* chore(tau2): keep trajectory prompt generic

* chore(tau2): format benchmark scripts

* bench(tau2): restore operation-family trajectory protocol

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-05-21 11:15:34 +08:00
fujiajie666 05fe851af3 delete files (#2148) 2026-05-20 18:02:03 +08:00
fujiajie666 cfb680ac99 update vaka exp (#2147) 2026-05-20 17:36:04 +08:00
fujiajie666 fc5ed5458a Support Vaka memory templates via a switch (#2130)
* 通过开关支持vaka的记忆模板

* 更新模板
2026-05-19 21:45:00 +08:00
02a1de8993 fix: guard against empty choices and message=None in LoCaMo LLM judge (#2096)
* fix: guard against empty choices and message=None in LLM judge responses

resp.choices[0].message.content.strip() raises IndexError (empty choices)
or AttributeError (message=None on filtered content) in the LoCaMo
evaluation harness, aborting benchmark runs silently.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: guard empty LLM judge content

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-05-18 12:00:39 +08:00
Qin Haojie e52d3446a0 feat(benchmark): update session load benchmark (#2031)
* feat(benchmark): add OpenViking server load benchmark

* refactor(benchmark): update session load benchmark
2026-05-14 14:03:08 +08:00
huangruitengandhuangruiteng c9ebba0734 feat(benchmark): add TAU-2 Memory V2 eval runner (#2003)
* benchmark: add tau2 eval scaffold

* benchmark: gate pending tau2 memory adapter

* benchmark: use litellm provider model default

* benchmark: fold preflight into tau2 runner

* benchmark: document tau2 dependency setup

* benchmark: simplify tau2 simulator patch

* benchmark: keep simulator patch prompt clean

* benchmark: clarify simulator patch config

* benchmark: clarify tau2 adapter boundary

* benchmark: wire tau2 memory v2 eval

* benchmark: harden tau2 memory agent tool calls

* benchmark: tolerate empty tau2 assistant responses

* benchmark: normalize tau2 llm environment

* benchmark: add tau2 memory prewrite strategy

* benchmark: support current tau2 runner api

* benchmark: align tau2 memory prewrite parity

* benchmark: make tau2 eval traces safer

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-05-13 15:42:06 +08:00
DuTao 6312e1e12e feat(bot):Support user-key OpenViking mode and align memory namespaces (#1994)
* fix emb

* eval

* memory uri

* test

* eval

* eval

* eval

* eval

* eval

* user-key

* fix pr

* fix pr
2026-05-13 14:00:58 +08:00
Evo b71b9a0c2d docs(benchmark/locomo): add mem0/supermemory/claudecode to README directory tree (#2005) 2026-05-13 11:12:38 +08:00
t0saki 0a6c36213f feat(benchmark): add Claude Code LoCoMo evaluation harness (#1990)
* feat(benchmark): add Claude Code LoCoMo evaluation

Add locomo benchmark for Claude Code, following the same pipeline as
openclaw/vikingbot: ingest → QA → judge → stat.

- ingest.py: send each session to `claude -p`, let auto-memory extract
- eval.py: run QA questions in isolated project dirs with auto-memory
- judge.py / stat_judge_result.py: reuse same grading logic
- run_full_eval.sh: orchestrate full pipeline

Supports custom API endpoint (ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN),
environment isolation via HOME override, and retry on timeout/error.

* feat(benchmark): add prompt-prefix support and ingest-phase statistics

- ingest.py: add --prompt-prefix flag to nudge auto-memory during ingest
- run_full_eval.sh: add --project-root, --home, --result-dir, --prompt-prefix
  flags for flexible environment isolation (bare vs prompted experiments)
- stat_judge_result.py: report ingest-phase token/cost/latency alongside
  QA statistics, with --ingest-csv auto-detection

* feat(benchmark): add OpenViking integration for LoCoMo eval

- eval.py: add --hooks-settings, --mcp-config, --ov-config for OV recall
  during QA (sets OPENVIKING env vars, injects OV hint in prompt)
- run_full_eval.sh: wire --input, --hooks-settings, --mcp-config,
  --ov-ingest-config, --ov-qa-config params through to eval/ingest
- Add OV config files: ov-hooks.json, ov-mcp.json, ov-ingest.conf, ov-qa.conf
- Add import_to_ov.py for direct LoCoMo → OpenViking data import via SDK

* feat(benchmark): update configuration files and enhance evaluation logging

* feat(benchmark): add configuration files for Claude Code LoCoMo evaluation and enhance eval script

* feat(benchmark): reorganize LoCoMo claude-code eval around 4 reproducible modes

Collapse months of rN iteration scripts into 4 entry points covering
the four published numbers:

- run_prompted.sh   - vanilla CC auto-memory (no OV)
- run_sdk_iso.sh    - openviking SDK pre-ingest, per-sample namespace
- run_sdk_noiso.sh  - openviking SDK pre-ingest, shared namespace
- run_e2e.sh        - claude -p stream-json multi-turn + auto-capture

All four go through the same eval.py/judge.py/stat_judge_result.py
pipeline. Hook + plugin lookup is parametrised via OPENVIKING_PLUGIN_DIR
so the configs are portable. Historical scripts/reports/intermediate
configs are moved under .tmp/legacy (gitignored) for reference.

* style(benchmark): apply ruff format + auto-fix F541 across LoCoMo claude-code suite

- ruff format on eval.py / ingest.py / judge.py / stat_judge_result.py
- drop extraneous f-prefixes flagged by F541 in stat_judge_result.py
  and eval.py
2026-05-12 17:12:04 +08:00
chenjw 44d3cc41b1 Feat/memory isolation 支持群聊模式 (#1711) 2026-05-06 10:45:06 +08:00
Zayn Jarvis 6ae4a9398b feat: replace controversial examples with neutral alternatives (#1844) 2026-05-04 12:06:58 +08:00
DuTao d4a5e5ea39 fix emb (#1825) 2026-04-30 17:33:15 +08:00
yangxinxin-7 be0b375fb4 Revert "feat: bailian (#1664)" (#1665)
This reverts commit e706a8cdc6.
2026-04-23 15:51:16 +08:00