Files
OpenViking/benchmark/longmemeval/openviking
9eac8a6d3d feat(sdk): sync go/ts/python SDKs with server find/search, recall, an… (#3737)
* feat(sdk): sync go/ts/python SDKs with server find/search, recall, and admin changes

Server-side changes recently landed that the language SDKs had drifted from:

- find/search results now return `tags` and no longer return
  `category`/`match_reason`/`relations`/`overview` (#3730). Go's strict
  struct was the only one broken; update MatchedContext accordingly.
- new admin endpoints for agent-evolution and per-account settings (#3695).
- public `search/recall` endpoint was missing from all SDKs.

Changes:
- python: add `level`/`since`/`until`/`time_field` to find/search; add an
  `extra` escape hatch to find/search/add_resource/write/batch_write so new
  server fields can be passed without an SDK bump (only forwarded when set,
  preserving `level=0`); add `recall` and the four admin methods.
- go: fix MatchedContext (add Tags, drop removed fields), add Recall and the
  four admin methods.
- typescript: type MatchedContext/FindResult, add RecallOptions, add `recall`
  and the four admin methods.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): unify options APIs and sync latest server interfaces

- migrate complex Python SDK calls to typed options dictionaries
- add dedicated context search and consistent extra-field handling
- align Go and TypeScript options with omission-aware serialization
- support session config, event tags, Agent Evolution date filters,
  OpenViking Assets, batch write, downloads, and create_parent
- refresh SDK tests and examples across all three languages

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): address options API review findings

- fix Go session extra merging and Python message precedence
- adapt LangChain calls to the Python options API
- migrate repository examples, tests, and documentation

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): complete options migration and message parity

- migrate remaining Python SDK benchmarks to options dictionaries
- normalize empty parts consistently for single and batch messages
- add regression guards for repository SDK call sites

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): align reindex options after main rebase

- preserve reindex tags in Python typed options
- add reindex extra support for Go and TypeScript
- reject official fields passed through extra across SDKs

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): support legacy keyword options

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

docs(sdk): use explicit Python SDK arguments

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): support set tags extra options

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): expose Go add resource options

Expose AddType and ProcessingMode through Go AddResourceOptions and serialize them to the resources API. Add a regression test covering the resulting request payload.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): flatten core Python client options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

docs(sdk): align Python examples with flattened options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): preserve core API compatibility

Co-authored-by: TRAE CLI <traecli@bytedance.com>

refactor(python-sdk): move resource hints to options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): align resource option callers

Co-authored-by: TRAE CLI <traecli@bytedance.com>

test(sdk): cover recursive reindex forwarding

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): preserve Go options compatibility

Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(python-sdk): expose message peer id

Co-authored-by: TRAE CLI <traecli@bytedance.com>

test(python-sdk): consolidate options coverage

Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(python-sdk): add parts and flatten image search

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* docs(sdk): align Python call examples

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-24 14:09:11 +08:00
..

LongMemEval OpenViking Benchmark

This directory contains the OpenViking evaluation flow for LongMemEval:

  1. import each user's haystack sessions into OpenViking;
  2. run one retrieval call per question;
  3. optionally rerank the retrieved memories;
  4. answer from the selected memory context;
  5. judge and summarize the result CSV.

The benchmark expects an OpenViking server to already be running. The commands below use the default local OpenViking client configuration. If you need a different endpoint, pass --openviking-url to both import and eval.

Data

Set the dataset path once:

DATA=/path/to/longmemeval_s_cleaned.json

The importer uses each sample's original question_id-derived user id, such as lm_user_<id>, so the eval script searches the same user space created during import.

Import

python benchmark/longmemeval/openviking/import_to_ov.py \
  --input "$DATA" \
  --parallel 16 \
  --submit-parallel 16 \
  --wait-mode deferred \
  --success-csv result/longmemeval_openviking_import_success.csv \
  --error-log result/longmemeval_openviking_import_errors.log

Use --force-ingest only when intentionally re-importing existing sessions.

Smoke Test

Run one question before a full evaluation:

python benchmark/longmemeval/openviking/run_eval.py "$DATA" \
  --output result/longmemeval_openviking_smoke.csv \
  --count 1 \
  --threads 1 \
  --single-search-context-limit 50 \
  --single-search-rerank-limit 10 \
  --single-search-max-context-chars 30000 \
  --debug-print-model-input

With --debug-print-model-input, the CSV includes the full answer prompt and retrieval trace. Check retrieved_uris_by_iteration to confirm rerank is enabled and context_uris contains the expected number of memories.

Full Eval

This example retrieves 50 memories, keeps the top 10 after rerank, and caps the memory text passed to the answer model at 30000 characters.

OUT=result/longmemeval_openviking_search50_rerank10_chars30000.csv

python benchmark/longmemeval/openviking/run_eval.py "$DATA" \
  --output "$OUT" \
  --threads 8 \
  --timeout 900 \
  --single-search-context-limit 50 \
  --single-search-rerank-limit 10 \
  --single-search-max-context-chars 30000

Set --single-search-rerank-limit 0 to disable rerank. Set --single-search-max-context-chars 0 to disable the character budget.

Judge And Stat

python benchmark/longmemeval/openviking/judge.py \
  --input "$OUT" \
  --parallel 40

python benchmark/longmemeval/openviking/stat_judge_result.py \
  --input "$OUT"

The judge writes result values back into the same CSV. Use --force to re-grade rows that already have a result. Use --strict-prompt when you want the stricter LongMemEval judge prompt instead of the default lenient prompt.

stat_judge_result.py prints overall accuracy, average memory token/character usage, and accuracy grouped by LongMemEval question type.