Files
OpenViking/benchmark
chenjwandTRAE CLI a843ab6bf2 session: add restricted Python DSL extraction protocol and make it the default (#4581)
* session: add restricted Python DSL extraction protocol and make it the default

Introduce a restricted Python memory SDK output protocol as an alternative to
the JSON extraction protocol, and switch the default to python. Both protocols
share the same ResolvedOperations post-processing, schema rules, and patch-repair
path via a new ExtractionOutputProtocol abstraction.

- Add extraction_output_protocol/{base,json,python}.py; python compiles a
  restricted AST into the same operations model as json.
- Default memory.extraction_output_format flips json -> python.
- Surface the offending source line on Python syntax errors and add targeted
  triple-quote retry guidance for string-literal breaks.
- Preserve every distinct fact on canonical merges; remove hardcoded memory
  type names from prompts so custom memory_types render dynamically.
- Downgrade benign batch-delete link-inheritance read failures to WARNING.
- Add memory_organization A/B benchmark and message_format pretty-printer.

Tests: extraction protocol, config loader, memory react suites pass.
Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: drop redundant entity split hint and duplicate abstractmethod

- entities.yaml: remove the size-triggered split hint; when to split/compact is
  decided at read time by memory_maintenance_notice, so the static schema
  description only keeps the identity semantics and fact-preservation rule.
- vlm/base.py: remove a duplicated @abstractmethod on get_completion_async.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* vlm: drop redundant *.vlm.call span decorators for trace parity

volcengine already dropped its @tracer("volcengine.vlm.call") wrapper to avoid
duplicate spans now that the request is logged via tracer.info(llm_input_messages=...).
Remove the symmetric litellm/openai decorators so all three backends behave the same.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* benchmark/locomo: fix commit_session kwarg for CLI AsyncHTTPClient

ov.AsyncHTTPClient resolves to openviking_cli.client._http_compat.AsyncHTTPClient,
whose commit_session takes a flat telemetry= kwarg and has no options= parameter.
Passing options={...} (the SDK-client shape) raised TypeError during import.
Use telemetry=True to match the CLI client, consistent with the other locomo
import scripts.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: don't parse error-target sentinels as URIs in extraction telemetry

The by-type extraction telemetry treated result.errors[].uri as a valid viking
URI and fell back to MemoryUpdater.memory_type_from_uri(), but that field is an
error *target* — it can be a sentinel like "unknown" or "events(page_id=100)".
VikingURI() then raised 'URI must start with viking://', turning a single
recorded extraction error into a crash of the whole long_term extraction step.
Count failed errors by the known uri->type map only, defaulting to "unknown".

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* bot: pass context_type via FindOptions after SDK find/search sync

The SDK find/search sync moved context_type from a top-level find() kwarg into
FindOptions. VikingBot still called client.find(context_type='memory') for peer
recall, so every per-turn type-quota recall raised 'unexpected keyword argument
context_type' and silently returned no memories. The answer agent then fell back
to manual multi-round search (iteration ~1.3 -> ~3.9) and accuracy dropped from
~83% to ~72-77%. Pass it via options={'context_type': 'memory'} instead.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* bot: adapt VikingClient.find to SDK FindOptions for context_type/filter

The SDK find/search sync moved context_type and filter out of top-level find()
kwargs into FindOptions. VikingClient.find still forwarded them as top-level
kwargs to the SDK client, so peer memory recall raised 'unexpected keyword
argument context_type' (and after the prior partial fix, 'options') and returned
no memories — the answer agent fell back to manual multi-round search, spiking
iteration ~1.3 -> ~4 and dropping accuracy ~83% -> ~76%.

Do the SDK adaptation once in VikingClient.find (pack context_type/filter into
options={...}); callers keep the stable VikingClient.find(context_type=...)
interface, so memory.py reverts to passing context_type= directly. Verified via
a single-question smoke: type_quota recall returns 13 memories, injection is
non-empty, iteration=1, answer correct.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: route event resolution repair through the output protocol

The event resolution-repair instruction was hardcoded to demand a JSON object,
but under the default Python protocol the repaired response is parsed by the
Python SDK compiler. When a first-pass event had out-of-bounds ranges, an
assistant-only span, or an ambiguous peer, the repair round returned JSON, the
compiler rejected it as an invalid program, retries were exhausted, and the
recoverable event memory was never written.

Add ExtractionOutputProtocol.render_resolution_repair(); JSON keeps the existing
JSON-object wording, Python asks for corrected sdk.create_events(...) calls.
_build_resolution_repair_instruction now delegates to the active protocol, like
patch-repair and the final instruction already do.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session,bot,benchmark: address PR review findings

- extract_loop: document 32768 as the extraction output floor (tuned for Doubao;
  lower-max models override via vlm.max_tokens) and extract
  _resolve_effective_max_output_tokens; ov.conf.example notes the override.
- python_protocol: alias non-identifier memory_type/field names on the Python DSL
  surface only (real names kept in URIs/storage/JSON); map aliases back when
  compiling, instead of hard-rejecting kebab-case custom schemas.
- run_full_eval.sh: move auto-commit + GIT_COMMIT_ID capture AFTER arg parsing so
  --auto-commit is honored and run metadata records the committed HEAD.
- litellm_vlm: strip Gemini cache_control from the already-sanitized messages so
  empty assistant turns are not reintroduced; sanitize_openai_messages passes
  through non-dict entries.
- Tests for each fix.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: reject Python DSL alias collisions instead of silently overwriting

_identifier_alias() is not one-to-one: memory_type 'project-notes' and
'project_notes' (or fields 'note-body'/'note_body') fold to the same DSL alias.
The alias->real dict comprehensions would silently drop one, making a schema/
field unreachable and routing writes to the wrong target. Add
_validate_alias_uniqueness(), invoked in render_contract() and the compiler
__init__ (so parse() paths without render are also guarded), which fails loudly
with a rename hint. Distinct-identifier names never collide, so real configs are
unaffected. Tests cover type, field, and no-render parse collisions.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: bound str.replace result size before allocation in Python DSL

The _MAX_EXPRESSION_SIZE guard covered * (repeat) and + (concat) but not the
whitelisted string methods: only join() had a projected-size check, so
('x'*1000).replace('x','y'*10000) could still allocate a >1MB result and bypass
the limit. Add _check_replace_size() that bounds source + occurrences*(len(new)
-len(old)) BEFORE calling str.replace (which builds the whole result in C), so
the oversized string is never allocated. replace is the only whitelisted method
that can materially inflate output (join already guarded; upper/lower/strip/
split/startswith/endswith do not grow). Tests cover an oversized replace being
rejected and a normal replace passing.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* session: close f-string width and str %-format inflation in Python DSL

The replace guard alone was insufficient: f-string format specs (f"{'x':>1000001}")
and str %-formatting ("%1000001s" % "x") also turn a small integer literal into an
arbitrarily large string with no repeat operator, bypassing _MAX_EXPRESSION_SIZE.
Neither has a legitimate use in memory content, so disallow them outright rather
than bounding width inflation: reject any non-empty f-string format spec and reject
str/bytes %-formatting (numeric % still allowed). Combined with the existing
*/+/join/replace pre-allocation checks, all small-input->large-output amplifiers
are now closed. Tests cover f-string width rejection, plain f-string, and str %.

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
2026-09-07 20:49:02 +08:00
..
2026-09-07 17:35:42 +08:00