* perf(vectordb): reuse pooled HTTP session per client
VikingDB clients issued every request through module-level
requests.request/post/get, which builds and discards a Session per
call. That means no connection pool and no keep-alive, so every
search/find pays a fresh TCP + TLS handshake.
Give each client one long-lived requests.Session backed by an
HTTPAdapter connection pool (mounted on both http:// and https://) and
route do_req through it, so repeated calls reuse warm connections.
Covers ClientForConsoleApi, ClientForDataApi, ClientForDataApiWithApiKey
and the private VikingDBClient.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* perf(vectordb): reuse one pooled client in VikingDBProject
has_collection/get_collection/_get_collections built a fresh
VikingDBClient per call, so each rebuilt a pooled Session whose
keep-alive connection died with the GC'd client. Hoist the client to an
instance attribute so all metadata calls share one warm pool.
Note: requests.Session persists Set-Cookie across calls (the old
one-shot requests.request did not). These APIs are HMAC/Bearer
authenticated and do not use cookies, so this is only relevant if a
fronting LB sets affinity cookies, which would now stick per client.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Reuse the server turn-budget planner for opt-in auto-commit retention, preserve manual full compaction and setup configuration, and handle zero retained turns when rebuilding pending tokens.
Refs #4415 (items 3 and 6).
Co-authored-by: tangtao <1024583279@qq.com>
`SearchRequest` constrains four of them with pydantic `Field`; the tool declared them
as plain ints and a plain list, so the MCP face accepted values the REST face rejects:
max_tokens REST ge=64 le=32000 tool: any int
dedup_turns REST ge=0 le=100 tool: any int
rewrite_max_bullets REST ge=1 le=20 tool: any int
exclude_uris REST max_length=200 tool: any length
Three of those are only permissive -- the values are clamped downstream. `exclude_uris`
is not. `normalize_exclude_uris` slices to MAX_EXCLUDE_URIS regardless, so the extra
exclusions were dropped in silence and the URIs the caller asked to exclude came back
in the results:
MAX_EXCLUDE_URIS = 200; asking to exclude 250 URIs
MCP path -> kept 200 of 250; silently dropped 50
REST path -> too_long: List should have at most 200 items after validation, not 250
Declare the bounds with Annotated/Field rather than checking them in the body: FastMCP
puts them in the tool's published input schema (minimum/maximum/maxItems), so the model
driving the tool sees them before it picks a value, and enforces them on call_tool,
which is the path an MCP client takes.
MAX_EXCLUDE_URIS is imported rather than repeated, so the cap and the slice that
motivated it cannot drift apart.
Independent of #4886: that one is the list-mode guard, this is the context-mode bounds.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(openclaw): clear reset context without changing session identity
* fix(sessions): unblock commits after reset boundary write failure
* refactor(sessions): read reset boundary from .done and skip duplicate empty resets
- _is_context_reset_archive reads the context_reset property directly instead
of inferring it from a missing overview
- reset on an already-reset empty session returns confirmation without
appending another empty archive directory
* refactor(sessions): reset boundary archive holds only .done
Terminal archives never have messages.jsonl read and a missing overview
already reads as empty, so the empty placeholder files were unused.
OpenAI Secure MCP Tunnel forwards X-Request-Id values formatted as
<uuid>/<suffix>. The strict charset [A-Za-z0-9._:-] rejected them with
HTTP 400, breaking MCP connectors created through the tunnel.
- allow '/' in the request id pattern (still capped at 128 chars)
- replace rejection with a regenerated uuid4 plus a warning log;
invalid raw values are never logged
Follow-up to #4794: install the cache scope at the fan-out owners via a
query_embed_cache_scope context manager (covers /recall and the MCP
search tool), cover the /skills/find two-find fan-out, shield the shared
embed task against waiter cancellation, key entries by embedder identity,
and reset the ContextVar token on scope exit.
Co-authored-by: pc.yu <nick@fourieralpha.com>
The no-API-key / non-trusted branch previously returned the request body
unchanged, so a Studio client could smuggle a forged openviking_connection
through --with-bot. The Bot gateway trusts loopback forwards from this
proxy, so that identity would be accepted.
Always drop client-supplied openviking_connection before attaching the
authenticated connection (or forwarding without one in legacy dev mode).
Follow-up to the #4650 / #4649 trust model.
* fix(vlm): route configured reasoning_effort through extra_body for non-OpenAI models
_is_reasoning_model() gates the native reasoning_effort kwarg on the
gpt-5/o-series prefixes, so for any other model name a configured
vlm.reasoning_effort was silently dropped from the request. OpenAI-
compatible providers then ran at their server-side default; for the
GLM-5.3 family that default is max, turning session-commit extraction
into minutes-long calls (#4686).
Route an explicitly configured reasoning_effort through extra_body when
the model is not an OpenAI reasoning family:
- only injects when the user actually set the field (a defaulted value
is never invented), so backends that never configure it see no change
- an explicit extra_request_body.reasoning_effort always wins
- DashScope endpoints keep the enable_thinking mechanism untouched
- gpt-5/o-series keep the native top-level kwarg
GLMVLM inherits the routing (issue MRE now sees the value reach the API).
Fixes#4686
* refactor(vlm): simplify reasoning effort routing
Use mutually exclusive provider handling and setdefault for explicit body
precedence, reading configuration presence without a duplicate flag.
* refactor(vlm): separate configured reasoning effort from OpenAI defaults
Build shared completion parameters once for text and vision, forwarding
explicit reasoning effort independently of model prefixes and thinking.
Preserve OpenAI defaults and extra-body precedence, and consolidate the
contract coverage into existing tests.
---------
Co-authored-by: now-ing <now-ing@users.noreply.github.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(bot): forward actor_peer_id through openviking_connection
The --with-bot proxy built openviking_connection without RequestContext.actor_peer_id, so vikingbot fell back to body user_id.
Fixes#4649
* fix(bot): use direct actor_peer_id access and tidy test
Address #4650 review nits: read RequestContext.actor_peer_id directly, drop vikingbot-side fallback assertion from the unit test, and ensure the test file ends with a newline.
With query expansion enabled (planned=3), the context face fans out
find calls per (query x scope); each find re-embeds the same query text,
so a request made 34-36 embedding calls for 3 distinct texts. Add a
request-scoped cache: the /search handler installs a ContextVar scope
whose dict is shared by reference across the gather fan-out, and
embed_compat awaits the first in-flight embed for the same (model,
prepared text) key. Paths without the scope fall back to the plain
embed, and resource embeds are never cached.
Co-authored-by: pc.yu <nick@fourieralpha.com>
GET /api/v1/watches (and every WatchManager read/mutate primitive) filtered
visibility by role, so an ADMIN caller saw every task in the account. Cloud
libraries register their per-user keys as ADMIN, so ListOpenVikingWatches
called with the same account but different user_id returned identical tasks.
Drop the ADMIN blanket-visibility branch in _check_permission: ROOT keeps the
system/scheduler bypass (WatchScheduler.schedule_task reads tasks as ROOT), but
ADMIN now falls through to the same owner check as USER (task.user_id ==
user_id). This isolates watch tasks per user within a shared account without
any cloud-side change; the scheduler's own execution path already passes each
task's stored user_id/original_role, so background runs are unaffected.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* feat(compile): let OV own external task lifecycle
Keep durable task state, recovery, query, and cancellation in OpenViking while external providers execute the workload.
* feat(compile): align external session protocol
Persist external session state in OV and adapt VikingBot to the documented create, status, and cancel contract.
* fix(compile): infer external API from base URL
Remove the redundant enable switch so generated Base Server configuration activates Compile directly from its configured endpoint.
* refactor(compile): simplify task lifecycle controls
Remove client-facing runtime and wait controls, and retire legacy OV routes in favor of the generic Task API. Bound transient status polling failures so unavailable providers fail the owned task.
* fix(compile): restore server runtime deadline
* fix(compile): enforce timeout in OV task polling
* feat(compile): expose args in CLI and SDKs
* fix(compile): correct provider retry boundaries
* fix(compile): bound cancellation convergence
* fix(compile): retry submit without runtime cap
* refactor(compile): keep provider credentials minimal
* fix(compile): poll tasks every 30 seconds by default
* refactor(compile): standardize runtime task protocol
* fix(compile): wait for remote cancellation to settle
* fix(queuefs): give SemanticMsg a per-instance creation timestamp
The dataclass default int(datetime.now().timestamp()) is evaluated once
at class-definition time, and the custom __init__ never assigns
timestamp, so every message created by a long-running process fell
through to the shared class attribute and reported the process-start
epoch instead of its own creation time. The frozen value leaked into
every serialized queue message (to_dict/asdict) and on-disk payload.
Assign the epoch in __init__ and make the class-level default a
default_factory so any construction path that skips the custom __init__
also gets a fresh timestamp. from_dict keeps preserving stored
timestamps verbatim; messages restored without one now get the
restore-time epoch instead of the process-start epoch.
* ci: retrigger after known test_fs_cp pathlock flake (run 34067470781)
---------
Co-authored-by: mac <bishopapril850965@yahoo.com>
* session: add restricted Python DSL extraction protocol and make it the default
Introduce a restricted Python memory SDK output protocol as an alternative to
the JSON extraction protocol, and switch the default to python. Both protocols
share the same ResolvedOperations post-processing, schema rules, and patch-repair
path via a new ExtractionOutputProtocol abstraction.
- Add extraction_output_protocol/{base,json,python}.py; python compiles a
restricted AST into the same operations model as json.
- Default memory.extraction_output_format flips json -> python.
- Surface the offending source line on Python syntax errors and add targeted
triple-quote retry guidance for string-literal breaks.
- Preserve every distinct fact on canonical merges; remove hardcoded memory
type names from prompts so custom memory_types render dynamically.
- Downgrade benign batch-delete link-inheritance read failures to WARNING.
- Add memory_organization A/B benchmark and message_format pretty-printer.
Tests: extraction protocol, config loader, memory react suites pass.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: drop redundant entity split hint and duplicate abstractmethod
- entities.yaml: remove the size-triggered split hint; when to split/compact is
decided at read time by memory_maintenance_notice, so the static schema
description only keeps the identity semantics and fact-preservation rule.
- vlm/base.py: remove a duplicated @abstractmethod on get_completion_async.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* vlm: drop redundant *.vlm.call span decorators for trace parity
volcengine already dropped its @tracer("volcengine.vlm.call") wrapper to avoid
duplicate spans now that the request is logged via tracer.info(llm_input_messages=...).
Remove the symmetric litellm/openai decorators so all three backends behave the same.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* benchmark/locomo: fix commit_session kwarg for CLI AsyncHTTPClient
ov.AsyncHTTPClient resolves to openviking_cli.client._http_compat.AsyncHTTPClient,
whose commit_session takes a flat telemetry= kwarg and has no options= parameter.
Passing options={...} (the SDK-client shape) raised TypeError during import.
Use telemetry=True to match the CLI client, consistent with the other locomo
import scripts.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: don't parse error-target sentinels as URIs in extraction telemetry
The by-type extraction telemetry treated result.errors[].uri as a valid viking
URI and fell back to MemoryUpdater.memory_type_from_uri(), but that field is an
error *target* — it can be a sentinel like "unknown" or "events(page_id=100)".
VikingURI() then raised 'URI must start with viking://', turning a single
recorded extraction error into a crash of the whole long_term extraction step.
Count failed errors by the known uri->type map only, defaulting to "unknown".
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* bot: pass context_type via FindOptions after SDK find/search sync
The SDK find/search sync moved context_type from a top-level find() kwarg into
FindOptions. VikingBot still called client.find(context_type='memory') for peer
recall, so every per-turn type-quota recall raised 'unexpected keyword argument
context_type' and silently returned no memories. The answer agent then fell back
to manual multi-round search (iteration ~1.3 -> ~3.9) and accuracy dropped from
~83% to ~72-77%. Pass it via options={'context_type': 'memory'} instead.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* bot: adapt VikingClient.find to SDK FindOptions for context_type/filter
The SDK find/search sync moved context_type and filter out of top-level find()
kwargs into FindOptions. VikingClient.find still forwarded them as top-level
kwargs to the SDK client, so peer memory recall raised 'unexpected keyword
argument context_type' (and after the prior partial fix, 'options') and returned
no memories — the answer agent fell back to manual multi-round search, spiking
iteration ~1.3 -> ~4 and dropping accuracy ~83% -> ~76%.
Do the SDK adaptation once in VikingClient.find (pack context_type/filter into
options={...}); callers keep the stable VikingClient.find(context_type=...)
interface, so memory.py reverts to passing context_type= directly. Verified via
a single-question smoke: type_quota recall returns 13 memories, injection is
non-empty, iteration=1, answer correct.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: route event resolution repair through the output protocol
The event resolution-repair instruction was hardcoded to demand a JSON object,
but under the default Python protocol the repaired response is parsed by the
Python SDK compiler. When a first-pass event had out-of-bounds ranges, an
assistant-only span, or an ambiguous peer, the repair round returned JSON, the
compiler rejected it as an invalid program, retries were exhausted, and the
recoverable event memory was never written.
Add ExtractionOutputProtocol.render_resolution_repair(); JSON keeps the existing
JSON-object wording, Python asks for corrected sdk.create_events(...) calls.
_build_resolution_repair_instruction now delegates to the active protocol, like
patch-repair and the final instruction already do.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session,bot,benchmark: address PR review findings
- extract_loop: document 32768 as the extraction output floor (tuned for Doubao;
lower-max models override via vlm.max_tokens) and extract
_resolve_effective_max_output_tokens; ov.conf.example notes the override.
- python_protocol: alias non-identifier memory_type/field names on the Python DSL
surface only (real names kept in URIs/storage/JSON); map aliases back when
compiling, instead of hard-rejecting kebab-case custom schemas.
- run_full_eval.sh: move auto-commit + GIT_COMMIT_ID capture AFTER arg parsing so
--auto-commit is honored and run metadata records the committed HEAD.
- litellm_vlm: strip Gemini cache_control from the already-sanitized messages so
empty assistant turns are not reintroduced; sanitize_openai_messages passes
through non-dict entries.
- Tests for each fix.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: reject Python DSL alias collisions instead of silently overwriting
_identifier_alias() is not one-to-one: memory_type 'project-notes' and
'project_notes' (or fields 'note-body'/'note_body') fold to the same DSL alias.
The alias->real dict comprehensions would silently drop one, making a schema/
field unreachable and routing writes to the wrong target. Add
_validate_alias_uniqueness(), invoked in render_contract() and the compiler
__init__ (so parse() paths without render are also guarded), which fails loudly
with a rename hint. Distinct-identifier names never collide, so real configs are
unaffected. Tests cover type, field, and no-render parse collisions.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: bound str.replace result size before allocation in Python DSL
The _MAX_EXPRESSION_SIZE guard covered * (repeat) and + (concat) but not the
whitelisted string methods: only join() had a projected-size check, so
('x'*1000).replace('x','y'*10000) could still allocate a >1MB result and bypass
the limit. Add _check_replace_size() that bounds source + occurrences*(len(new)
-len(old)) BEFORE calling str.replace (which builds the whole result in C), so
the oversized string is never allocated. replace is the only whitelisted method
that can materially inflate output (join already guarded; upper/lower/strip/
split/startswith/endswith do not grow). Tests cover an oversized replace being
rejected and a normal replace passing.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* session: close f-string width and str %-format inflation in Python DSL
The replace guard alone was insufficient: f-string format specs (f"{'x':>1000001}")
and str %-formatting ("%1000001s" % "x") also turn a small integer literal into an
arbitrarily large string with no repeat operator, bypassing _MAX_EXPRESSION_SIZE.
Neither has a legitimate use in memory content, so disallow them outright rather
than bounding width inflation: reject any non-empty f-string format spec and reject
str/bytes %-formatting (numeric % still allowed). Combined with the existing
*/+/join/replace pre-allocation checks, all small-input->large-output amplifiers
are now closed. Tests cover f-string width rejection, plain f-string, and str %.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Replace the derived boolean with an extensible mode so restricted inheritance can become a single additional state.
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* fix(agent-evolution): preserve legacy experience lineage tags
* fix(tags): relax character restrictions while retaining length limits
* fix(tags): remove search tag format and length restrictions
* fix(tags): retain only comma validation for search tags
* fix(tags): limit relaxation to charset and length checks