* perf(vectordb): reuse pooled HTTP session per client
VikingDB clients issued every request through module-level
requests.request/post/get, which builds and discards a Session per
call. That means no connection pool and no keep-alive, so every
search/find pays a fresh TCP + TLS handshake.
Give each client one long-lived requests.Session backed by an
HTTPAdapter connection pool (mounted on both http:// and https://) and
route do_req through it, so repeated calls reuse warm connections.
Covers ClientForConsoleApi, ClientForDataApi, ClientForDataApiWithApiKey
and the private VikingDBClient.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
* perf(vectordb): reuse one pooled client in VikingDBProject
has_collection/get_collection/_get_collections built a fresh
VikingDBClient per call, so each rebuilt a pooled Session whose
keep-alive connection died with the GC'd client. Hoist the client to an
instance attribute so all metadata calls share one warm pool.
Note: requests.Session persists Set-Cookie across calls (the old
one-shot requests.request did not). These APIs are HMAC/Bearer
authenticated and do not use cookies, so this is only relevant if a
fronting LB sets affinity cookies, which would now stick per client.
Co-authored-by: TRAE CLI <traecli@bytedance.com>
---------
Co-authored-by: TRAE CLI <traecli@bytedance.com>
Reuse the server turn-budget planner for opt-in auto-commit retention, preserve manual full compaction and setup configuration, and handle zero retained turns when rebuilding pending tokens.
Refs #4415 (items 3 and 6).
Co-authored-by: tangtao <1024583279@qq.com>
* fix(studio): group agent session messages and improve chat readability
* feat(studio): refine sidebar icon system
* fix(studio): align session images with server format and localize errors
* fix(studio): unify active sidebar icon color
`SearchRequest` constrains four of them with pydantic `Field`; the tool declared them
as plain ints and a plain list, so the MCP face accepted values the REST face rejects:
max_tokens REST ge=64 le=32000 tool: any int
dedup_turns REST ge=0 le=100 tool: any int
rewrite_max_bullets REST ge=1 le=20 tool: any int
exclude_uris REST max_length=200 tool: any length
Three of those are only permissive -- the values are clamped downstream. `exclude_uris`
is not. `normalize_exclude_uris` slices to MAX_EXCLUDE_URIS regardless, so the extra
exclusions were dropped in silence and the URIs the caller asked to exclude came back
in the results:
MAX_EXCLUDE_URIS = 200; asking to exclude 250 URIs
MCP path -> kept 200 of 250; silently dropped 50
REST path -> too_long: List should have at most 200 items after validation, not 250
Declare the bounds with Annotated/Field rather than checking them in the body: FastMCP
puts them in the tool's published input schema (minimum/maximum/maxItems), so the model
driving the tool sees them before it picks a value, and enforces them on call_tool,
which is the path an MCP client takes.
MAX_EXCLUDE_URIS is imported rather than repeated, so the cap and the slice that
motivated it cannot drift apart.
Independent of #4886: that one is the list-mode guard, this is the context-mode bounds.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(pi-plugin): preserve camelCase tool result messages
* test(pi-plugin): cover native tool result capture and incremental sync
Refactor tests for extractBranchCapturePayloads to handle various tool result scenarios and ensure correct payload extraction.
* docs: stop the integration table from zebra-striping its second row
* docs: title-case the community heading and add the contributor graph
* docs: lay out the contributor graph in 15 columns
* docs: regroup integration tables and rename helper section
* docs: shrink integration cell captions
* docs: restore caption size, shorten Hermes caption
* docs: normalize integration logo sizes, tighten pi logo viewBox
* docs: add normalized integration logo set
* docs: refresh studio hero with light and dark variants
* docs: use OpenAI mark for the Codex entry
* docs: add DSH and Doubao Work to the integration table
* docs(readme): clarify value, concepts, and getting started
* docs(readme): fit concept and integration tables on narrow screens
* docs(readme): preserve original structure and header
* docs(readme): preserve research copy and Claude label
* docs(readme): keep Agent Plugins label on one line
* docs(readme): emphasize direct filesystem operations
* docs(readme): add new-tab attributes to Studio image links
* docs(readme): feature context compilation in Why section
* docs(readme): label DeerFlow integration as Plugin plus MCP
* docs(readme): keep MCP client labels on one line
* fix(openclaw): clear reset context without changing session identity
* fix(sessions): unblock commits after reset boundary write failure
* refactor(sessions): read reset boundary from .done and skip duplicate empty resets
- _is_context_reset_archive reads the context_reset property directly instead
of inferring it from a missing overview
- reset on an already-reset empty session returns confirmation without
appending another empty archive directory
* refactor(sessions): reset boundary archive holds only .done
Terminal archives never have messages.jsonl read and a missing overview
already reads as empty, so the empty placeholder files were unused.
* feat(pi-experimental): fork pi extension into experimental context-management skeleton
* feat(pi-experimental): add context-window core, OV archive client methods and contextWindow config
- lib/context-window-core.mjs: pure, io-injected core for agent-driven context
windows (tool-call-anchored cut, frozen window header, blocking reset
pipeline with a single deadline, pending-overview refresh, reminder ladder,
pi-compaction fallback rules)
- client.ts: sessionRootUri, readArchiveOverview/readArchiveMessages via
content/read, listSessionArchives, grepSessionArchives, getTask; the
/sessions/{id}/archives route is not used (blocked on the target gateway)
- config.ts/config.json: contextWindow block with clamps and env overrides;
dead captureMode key removed
* feat(pi-experimental): wire agent-driven context windows into pi
- context-window.ts adapter binds the core to SyncManager/OVClient/pi
- tools.ts: new_context, history (list_windows/list_items/read_item/
search_contents) and get_context_remaining with Codex-style names
- index.ts: offline restore before the health check, cut before recall,
per-prompt [context-status] message, one-shot reminders, pi-compaction
fallback through the core
- core follow-ups: abort checks before the handoff post, flush budget floor,
syncBranch in handleBeforeCompact, lazy session validation, persisted
previous overview
- docs: README, CONTEXT-WINDOW.md, agent-integrations pages; CI test glob
- scripts/e2e-window.mjs live gate replaces the takeover-era e2e-live
* fix(pi-experimental): apply three-lens review findings
- coexistence guard keyed on tool sourceInfo.path and re-checked on
start/before_agent_start/turn_end
- reminders: ignore aborted/error turns as first observation, gate the idle
note on the gap the user just returned from, threshold-free guidance
- core: persisted archives ledger for window/archive mapping, stale reminder
detection for custom-role messages, task-aware overview wait, retracted
handoff on post-handoff refusals, notes cleared on compaction fallback
- tools: refusal instructions, history fails closed on unreadable listings,
grep line/item index kept stable, viking_search scope text without ~
- config: recentResetGuardMs exposed, dead keys removed; docs synced
* fix(pi-experimental): status line metrics from the branch on a fresh process, archive line index stability, wording nits
* test(pi-experimental): reasoning level, deterministic tool output and keepable OV session in the e2e gate
- E2E_LLM_REASONING=off|minimal|low|medium|high|xhigh marks the model as
reasoning-capable and sets pi's defaultThinkingLevel; a custom relay also
needs compat.supportsReasoningEffort, which URL auto-detection cannot infer
- T1 now reads a seeded release.md, so the archive reliably carries a
[tool-result ...] entry instead of depending on the model reaching for a tool
- E2E_KEEP_OV_SESSION=1 skips the session delete so the archives stay readable
for a demo
* test(pi-experimental): add a long-context scenario where the agent resets under pressure
E2E_WINDOW_LONG=1 seeds the extension's own sources (22 files, ~105k tokens of
material) into the workspace and asks for a file-by-file inventory, without ever
mentioning the context tools. The gate then checks what the agent did on its
own: how full the window got, whether it reset, whether the cut was legal, and
whether the work continued across the boundary.
- peak pressure is measured from the provider payloads, not only from the
per-prompt [context-status] line, which undersamples a tool-heavy turn
- runTurn takes a timeout; the long turns get 25 minutes instead of 10
- judgement-dependent checks warn, harness behaviour still fails the gate
Observed on doubao-seed-2-1-pro with reasoning high: 96 requests, peak 47% of a
128k window, three self-initiated resets, 22/22 files inventoried.
* docs(pi-experimental): publish redacted demo evidence, drop the hardcoded relay
- demo-evidence/pi-ctxwin-demo.zip: two real runs against a live OpenViking
server and a live model, with transcripts, the provider payloads either side
of every reset and the archives pulled back off the server. Secrets file
absent, relay hostname and operator username replaced; REDACTIONS.md inside
the archive lists every substitution
- e2e-window.mjs no longer defaults E2E_LLM_BASE_URL / E2E_LLM_MODEL to a
private relay: both are now required, so no endpoint of anyone's is baked in
- CONTEXT-WINDOW.md documents the reasoning, long-context and keep-session
knobs added with those scenarios
* docs(pi-experimental): name the real endpoint and model in the demo evidence
The published archive now says what the runs actually used — Volcengine's
Doubao 2.1 Pro (doubao-seed-2-1-pro-260628) on Ark at
https://ark.cn-beijing.volces.com/api/v3 — instead of an anonymous relay
placeholder. The runs went through a private proxy in front of Ark; the same
key reaches the official endpoint directly, and REDACTIONS.md says so.
* docs(pi-experimental): say which repo-level hooks this directory deliberately does not touch
The extension keeps its diff inside its own folder, so the CI test glob and the
shared-file sync TARGETS entry are not added. Record both, plus the local
.gitignore that re-includes lib/, so a maintainer promoting this out of
experimental status knows what to wire up.
* docs(readme): fix publication citation line breaks
* docs(readme): present agent integrations in a logo table
* docs(readme): align integration logos across each row
* docs(readme): shorten Claude label and link integration details
* docs(readme): add DeerFlow and general integration logos
* docs(readme): keep integration captions and mode names on one line
* docs(readme): prevent CJK MCP labels from wrapping
* docs(readme): remove redundant MCP integration note
* docs(readme): add research publications and conference milestones
* docs(readme): connect publications to agent memory and retrieval
* docs(readme): add VikingRAG research and submission status
OpenAI Secure MCP Tunnel forwards X-Request-Id values formatted as
<uuid>/<suffix>. The strict charset [A-Za-z0-9._:-] rejected them with
HTTP 400, breaking MCP connectors created through the tunnel.
- allow '/' in the request id pattern (still capped at 128 chars)
- replace rejection with a regenerated uuid4 plus a warning log;
invalid raw values are never logged
* fix(plugins): drain the pending queue in-process so a transient write failure self-heals
The dsh memory plugin latches capture and commit on the first retryable
write failure (hasPendingWrites) and only reset the latch at session
init, so the long-lived dsh process stayed stuck until restart.
Add a per-process single-flight drainer (default 60s, env
OPENVIKING_PENDING_DRAIN_INTERVAL_MS) that follows the session-start
flow: probe health, replay the queue without consuming retry budgets,
then re-derive every session's latch from the queue. replayPending gains
an optional consumeRetries flag (default true, byte-compatible):
drainers release a failed claim back to its original filename instead of
incrementing the retry count, so the session-start path keeps owning all
retry accounting and D4 deletions. Latch and health transitions are
logged once per flip for observability.
* test(plugins): cover the drainer and non-consuming replay mode
Add pending-queue coverage for consumeRetries:false (retryable failures
stay retryable and ordered, non-retryable and exhausted entries still
delete, commitSession failures keep the run going, default mode
unchanged) and runtime drainer coverage (recovery clears the latch,
outages keep it and leave entries retryable, empty queue means zero
HTTP, commit resumes after the drain, single-flight, per-session latch
isolation, interval wiring with env fallback).
---------
Co-authored-by: pc.yu <nick@fourieralpha.com>
Follow-up to #4794: install the cache scope at the fan-out owners via a
query_embed_cache_scope context manager (covers /recall and the MCP
search tool), cover the /skills/find two-find fan-out, shield the shared
embed task against waiter cancellation, key entries by embedder identity,
and reset the ContextVar token on scope exit.
Co-authored-by: pc.yu <nick@fourieralpha.com>
The send:// branch of _parse_data_uri joined the URI remainder onto
get_data_path()/images/ with no validation, so a crafted outbound URI
such as send://../../secrets.txt read an arbitrary server file, and the
channel image delivery path would attach and exfiltrate it to the chat
platform. The URI text is model-authored at runtime, i.e. untrusted
exactly like the resource content fenced on the extraction paths
(#4292).
send:// remainders must now be bare image filenames ([A-Za-z0-9][A-Za-z0-9._-]{0,200}); anything containing path
separators, leading dots, or escapes raises ValueError before any file
is opened. #4660 widens the delivery surface to Discord/Slack/Telegram,
so the underlying loader needs this regardless of which channels ship.
Co-authored-by: mac <bishopapril850965@yahoo.com>