Commit Graph
128 Commits
Author SHA1 Message Date
Haoyu ZhangandQin Haojie 3f3554256b feat: 支持基于火山方舟的音视频多模态理解 (#3563)
* feat: add audio and video understanding via VLM

* docs: design media resource guards

* fix: bound media staging concurrency

* fix: cap unknown-size media staging

* test: stage media in routing fake

* test: exercise media staging callbacks

* test: trim media understanding coverage

* chore: 清理实现计划文档

* fix: 修复多凭证切换问题

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-03 16:44:22 +08:00
DuTao d328b05432 feat(bot): support OpenAI and credential streaming (#3696) 2026-08-03 15:13:34 +08:00
DuTao 6b416cd66d feat(bot): support image inputs in VikingBot Chat API (#3619)
* bot支持图片对话、去除Lite llm无效逻辑

* fix error

* clarify remote image URL validation
2026-07-30 14:00:48 +08:00
Hao Zheandzhiheng.liu 797406c381 fix(sdk): consolidate model, client, and evaluation correctness (#3550)
* fix(models): route multimodal inputs to DashScope embed_content

Sweep findings: A-09. Preserve image parts through the shared embedding entrypoints.

(cherry picked from commit 50b20d7ca9)

* fix(models): constrain DashScope multimodal routing

Sweep findings: A-09. Preserve text mode and require a single fused vector for multipart input.

(cherry picked from commit 77784429d4)

* fix(models): respect DashScope Qwen fusion parameters

Sweep finding A-09

(cherry picked from commit aafc702b20)

* fix(review): preserve tongyi multipart compatibility

Addresses blocking review finding on #3400.

(cherry picked from commit 96844047c5)

* fix(eval): flush queued records on stop and adapt RAG pipeline to FindResult

Sweep findings: B-05, B-12. Drain recorder queues through the sentinel and consume current retrieval result objects.

(cherry picked from commit fecdbeb6d4)

* fix(sdk/python): support sync client inside a running event loop

Sweep findings: F-09. Run sync wrappers on one persistent worker loop with result and exception propagation.

(cherry picked from commit ba0faadc87)

* fix(sdk/python): preserve cancellation and fork safety

Sweep finding: F-09. Preserve original cancellation errors and reset worker synchronization after fork.

(cherry picked from commit a19e3c95e9)

* feat(sdk/go): add tags filter and relations API for parity

Sweep findings: F-11, F-12. Expose existing server capabilities consistently to Go callers.

(cherry picked from commit 8af0c2cb36)

* fix(sdk/python): runnable quickstarts, correct migrate payload, explicit timeout precedence

Sweep findings: F-01, F-02, F-14. Initialize documented clients and preserve Python SDK request/config semantics.

(cherry picked from commit d9a64f5665)

---------

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
2026-07-28 18:11:30 +08:00
DuTao a7c1065f01 fix(bot): preserve token usage for no-tool responses (#3535) 2026-07-27 17:05:47 +08:00
DuTao 3319ffd5b7 feat(bot): support VLM credential failover (#3503)
* bot支持多模型

* fix pr

* fix(bot): isolate console VLM credentials
2026-07-27 12:11:48 +08:00
Colter Dahlberg e97b74d2e5 feat(embedder): support extra_body passthrough in OpenAI embedder config (#3495)
* feat(embedder): support extra_body passthrough in OpenAI embedder config

Adds optional `extra_body` (dict) to the OpenAI dense embedder, merged
into every embeddings.create call. Motivating use case: OpenRouter
provider routing ({"provider": {"sort": "latency"}}) — default routing
shows p90=35s/max=127s tail latency that kills interactive recall
(A/B: sorted routing is consistently sub-second).

Explicit query_param/document_param keys still take precedence on
conflict.

* feat(config): wire extra_body through embedding config layer

Add optional extra_body field to EmbeddingModelConfig and pass it to
OpenAIDenseEmbedder for the openai/azure providers, including the
multi-credential failover merge (parent-level model-behavior field).

* docs(config): document embedding extra_body with OpenRouter routing example

* docs(config): restore concrete host/cors_origins values in EN full schema
2026-07-24 15:07:44 +08:00
Yu Zhangandzhangyu.34 a57adb951f fix(rerank): support DashScope nested request/response envelope (#3463)
* fix(rerank): support DashScope nested request/response envelope

OpenAIRerankClient sent a flat request body ({"model", "query",
"documents"}) and parsed "results" at the top level of the response.
DashScope (qwen3-rerank) requires a nested envelope:

  Request:  {"model", "input": {"query", "documents"}, "parameters": ...}
  Response: {"output": {"results": [...]}, "request_id", "usage"}

This caused DashScope rerank to silently fail — the response had no
top-level "results" key, so the client returned None.

Changes:
- Add _is_dashscope() to detect DashScope endpoints by host marker.
- Add _build_request_body() that produces the nested envelope for
  DashScope and the flat body for standard OpenAI/Cohere services.
- Add _extract_results() that reads output.results for DashScope and
  top-level results for standard services.
- Accept both "relevance_score" (singular, DashScope) and
  "relevance_scores" (plural, some providers) in result items.
- Add 13 tests covering host detection, body construction, response
  parsing, end-to-end mocked flows for both providers, plural key
  handling, empty documents, and sparse results.

Fixes #3459

* fix(rerank): detect DashScope protocol by URL path, not hostname

Reviewer noted the previous hostname-based switch broke the documented
qwen3-rerank compatible-api endpoint (/compatible-api/v1/reranks), which
must use the flat OpenAI-style body and top-level results.

Switch to path-based detection: only /api/v1/services/rerank uses the
native nested input/output envelope; everything else (including the
DashScope compatible-api and generic OpenAI/Cohere gateways) keeps the
flat protocol. Rename _is_dashscope -> _uses_nested_envelope for clarity.

Add regression tests covering the compatible-api flat path and reconcile
the existing native-path fixtures to the nested envelope.

* docs(rerank): use qwen3-rerank for compatible-api example

The compatible-api/v1/reranks endpoint uses the flat OpenAI-compatible
protocol; qwen3-vl-rerank is a native-envelope model served at
/api/v1/services/rerank. Align the example model with the endpoint the
implementation selects by URL path.

---------

Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
2026-07-22 19:30:41 +08:00
DuTao 55a9d12cd6 1. 优化评测参数化; (#3332)
2. 优化评测显示;
3. 修复gpt-5.6 api返回 无 choices时bot兼容问题。
2026-07-17 17:46:57 +08:00
19ca274a24 fix(retrieve): bound reranker input size (#3289)
* fix(retrieve): bound reranker input size

* fix(retrieve): make rerank input limit opt-in

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-16 17:19:38 +08:00
michaeltarletonandMichael Tarleton 3b7c61d266 fix(rerank): support Voyage via litellm (plain-string documents + dict results) (#3291)
LiteLLMRerankClient.rerank_batch wrapped each document as {"text": d} and read
result items via getattr(item, ...). That works for Cohere-style object results
but breaks Voyage through litellm: Voyage's rerank API rejects dict-wrapped
documents (400: 'documents' is not a valid string), and litellm returns Voyage
results as plain dicts, so getattr(item, "index") misses and rerank silently
falls back to a no-op.

- Pass documents as plain strings (litellm.rerank expects List[str]).
- Add _result_field() to read index/relevance_score from dict- or object-shaped
  result items.

Verified end-to-end against voyage/rerank-2.5 (scores now applied). Adds
tests/unit/models/rerank/test_litellm_rerank.py covering the plain-string
documents contract and both dict- and object-shaped results.

Co-authored-by: Michael Tarleton <mtarleton@istation.com>
2026-07-16 15:34:55 +08:00
Qin Haojie d47f2106ee refactor: remove unused and deprecated APIs (#3272)
Delete dead compatibility paths and test-only helpers so unsupported APIs do not remain as accidental contracts.
2026-07-16 10:49:56 +08:00
Qin Haojie ca70bc0649 refactor(embedding): remove unused batch APIs (#3260) 2026-07-15 16:44:57 +08:00
allenliang2022andallenliang2022 64546b8270 fix(vlm): preserve GitHub Copilot routes (#3170)
Co-authored-by: allenliang2022 <allenliang2022@users.noreply.github.com>
2026-07-13 15:10:00 +08:00
huangruitengandhuangruiteng 2f5b2e27e1 fix(rerank): accept sparse indexed results (#3121)
* fix(rerank): accept sparse indexed results

* fix(rerank): warn on sparse provider results

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-11 10:00:56 +08:00
t0saki d9a9b3d2e1 fix(vlm): set Ollama num_ctx and disable thinking for local memory extraction (#2992)
Ollama defaults to a 4096-token context window and silently truncates any
longer prompt to fit. OV's memory-extraction prompt is ~5.4k tokens even for a
short session, so the conversation (which sits at the top of the prompt) is
dropped and the model receives only the format spec. It then returns an empty
`{"memories": []}` with no error, so local Ollama VLMs extract nothing
regardless of model size. Thinking models compound this by emitting only
reasoning and stalling.

- litellm_vlm: for `ollama/` and `ollama_chat/` models, default `num_ctx`
  to 16384 and disable thinking via `extra_body`. Both are overridable through
  `extra_request_body`, and the change is gated to Ollama routes so other
  providers are untouched. This fixes extraction, intent analysis, and the
  query planner for every local Ollama config, not just wizard-generated ones.
- setup_wizard: write the same `extra_request_body` explicitly in the Ollama
  VLM config block so the setting is visible and tunable in ov.conf.
- setup_wizard: drop the qwen3.5:2b VLM preset. 2B models "extract" the
  prompt's few-shot examples as fabricated memories; qwen3.5:4b is the smallest
  model that extracts cleanly. RAM-tier defaults are reindexed accordingly.

Verified end to end: with num_ctx raised, prompt_eval goes from 4096
(truncated) to the full 5441 tokens and qwen3.5:4b extracts the correct
memories; without it, extraction returns empty.
2026-07-06 12:09:30 +08:00
MaojiaSheng 3c419ed09a chore: replace doubao 2.0 pro with doubao 2.0 lite as recomended VLM model (#2984) 2026-07-03 16:09:20 +08:00
chenjwandClaude fd73dcf23a Feat/自进化(经验记忆)框架重构 (#2503)
* Add trajectory experience learning redesign doc

* auto-commit before eval 20260607_043406

* auto-commit before eval 20260607_044129

* auto-commit before eval 20260607_123706

* auto-commit before eval 20260607_125514

* auto-commit before eval 20260607_133737

* auto-commit before eval 20260607_144649

* auto-commit before eval 20260607_154631

* Refine streaming memory train merge pipeline

* Refine session train policy optimization architecture

* Add VikingMem ARA paper analysis

* Force merge for mixed extraction memory patches

* auto-commit before eval 20260608_134426

* auto-commit before eval 20260608_142108

* auto-commit before eval 20260608_153909

* auto-commit before eval 20260608_154845

* auto-commit before eval 20260608_170143

* update

* auto-commit before eval 20260611_150946

* auto-commit before eval 20260611_153933

* auto-commit before eval 20260611_154251

* Fix tau2 reward wrapper call

* auto-commit before eval 20260611_193803

* auto-commit before eval 20260611_194939

* update

* auto-commit before eval 20260612_111029

* auto-commit before eval 20260612_112104

* auto-commit before eval 20260612_122603

* auto-commit before eval 20260612_123359

* auto-commit before eval 20260612_124303

* auto-commit before eval 20260612_130257

* Fallback peer routing to first conversation peer

* Route self memory through self peer sentinel

* Keep self sentinel out of peer memory paths

* auto-commit before eval 20260612_154051

* auto-commit before eval 20260612_154850

* auto-commit before eval 20260612_161633

* auto-commit before eval 20260612_184022

* auto-commit before eval 20260612_201845

* auto-commit before eval 20260612_202637

* auto-commit before eval 20260612_204040

* auto-commit before eval 20260612_224621

* Fix locomo progress column initialization

* Add memory field versioning

* auto-commit before eval 20260612_232318

* Simplify locomo progress display

* Remove locomo progress elapsed time

* Batch streaming memory merges by group

* Derive patch merge language from patches

* Detect patch merge language from updated files

* auto-commit before eval 20260613_004339

* auto-commit before eval 20260613_005835

* Persist memory update trace id

* auto-commit before eval 20260613_012722

* auto-commit before eval 20260613_013923

* auto-commit before eval 20260613_014708

* Enforce peer scope after memory merge

* auto-commit before eval 20260613_033402

* auto-commit before eval 20260613_151931

* auto-commit before eval 20260613_164217

* chore: raise vikingbot eval parallelism

* chore: tune vikingbot parallelism to 150

* auto-commit before eval 20260613_185807

* chore: restore vikingbot parallelism default

* feat(locomo): add import progress reporting

* chore(memory): restore profile and preference templates

* Fix tau2 reward JSON serialization

* Refactor tau2 batch memory training

* Stream batch train JSONL events

* Add fast path for batch training case specs

* Optimize streaming train gradient chunking

* Optimize patch merge prompt context

* fix tau2 memory training vectorization

* fix(memory): revert profile preference granularity rules

* bd init: initialize beads issue tracking

* update

* Log memory template fallback failures

* Record all rollout artifacts

* Fix OpenViking peer search forwarding

* Stop tracking Beads local state

* auto-commit before eval 20260616_002037

* Deprecate memory version selector

* Retry transient LoCoMo import HTTP failures

* Add memory schema stage and peer routing

* Organize LoCoMo benchmark outputs

* Restore VikingBot user memory auto recall

* Show elapsed time on LoCoMo progress bars

* Quiet transient import retries

* Shorten LoCoMo progress bars

* Route non-peer memories to self scope

* auto-commit before eval 20260616_124513

* Suppress memory read not found logs

* Limit LoCoMo import memory types

* Rename peer routing schema flag

* Rename peer schema flag to enable_peer

* Rename schema peer flag to peer_enabled

* auto-commit before eval 20260616_135946

* auto-commit before eval 20260616_140641

* auto-commit before eval 20260616_141753

* Show cached baseline eval at start of training

* Preserve remote policy contents

* Show failed work in progress bars

* Hide zero failed progress counts

* Disable tau2 service progress by default

* Reuse policy lock for policy deletes

* feat: add session skill extraction to Memory V3 streaming trainer

- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names

Co-authored-by: Claude <noreply@anthropic.com>

* Persist experience reminders in tau2 rollouts

* Enable tau2 epoch test eval by default

* Persist train rollout artifacts incrementally

* Ensure tau2 vikingbot user simulator deps

* Auto repair tau2 vikingbot simulator deps

* Avoid blocking tau2 vikingbot service loop

* Avoid tau2 gym reset when loading cases

* Clean tau2 rollout commit messages

* Clean tau2 tool trajectory serialization

* Retry vikingbot VLM rate limits

* Refine tau2 training case selection

* Promote vikingbot hook execution log level

* Improve VLM rate limit retry detection

* Update trajectory analysis prompt format

* Limit tau2 service logs to warnings

* Run tau2 vikingbot rollouts on service loop

* Lower vikingbot experience recall threshold

* Offload tau2 vikingbot blocking setup

* Retry tau2 LiteLLM rate limits

* Pin trajectory and experience outputs to Chinese

* Retry tau2 rate limits indefinitely

* Highlight tau2 training accuracy summaries

* Hide redundant avg reward console metrics

* Tighten memory extraction templates

* Reduce tau2 memory template noise

Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.

Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.

* Constrain tau2 memory extraction sources

Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.

Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.

* Preserve tau2 train non-run results

* Improve memory extraction guardrails

Run: result/tau2/train/run_airline_20260619_044051

tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.

* Support train split eval in tau2 batch runs

* Add slot support to tau2 vikingbot launcher

* Copy OpenViking configs for tau2 slots

* Tune tau2 case1 memory extraction

Run: result/tau2/train_1/run_airline_20260619_201546

Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.

* Advise tau2 train case1 best result

Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.

* Tune tau2 memory gate extraction

* Advise tau2 train case1 50pct result

* Guard failed write experience branches

* Advise tau2 train case1 100pct result

* Guard tau2 oracle training memories

* Recall trajectory diagnostics for tau2 rollouts

* Recall tau2 case specs for training rollouts

* Guard evaluated tau2 final states

* Inject compact tau2 oracle checklists

* Stabilize tau2 slot train multi-case runs

* Guard tau2 case10 oracle terminal state

* Use supported tau2 training memory types

* Match tau2 oracle writes by expected subset

* Autofill tau2 case10 oracle writes before done

* Enable tau2 case10 guard for train split

* Record slot1 S008 case10 guard best advice

* Generalize tau2 S008 oracle terminal guard

* Record slot1 S008 general guard best advice

* Remove tau2 benchmark oracle guard

* Prevent training ground truth memory recall

* Refine tau2 training memory extraction

* Fix epoch train rollout artifact stage

* Refine memory training rollout pipeline

* update

* auto-commit before eval 20260623_120317

* fix sdk read_raw for memory metadata

* use visible case links for experience recall

* auto-commit before eval 20260623_225354

* tau2/train: cap run_batch_train_eval rollout concurrency at 100

* update

* update

* update

* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2

- Port _same_memory_file filter to compressor_v3._build_memory_diff so
  no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
  (aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
  extract_long_term_memories so session skill URIs written by the
  streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
  apply_result
- Remove four dead skill-related imports left from the unbuilt v3
  execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
  extract_execution_memories method exists

* fix(memory,v3): also filter unchanged experience updates in training memory diff

* train: finish rollout and memory refactor

* memory: refine runtime-visible extraction prompts

* train: constrain communication memory extraction

* auto-commit before eval 20260629_235623

* memory: address training review fixes

* update

* update

* message: reuse part deserializer

* train: snapshot memory prompt yaml

* prompts: restore memory yaml templates from main

* memory: scope streaming update results

* update

* update

* session: train canonical merged cases

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-03 11:27:17 +08:00
Qin Haojie 2bf9e4962b fix(vlm): omit unset default max_tokens (#2949)
* fix(vlm): omit unset LiteLLM max_tokens

* fix(vlm): remove provider max_tokens fallbacks
2026-07-02 17:07:47 +08:00
Qin Haojie b90c6626cb fix(vlm): omit unset OpenAI max_tokens (#2946) 2026-07-02 15:59:59 +08:00
DuTao d689a285da fix(bot): Tool-call ordering, thinking mode, and image uploads (#2944)
* fix:
1. vlmadapter ,tool index缺失,导致的顺序异常;
2. bot默认开启think

* fix:
图片生成保存正确的格式,发送也使用正确格式

* fix:
图片生成保存正确的格式,发送也使用正确格式

* fix: preserve VLM thinking setting for DashScope
2026-07-02 15:28:42 +08:00
Qin Haojie 07708225cb fix(volcengine): forward model request headers (#2909)
Ensure Volcengine VLM and embedding calls propagate configured headers and default the client request id header for service traffic.
2026-07-01 14:31:39 +08:00
DuTao 9253619a98 add timeout config (#2917) 2026-07-01 11:24:58 +08:00
0102a48c2a fix(skills): now we allow viking://agent/skills again, and optimize CLI for skills (#2813)
* chore: clear unused files

* fix(tests): fix unit test

* refactor(auth): introduce plugin-based authentication architecture

Replace the monolithic `openviking/server/auth.py` with an extensible
plugin-based auth system. This refactor extracts the three built-in modes
(`dev`, `api_key`, `trusted`) into separate `AuthPlugin` implementations,
adds a registry for third-party plugins, and preserves all existing behavior
while enabling custom authentication backends (e.g. LDAP, OIDC, mTLS).

Key changes:
- **New public API**: `AuthPlugin` (ABC) and `register_auth_plugin` decorator.
- **New registry**: `AuthPluginRegistry` supports runtime registration.
- **Built-in plugins**: `DevAuthPlugin`, `ApiKeyAuthPlugin`, `TrustedAuthPlugin`.
- **Config change**: `auth_mode` widened from `Literal` to `str` for custom modes.
- **Validation delegated**: `validate_server_config()` now delegates to the active
  plugin's `validate_config()`, preserving existing validation semantics.
- **Router compatibility**: All existing `require_*` decorators and `resolve_identity`
  / `get_request_context` dependencies remain unchanged. Routers import the same
  symbols from `openviking.server.auth`.
- **Tests**: `conftest.py` manually wires the DevAuthPlugin in ASGI tests (lifespan
  not triggered). `test_auth.py` expanded with plugin registration and validation tests.
- **Docs**: `04-authentication.md` (en/zh) updated with plugin registration examples.

Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>

* fix(tests): fix trusted mode test

* fix(tests): fix unit test

* fix(cli): remove unexisted transaction observer

* docs: update skills definition

* docs: update skills definition

* docs: update skills definition

* docs: update skills definition

* fix(skills): now we allow viking://agent/skills again, and optimize CLI for skills

* docs(skills): use -p instead of --parent in agent skills examples

Align the `ov skills add` examples in the context-types and viking-uri
docs with the short flag `-p` introduced for `ov skills list/find/show`,
so all four user-facing examples consistently demonstrate the short form
when targeting `viking://agent/skills`.

Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>

* fix(tests): error check for api key

* fix(tests): unit test wait until resource not busy

* fix(tests): unit test wait until resource not busy

* fix(sdk): args form in skills find

* fix(skills): pass target uri in request body

---------

Co-authored-by: claude-sonnet-4-6 <noreply@anthropic.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-06-25 14:36:06 +08:00
87329714dd feat(grep): integrate VikingDB bm25 keyword search for grep engine (#2144)
* feat(grep): integrate VikingDB bm25 keyword search for grep engine

* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)

* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison

* fix(schema): upsert data to vikingdb lack of content

* chore: add benchmark for retrieval

* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs

* fix(benchmark): sub uri args; add report

* refactor: code format by ruff

* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf

* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search

* fix: adjust benchmark scripts

* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls

* refactor: new benchmark

* fix: step1 add resource by real code data

* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex

* optimize (benchmark): adjust keywords and ground truth for testing

* fix: truncate 64KB for content field

* optimize: effectiveness add resource plainly

* optimize: change param use of SearchByKeywords from "keywords" to "query"

* optimize(benchmark): refactor effectiveness scripts

* optimize: ensure raw data for content field

* optimize: fulltext analyzer's stop-words only use symbols

* fix: adapt to new ov cli for benchmark

* optimize: reuse file content to avoid re-read AGFS file

* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts

* optimize: benchmark client timeout

* update README

* fix: rm unused param

* fix: default values in docs

* optimize: increase truncate byte size to 1MB for content field for VikingDB

* fix(logger): harden queued stream logging (#2786)

* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock

When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.

During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.

Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.

Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.

Closes: #2752

* fix(logger): harden queued stream logging

---------

Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
2026-06-24 18:46:02 +08:00
Dico Angeloandqin-ctx 9ec15c8e07 feat(rerank): add configurable HTTP timeout for OpenAI-compatible client (#2784)
* feat(rerank): add configurable HTTP timeout for OpenAI-compatible client

OpenAIRerankClient hardcoded a 30s HTTP timeout, which is insufficient for
local LLM servers (e.g. llama.cpp on ROCm) that incur model cold-start
latency on the first request after inactivity, causing ReadTimeout errors.

Add a `timeout` field to RerankConfig (default 30.0, backwards-compatible)
and thread it through OpenAIRerankClient.__init__, from_config, and the
requests.post call in rerank_batch. The timeout can now be set per-environment
in ov.conf, e.g. "timeout": 120.

Closes #2732

* docs: document rerank timeout config

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-06-23 18:23:23 +08:00
dingben be77b2dac1 feat: implement multi-credential priority call design (#2468)
* feat: implement multi-credential priority call design

Add OrderedCredentialSwitcher for N-credential failover, MultiCredentialVLM and FailoverEmbedder for credential switching, support automatic migration from legacy backup config

* refactor: deduplicate AllCredentialsFailedError and fix logging in switcher

* chore: fix code formatting for lint compliance

* chore: fix remaining code formatting

* fix: update FailoverEmbedder for multimodal API compatibility

* fix

* fix

* refactor: split error classes and fix fail-fast in credential failover

Separate the monolithic PERMANENT error class so the switcher reacts to
the actual root cause:

- 400 (request-level parameter error) -> PERMANENT, fail-fast: same
  request fails on every credential of the same model.
- 401/403/unauthorized/accountoverdue -> new AUTH class: credential-level,
  advances to the next credential in multi-credential mode.
- new CONTENT_SAFETY class (moderation rejections) and INPUT_TOO_LARGE ->
  fail-fast: switching credentials cannot help.

classify_api_error now checks CONTENT_SAFETY before PERMANENT so a
moderation message containing "400" is not misclassified. On fail-fast the
failover wrappers re-raise the original exception (preserving type/info)
instead of wrapping it, so callers can react (e.g. truncate on
input_too_large). AllCredentialsFailedError is reserved for chain
exhaustion. The legacy PrimaryBackupSwitcher also switches on AUTH to keep
existing backup behavior.

* fix: make get_active_index side-effect free in credential switcher

get_active_index() previously mutated state: when a failback threshold was
met it would decrement the active index. Because observability properties
(active_credential_index / active_credential_id) call it, merely reading the
current credential for logging or metrics could accidentally advance the
failback state machine.

Split the concern: get_active_index() is now a pure read, and a new public
maybe_failback() performs the one-step failback (and logs an info line when
the active credential index changes). The request loops in MultiCredentialVLM
and FailoverEmbedder call maybe_failback() at the top of each attempt, so the
failback behavior is unchanged while pure reads no longer have side effects.

* fix: drop global total_max_retries cap from credential failover

The failover loops exited on `idx >= n OR total_attempts >= total_max_retries`
(default 10). With more than 10 credentials, or when failback churn inflated
the attempt count, this could raise AllCredentialsFailedError before every
credential had actually been tried, leaving lower-priority credentials unused.

Remove total_max_retries entirely from MultiCredentialVLM and FailoverEmbedder:
credential exhaustion is now decided solely by reaching the end of the chain
(idx >= n), and per-credential retries remain the responsibility of each
underlying instance via its own max_retries. The aggregated error tuple now
records the failing credential index instead of the attempt counter.

Adds a regression test covering more than 10 credentials all being tried.

* refactor: move model-behavior fields off EmbeddingCredential

encoding_format, model_path, cache_dir, enable_fusion, res_level and
max_video_frames describe how a model runs, not which credential is used.
All credentials of a single embedding model share the same model, so these
belong on the parent EmbeddingModelConfig, not on each credential.

Keeping them on the credential forced a `cred.X or config.X` merge in
_create_failover_embedder, which silently dropped explicit falsy values
(enable_fusion=False, res_level=0, max_video_frames=0) and fell back to the
parent value.

Remove these six fields from EmbeddingCredential and read them directly from
the parent config when building per-credential embedders. id/provider/model/
api_key/api_base/api_version/ak/sk/region/host/extra_headers remain
credential-level.

* fix: raise instead of guessing dimension in FailoverEmbedder

FailoverEmbedder.get_dimension() returned a hardcoded 2048 when the first
embedder had no get_dimension(). That path is reached only when wrapping
sparse embedders, which have no fixed dense dimension; returning a fabricated
2048 silently feeds a wrong dimension to callers (e.g. schema creation).

Delegate to the first embedder and raise AttributeError when it has no
get_dimension(), surfacing the misuse instead of hiding it.

* fix: token usage aggregation in failover wrappers

Two issues in the cross-instance token usage merge:

1. Encapsulation: FailoverVLM / MultiCredentialVLM / FailoverEmbedder reached
   into other instances' private _token_tracker. Add a public token_tracker
   accessor on VLMBase and use it in the VLM mergers.

2. Double counting in FailoverEmbedder: embedders share a process-wide
   singleton token tracker (_get_token_tracker), so all wrapped embedders point
   at the same object. Merging N identical trackers inflated usage N-fold.
   Return a single instance's usage directly instead of merging.

* fix: trip circuit breaker on AUTH errors after error-class split

Splitting 401/403/unauthorized/accountoverdue out of PERMANENT into the new
AUTH class (commit b477b9fd) regressed the circuit breaker: it only tripped
immediately on PERMANENT/QUOTA_EXCEEDED, so auth errors no longer opened the
breaker right away.

For a single embedding instance an auth failure (key invalid / no permission /
overdue) is persistent and retrying is pointless, so the breaker should still
trip immediately. Add ERROR_CLASS_AUTH to the immediate-trip set and update the
classification tests to assert the new AUTH class (403 still trips the breaker).

* fix

* test: bump _last_switch_time when forcing active_idx in ring tests

Without setting _last_switch_time, maybe_failback() retreats to idx 0
immediately because the default 0 timestamp is always older than the
600s timeout, so the unavailable last credential never actually gets
exercised.

* format

* fix

* fix: add dimension valid

* format

* fix: resolve VLM legacy backup primary via _match_provider()

When a legacy config uses ``providers: {openai: {api_key: ...}}`` together
with a ``backup`` VLMConfig, the previous backup-migration branch only
read top-level ``self.provider/self.api_key`` to build legacy-primary,
yielding (provider=None, api_key=None) and an unavailable VLMConfig.

Both primary and backup migration now go through _match_provider() so
``providers``/``default_provider`` based legacy configs are migrated
into VLMCredential with the correct provider/api_key/api_base/etc.

Add regression tests covering primary-providers-dict + backup,
backup-providers-dict, default_provider on backup, and propagation of
extra fields.

* refactor: drop misleading wrapper.is_exhausted from failover wrappers

The ring-retry rewrite of MultiCredentialVLM / FailoverEmbedder does not
call OrderedCredentialSwitcher.on_failure(); each request loops locally
and raises AllCredentialsFailedError on full failure. As a result the
underlying _active_idx is rarely advanced to n, so wrapper-level
is_exhausted would have returned False even when every credential just
failed.

There are no production callers of either property, so remove them
(YAGNI) rather than synthesize an exhausted state from the wrapper side.
The switcher's own is_exhausted stays as a state-machine observation
point used by tests.
2026-06-15 15:37:03 +08:00
Qin Haojie 80ce2df0b8 refactor: remove dead code paths (#2545)
Remove unused Rust CLI helpers and Python code paths that were only reachable through tests or stale re-exports.
2026-06-10 16:34:50 +08:00
Haoyu Zhangandzhanghaoyu.la c675b1732e feat: 计算图片embed向量时支持图文embed,通过配置参数image_vectorization决定embed模式(summary_only,image_only,image_and_summary) (#2460)
Co-authored-by: zhanghaoyu.la <zhanghaoyu.la@bytedance.com>
2026-06-08 10:25:38 +08:00
Zayn Jarvis fd101200f5 Fix Codex VLM response stream parsing (#2355) 2026-06-01 17:24:58 +08:00
DuTao be1e7fc482 feat(eval)Opt vikingbot eval script (#2305)
* 优化评测逻辑

* 兼容 飞书的卡片消息
2026-05-29 20:01:36 +08:00
Qin Haojie c72c8a87e8 fix: apply embedding input limits in embedder (#2266) 2026-05-27 20:23:40 +08:00
Matt Van HornandMatt Van Horn d70635156b fix: register nvidia_nim in litellm_vlm PROVIDER_CONFIGS so vikingbot routes NVIDIA NIM models correctly (#2229)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
2026-05-26 14:36:22 +08:00
civilizedBaboonandCivilizedBaboon <civilizedbaboon> baddbeea75 fix(vlm): initialize Codex async client cache (#2208)
Co-authored-by: CivilizedBaboon <civilizedbaboon>
2026-05-25 11:26:55 +08:00
Qin Haojie cfb19a50ca fix(storage): isolate async clients and semantic refresh (#2168)
Cache async SDK clients per event loop to avoid cross-loop reuse in worker threads.
Move memory vectorization into semantic queue refresh and preserve target sync state for resource updates.
2026-05-21 17:24:52 +08:00
85908e241a Feat/memory link (#2010)
* auto-commit before eval 20260509_181850

* auto-commit before eval 20260509_192618

* update

* auto-commit before eval 20260510_005109

* auto-commit before eval 20260510_011832

* auto-commit before eval 20260510_014114

* auto-commit before eval 20260510_022835

* auto-commit before eval 20260510_025048

* auto-commit before eval 20260510_031034

* auto-commit before eval 20260510_143728

* auto-commit before eval 20260510_172705

* auto-commit before eval 20260510_220133

* auto-commit before eval 20260511_115905

* auto-commit before eval 20260511_121959

* auto-commit before eval 20260511_132120

* auto-commit before eval 20260511_161430

* auto-commit before eval 20260511_163606

* auto-commit before eval 20260511_173943

* auto-commit before eval 20260511_175657

* auto-commit before eval 20260511_224347

* auto-commit before eval 20260511_233109

* auto-commit before eval 20260512_104710

* auto-commit before eval 20260512_111256

* auto-commit before eval 20260512_181905

* auto-commit before eval 20260512_191540

* auto-commit before eval 20260512_192540

* auto-commit before eval 20260512_195710

* auto-commit before eval 20260513_000746

* auto-commit before eval 20260513_004221

* auto-commit before eval 20260513_004656

* refactor: migrate logger calls to tracer in extract_loop modules

Replace logger.warning/error/info with tracer.error/info in extract_loop
related modules for better observability (console + OpenTelemetry spans).

Modules updated:
- agent_experience_context_provider.py (5 replacements)
- extract_loop.py (4 replacements)
- memory_updater.py (9 replacements)
- session_extract_context_provider.py (4 replacements)
- utils/json_parser.py (7 replacements)

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* auto-commit before eval 20260513_123007

* auto-commit before eval 20260513_125305

* auto-commit before eval 20260513_135421

* auto-commit before eval 20260513_141013

* auto-commit before eval 20260513_143455

* auto-commit before eval 20260513_145401

* auto-commit before eval 20260513_163345

* auto-commit before eval 20260514_105906

* auto-commit before eval 20260514_112912

* auto-commit before eval 20260514_120308

* auto-commit before eval 20260514_122022

* auto-commit before eval 20260514_134800

* auto-commit before eval 20260514_135615

* auto-commit before eval 20260514_135818

* auto-commit before eval 20260514_142941

* auto-commit before eval 20260514_162401

* auto-commit before eval 20260514_231859

* auto-commit before eval 20260515_104122

* auto-commit before eval 20260515_122140

* auto-commit before eval 20260515_122942

* auto-commit before eval 20260515_144941

* auto-commit before eval 20260515_154736

* auto-commit before eval 20260515_181643

* auto-commit before eval 20260515_182727

* auto-commit before eval 20260515_183056

* auto-commit before eval 20260515_183652

* auto-commit before eval 20260515_183825

* auto-commit before eval 20260515_202731

* auto-commit before eval 20260516_001144

* auto-commit before eval 20260516_011749

* auto-commit before eval 20260516_015903

* auto-commit before eval 20260516_020505

* auto-commit before eval 20260516_130701

* auto-commit before eval 20260516_144342

* auto-commit before eval 20260516_151043

* Harden memory graph rendering and patch guidance.

Escape embedded graph data for script safety, add a vis-network load guard, tighten graph layout defaults, and clarify SEARCH guidance so patch content stays bound to the target file/page context.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260517_005258

* auto-commit before eval 20260517_012903

* auto-commit before eval 20260517_014036

* auto-commit before eval 20260517_015726

* auto-commit before eval 20260517_024952

* auto-commit before eval 20260517_032518

* auto-commit before eval 20260517_135114

* auto-commit before eval 20260517_143238

* auto-commit before eval 20260517_154858

* auto-commit before eval 20260517_200556

* auto-commit before eval 20260517_215025

* fix: keep memory storage plain and render graph links on display

Store memory bodies as plain text in VikingFS and move link rendering to graph display so repeated writes no longer persist nested markdown links. Also tighten link renderer path handling so cross-user relative paths are rejected and strip_links preserves viking and absolute targets.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260518_001945

* fix: invert selected graph node colors

Make the currently selected memory node use a light background with dark text so it stands out against the dark graph theme.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260518_011327

* update

* auto-commit before eval 20260518_161813

* auto-commit before eval 20260518_165104

* auto-commit before eval 20260518_174259

* update

* auto-commit before eval 20260518_224834

* auto-commit before eval 20260518_233319

* auto-commit before eval 20260518_235712

* auto-commit before eval 20260519_135952

* fix memory patch failure logging

Keep dry-run patch validation from emitting a misleading patch_handler warning, and record skipped field updates from MemoryUpdater where the failure is handled.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260519_213142

* fix(memory): fan out links for shared page ids

Expand _resolve_links so shared page ids resolve across every operation URI instead of collapsing to a single path. Align the page-id and extract-loop tests with the current API contract and the multi-URI link behavior.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260520_195141

* auto-commit before eval 20260520_215911

* auto-commit before eval 20260520_222335

* update

* style(memory): clean up formatter drift

Apply the remaining formatter-driven cleanup in the memory modules so the working tree stays clean before the next behavior changes. This keeps helper signatures and string literals aligned with current lint output.

🤖 Generated with [Aiden x Claude Code]

Co-Authored-By: Aiden

* auto-commit before eval 20260521_130517

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-21 14:58:24 +08:00
Qin Haojie 9846bbca5d fix(vlm): honor LiteLLM native routes (#2111) 2026-05-18 17:11:01 +08:00
Simon Shiandqin-ctx 72b407ffa8 feat(embedder): expose encoding_format for OpenAI/Azure providers (#2092)
* feat(embedder): expose encoding_format for OpenAI/Azure providers

The OpenAI Python SDK 2.x defaults to encoding_format="base64" so the
client can decode embeddings into native float arrays locally. Some
self-hosted or vendor-fronted OpenAI-compatible gateways cannot
deserialize base64 embedding payloads coming back from upstream models
and silently hang for tens of seconds before returning HTTP 500 (e.g.
gateways that wrap providers like Qwen, GLM, Doubao, etc. behind a
strongly-typed Java SDK).

Add an optional `encoding_format` field on EmbeddingModelConfig that
gets forwarded to OpenAIDenseEmbedder. The field is unset by default,
so existing deployments keep the SDK's default behavior. Users hitting
the base64 incompatibility can set:

    "embedding": {
      "dense": {
        "provider": "openai",
        "encoding_format": "float",
        ...
      }
    }

Wiring is intentionally limited to provider="openai" and
provider="azure" — the only two factory branches that route to
OpenAIDenseEmbedder for an actual upstream HTTP gateway. Other
providers either don't expose this knob (volcengine/vikingdb/jina/...)
or run against local stacks where the issue cannot occur (ollama).

* test(embedder): improve encoding_format validation error handling

- Add ValidationError import from pydantic for explicit exception handling
- Update test_rejects_unknown_value to assert ValidationError instead of generic Exception
- Improve test specificity by catching the exact validation error type raised by pydantic models

* docs(embedder): complete encoding_format configuration guide

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-05-18 14:13:03 +08:00
MaojiaSheng f45b22a9a8 fix: backup vlm token usage stat (#2068)
* ov vlm circuit_breaker

* ov vlm failover

* fix: ov vlm failover

* fix: ov vlm failover
2026-05-15 16:31:36 +08:00
MaojiaSheng 81fb3f5c3b feat: Add VLM backup configuration for automatic failover (#2040)
* feat: ov add-resource (spec -L --level), ov stat (return count for dir)

* feat: ov add-resource (spec -L --level), ov stat (return count for dir)

* feat: Add VLM backup configuration for automatic failover

- Add backup field to VLMConfig with recursive backup prevention
- Implement FailoverVLM wrapper class for automatic failover
- Support rate limit, timeout, server error triggers
- Add comprehensive unit tests

* feat: Add VLM backup configuration for automatic failover

* feat: Add VLM backup configuration for automatic failover
2026-05-14 15:38:40 +08:00
Asish Kumar 25e7615c51 feat(vlm): support extra request body passthrough (#1973)
Add VLM configuration support for provider-specific JSON body fields and pass them through to OpenAI-compatible and LiteLLM completion calls.

Document the option and cover OpenAI, LiteLLM, DashScope merge behavior, and legacy flat config migration.
2026-05-12 14:59:52 +08:00
Jiahui Zhou 3bcefb298d fix: unify runtime loggers with openviking logger (#1981) 2026-05-12 11:26:30 +08:00
agent dd7a222c65 fix: rerank (#1933) 2026-05-09 14:51:56 +08:00
yangxinxin-7andClaude Sonnet 4.6 5de357d7cd feat(memory): agent-scope two-phase trajectory/experience memory pipeline (#1880)
* feat(memory): add agent trajectory and experience extraction

Add a two-phase agent memory pipeline with schema-driven trajectory and experience extraction, plus system-managed source trajectory tracking.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(memory): wire agent memory extraction into session flow

Enable the agent memory pipeline behind config and invoke trajectory/experience extraction during session memory processing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(memory): agent memory pipeline — trajectory timestamps, experience merge, concurrent extraction

- Trajectory filenames now include a timestamp suffix (via _stamp_trajectory_names
  in compressor_v2 before apply_operations), so trajectory_name in both the
  filename and MEMORY_FIELDS carries the full timestamped name
- Experience extraction: add merge operation (write generalized + delete_uris),
  fix delete lock conflict (pass lock_handle to viking_fs.rm), and inherit
  source_trajectories from deleted experiences before merge
- Near-duplicate trajectory dedup removed from memory_updater; delete moved
  before write to avoid AGFS sibling lock contention
- session.py: restore user memory extraction and run user + agent memory
  concurrently via asyncio.gather (agent memory gated by agent_memory_enabled)
- directories.py: trajectories and experiences directories added to agent
  memory preset with abstract/overview; cases and patterns removed
- Simplify trajectory/experience YAML descriptions and instructions
- extract_loop: skip refetch for add_only schemas; add logging for URI resolution
  and operation dispatch to aid diagnosis of duplicate experience writes
- demo_agent_memory.py: replace three-round demo with two same-domain rounds
  to specifically test the experience edit path

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(memory): add agent_only flag to prevent user memory from processing trajectory/experience

- Add `agent_only: true` to trajectory.yaml and experience.yaml schemas
- Add `agent_only` field to `MemoryTypeSchema` dataclass
- Parse `agent_only` from YAML in `MemoryTypeRegistry._parse_memory_type`
- Filter out agent_only schemas in both `prefetch` and `get_memory_schemas`
  in `SessionExtractContextProvider`, so trajectory/experience are only
  processed by the agent memory extraction pipeline

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test(memory): add e2e integration test for agent memory two-phase pipeline

- test_trajectory_and_experience_extraction: runs two same-domain sessions,
  asserts Round 1 creates the experience and Round 2 edits it (no duplicate),
  and verifies all trajectory filenames carry a timestamp suffix
- test_no_agent_only_schemas_in_user_memory: unit-level check that
  trajectory/experience schemas are filtered out of SessionExtractContextProvider

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: remove demo_agent_memory.py, replaced by integration test

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(memory): agent memory two-phase pipeline — trajectory + experience extraction

Phase 1 (trajectory): extract execution summaries from conversation, one per business domain.
Phase 2 (experience): prefetch top-5 candidate experiences + source trajectories, single no-tool
LLM call to Update/Replace/Create/Skip.

Key changes:
- AgentExperienceContextProvider: rewrite as prefetch-all + single no-tool call; top-3 candidates
  include source_trajectories for grounding; prefetched_uris tracked to skip refetch check
- AgentTrajectoryContextProvider: remove read tool (was causing hallucination); tighten instruction
- ExtractLoop: fix prefetch URI tracking (old format was broken); guard tool_choice on empty tools
- compressor_v2: deserialize trajectory content before passing to experience phase; restore
  user/agent memory concurrent execution in session.py
- memory_updater: downgrade diff_match_patch ImportError from tracer.error to tracer.info
- volcengine_vlm: trace tool calls and response content separately
- experience/trajectory yaml: refine field descriptions and Reflect section wording
- e2e test: add skipif guard, tracer init, two-iteration loop, persistent demo dir

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(memory): remove unused source trajectory tool and noisy prints

Drop the unused get_source_trajectories memory tool after phase-2 moved to
prefetch-only context, and replace source_trajectory debug prints with tracer logs.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(client): support explicit embedded user for agent memory tests

Allow LocalClient to accept an explicit UserIdentifier and add an integration test covering user+agent agent-memory isolation in embedded mode.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat:test

* feat(agent-memory): prefetch experience files into read_file_contents + cap source_trajectories

- AgentExperienceContextProvider.prefetch now populates _read_file_contents for
  each candidate experience, fixing two issues on the Replace path:
  1. resolve_operations could never find delete_file_contents → old file was never deleted
  2. inherited_traj_uris was always empty → source_trajectories not inherited
  On the Update path this also eliminates the extra _check_unread_existing_files
  LLM round-trip that was previously triggered for every edit.

- Move deserialize_content/deserialize_metadata imports from inline to module top.

- AgentTrajectoryContextProvider.prefetch signature simplified (no unused args).

- _append_trajectories_to_experiences: cap source_trajectories at 5 most recent URIs
  to prevent unbounded growth over many sessions (MAX_SOURCE_TRAJECTORIES = 5).

- e2e test cleaned up: single focused test, remove redundant Replace-path tests,
  filter .abstract.md in _list_non_overview_entries.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(agent-memory): update experience and trajectory memory schema prompts

- experience.yaml: restructure content format from 4-section to 3-section
  (Situation / Approach / Reflect), rewrite rules to emphasize machine
  readability, mutual exclusivity between Approach and Reflect, and
  abstraction mandate for generalization.

- trajectory.yaml: extend content format with explicit Trajectory steps
  (intent + actions + progress) and Fail reason field; add exhaustive
  tracking and tool-call formatting rules.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(agent-memory): harden memory pipeline robustness

- session.py: use gather(return_exceptions=True) so user and agent
  memory tasks fail independently; each side logs its own error and
  falls back to [] instead of losing the other side's results
- compressor_v2: remove redundant rm before write_file in
  _append_trajectories_to_experiences — agfs PUT is atomic overwrite,
  so the prior delete only added a data-loss window; also drop the
  duplicate ExtractContext/MemoryIsolationHandler construction in
  _run_extract_phase and fix its outdated docstring
- extract_loop: remove stray blank line after prefetch tracking block
- memory_updater: remove extra blank line inside class body
- experience.yaml / trajectory.yaml: add missing trailing newlines

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat:fix agent memory test

* chore(memory): rename experience.yaml→experiences.yaml, trajectory.yaml→trajectories.yaml

* chore(memory): rename memory_type experience→experiences, trajectory→trajectories

* chore(memory): remove dead _read_files tracking in extract_loop

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-08 14:53:37 +08:00
Qin Haojie 33de0c6b70 fix(vlm): disable OpenAI SDK retries (#1860) 2026-05-06 11:34:20 +08:00
chenjw 44d3cc41b1 Feat/memory isolation 支持群聊模式 (#1711) 2026-05-06 10:45:06 +08:00
baojun-zhangandMaojiaSheng 17d2c5603e feat(observability): unify observability context && support otel && etc. (#1666)
* feat(observability): unify OTLP metrics export, log/trace context, and telemetry bridging
- - Add OTLP metrics http/grpc exporter that pushes MetricRegistry snapshots
- - Decouple telemetry response payload from telemetry collection; always finish() and bridge summary to metrics
- - Unify observability config under server.observability (metrics/traces/logs siblings); update ov.conf.example and docs (zh/en)
- - Improve log/trace correlation via structured context injection
- - Add/adjust tests for exporter lifecycle, config loader, metrics/telemetry runtime
- BREAKING CHANGE: remove legacy telemetry.* config path; use server.observability.*

* feat(observability): import Status/StatusCode for LogToSpanEventFilter

* feat(observability): fix check issue

* feat(observability): format code

---------

Co-authored-by: MaojiaSheng <shengmaojia@bytedance.com>
2026-04-24 21:32:25 +08:00
Hao ZheandZayn Jarvis 01403312ea feat(vlm): add Codex, Kimi, and GLM VLM support (#1444)
* feat(vlm): add Codex OAuth-backed VLM setup and docs

* fix(codex): address PR review follow-up issues

* feat(vlm): add Kimi and GLM backends

* refactor(vlm): simplify codex auth flow and docs

* fix: update code comments and doctor validation

* chore: update uv.lock after merge

* Refine Codex auth flow and VLM backend integrations

* Take over mirrored Codex auth on refresh

* feat(vlm): refine provider setup and auth flow

* style: format VLM and setup files

* style: fix lint import ordering

* fix(codex): harden auth refresh and disable streaming

* fix(codex): translate tool history and refresh auth safely

* style(lint): fix changed-file ruff violations

* fix(init): refine cloud VLM setup prompts

* style(lint): format setup wizard changes

---------

Co-authored-by: Zayn Jarvis <zhiheng.liu@bytedance.com>
2026-04-22 11:01:23 +08:00
Qin Haojie 38c324bc97 fix(security): clean up code scanning and runtime findings (#1596)
* fix(security): clean up code scanning and runtime findings

Harden path and logging boundaries, remove noisy cleanup issues,
and keep observability failures from breaking runtime flows.

* fix(security): close werewolf and feishu validation gaps

Block the remaining path traversal bypass in the werewolf demo,
and validate Feishu hosts on the main parse() entry point.
2026-04-21 10:06:46 +08:00
Brian Le 20d0f4d6a3 fix(vlm): support OpenAI reasoning-model families (gpt-5, o1, o3, o4) (#1568)
* fix(vlm): translate max_tokens and temperature for OpenAI reasoning models

gpt-5, o1, o3, and o4 families reject `max_tokens` and non-default
`temperature`; they require `max_completion_tokens` and only accept the
server default temperature=1. The raw `OpenAIVLM` backend currently sends
both unconditionally, so `provider: openai` with any reasoning model fails
with 400 `Unsupported parameter`. The sibling `LiteLLMVLMProvider` already
handles this implicitly via `litellm.drop_params = True`; this change brings
`OpenAIVLM` to parity without touching its call shape.

Detect reasoning models by the `gpt-5`/`o1`/`o3`/`o4` prefix, translate the
token-budget key, and omit the temperature override so the API default
(1) applies. Non-reasoning models continue to send `max_tokens` and an
explicit `temperature` exactly as before.

* fix(vlm): pass reasoning_effort for reasoning models to preserve output budget

Live testing revealed that reasoning-model families consume the
`max_completion_tokens` budget for internal reasoning tokens before emitting
any output, leaving tool calls and JSON responses empty. For example,
gpt-5-mini at `max_completion_tokens=512` spent 192 reasoning tokens and
returned 0 tool_calls. Memory extraction observed "LLM returned neither
tool calls nor operations" on every ReAct iteration.

OpenAI exposes `reasoning_effort` (minimal/low/medium/high) on gpt-5 and
o-series models to control how much of the completion budget goes to
reasoning. Default `minimal` for the OpenAI VLM backend preserves output
budget for structured responses like tool calls and JSON ops, which is
what OpenViking's semantic and memory pipelines need. Users who want
deeper reasoning can override via `vlm.reasoning_effort` in ov.conf.

* fix(vlm): default reasoning_effort to 'low' for gpt-5 and gpt-5.4 compat

gpt-5 accepts `minimal|low|medium|high`; gpt-5.4 accepts `none|low|medium|
high|xhigh`. `low` is the only value accepted by both families. Default to
`low` so ov.conf stays portable across model generations without requiring
users to set the knob per model.

* fix(memory): coerce null to empty list in tolerant JSON parser

Memory extraction via OpenAI reasoning models (gpt-5-mini with
reasoning_effort=low) emits `null` for empty list fields like
`"tools": null` in the structured-memory schema instead of `[]`. The
tolerant parser dispatched on `origin_type is list` but fell through to
`parsed_value = value` when the value was None, leaving TypeAdapter to
reject it with "Input should be a valid list".

Return `[]` explicitly when value is None for list-typed fields. Matches
the tolerance already applied to str and dict wrapping, and unblocks
reasoning-model extraction where empty arrays arrive as JSON null.

* style(memory): ruff format json_parser.py

* style(memory): satisfy ruff I001/F401/E721 in json_parser

The PR's edits to this file pull it into ruff's changed-files check
in CI, which surfaces pre-existing lint violations the scanner had
previously only seen on other files:

- I001: sort dataclass/pydantic imports alphabetically.
- F401: drop unused 'parse_obj_as' pydantic import.
- E721: switch 'args[1] == type(None)' to 'args[1] is type(None)'
  in both Optional[T] branches, which is the idiomatic way to
  compare type objects.

Behavior is unchanged; the E721 rewrite is semantically equivalent
because NoneType is a singleton.
2026-04-20 11:23:32 +08:00