Commit Graph
209 Commits
Author SHA1 Message Date
Zonas ZhouandClaude 6e77291265 feat(pdf): refactor MinerU parsing to the official file_parse API (#3953)
* feat(pdf): refactor MinerU parsing to the official file_parse API

* feat(pdf): remove mineru_api_key from configuration and examples

* feat(pdf): preflight MinerU /health during service initialization

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-17 13:48:49 +08:00
Qin Haojie 46f0d60c60 fix(pack): restore account backups without deleting target-only data (#4003) 2026-08-14 16:26:32 +08:00
chenjwandqin-ctx ed1bd4b897 refactor(memory): 统一 V3 提取并提升会话提交与评测稳定性 (#3346)
* fix(memory): disable unsupported tool and skill extraction

* refactor(memory): retire SessionCompressorV2

* docs: design service import cycle fix

* fix(import): break QueueFS service import cycle

* update

* docs: design memory overview lock coverage fix

* fix(memory): cover overview files in update leases

* docs: design session commit default concurrency 50

* perf(queue): raise session commit concurrency to 50

* docs: revise session commit concurrency design

* docs: plan session commit default 8

* perf(queue): default session commit concurrency to 8

* fix(bot): disable cron during eval chat

* docs: design memory link lock stabilization

* docs: plan memory link lock stabilization

* fix(memory): stabilize link update lock coverage

* docs: cover remapped post-group link locks

* docs: design plain-content patch validation

* docs: design first failing patch diagnostics

* fix: report actual failing patch block

* fix(memory): remap replacement links before locking

* fix(bot): include trusted identity in health probe

* test: consolidate memory contract coverage

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-13 21:48:49 +08:00
Zayn JarvisandClaude Opus 5 4920297ccc feat(mcp): add write/edit/tree tools for viking:// as agent working directory (#3936)
* fix(storage): keep non-memory appends free of memory trailers

ContentWriteCoordinator._write_in_place routed every append through
MemoryFileUtils, which strips the existing trailing newline and appends
a reserved MEMORY_FIELDS metadata trailer, even for resource/skill files
where MEMORY_FIELDS is not a reserved format (see content_visibility).
Append to non-memory files now concatenates raw content instead, matching
POSIX append semantics and the documented visibility rules.

* feat(mcp): add write tool with exact-string edit support

Agents could not use viking:// as a working directory through MCP: no
tool could create or update file content. Add a write tool covering full
writes (mode=replace as create-or-overwrite, append, strict create) and
targeted edits (a list of {old_string, new_string, replace_all}
exact-string replacements applied in order, all-or-nothing), following
the Write/Edit conventions of common agent harnesses.

Edits read via read_visible and write back through the content-write
coordinator, so memory metadata trailers are preserved and semantic /
vector re-indexing triggers as with any other write. Parent directories
are created automatically by the storage layer. Descriptions spell out
writable scopes (resources, user memories/resources, agent) and the
wait=true knob for read-after-write search consistency.

Also update the stale tool-count comment in app.py and the MCP tool
tables in the en/zh guides (13 -> 14 tools).

* feat(mcp): add tree tool, split targeted edits into edit tool

tree renders the recursive directory tree under a viking:// URI,
indented by depth with file sizes, for whole-layout orientation;
level_limit/node_limit bound the output and include_abstract adds
per-file summaries. Missing directories report "(nothing under ...)"
instead of an error, matching the read tool's convention.

edit(uri, old_string, new_string, replace_all) takes over the targeted
exact-string replacement that previously lived in write's edits array,
matching the classic Edit tool signature harnesses already train on.
write now only does full-content writes (content + mode), removing the
mutually-exclusive content/edits schema ambiguity. Edits still read via
read_visible and write back through the content-write coordinator, so
memory metadata trailers are preserved and re-indexing triggers as with
any other write.

* test(plugin): update canonical MCP tool list for tree/write/edit

The marketplace test pins the server-registered MCP tool list; add the
new tree, write, and edit tools to fix plugin-tests CI.

* feat(storage): support plain files at the user scope root

Agents treating viking:// as a working directory naturally drop files
like viking://user/zeus-persona.md at the user root, but the write
coordinator only accepted the memories/ and resources/ subtrees.

Two changes make that work:

- Namespace shorthand: a dotted first segment under viking://user/ is a
  file name, not a user id (canonical user ids are dot-free by
  convention), so viking://user/zeus-persona.md now canonicalizes to
  viking://user/<current-user>/zeus-persona.md, matching how the
  reserved memories/resources/skills segments already shorthand.
  Dot-free segments still address an explicit user, and an exact match
  with the current user id still wins.

- Coordinator: plain files directly under the user root (or in
  non-managed subdirectories) anchor their semantic refresh at the
  parent directory. The managed subtrees skills/, peers/, privacy/ and
  sessions/ remain read-only with an actionable error message.

* fix(namespace): narrow user-root shorthand to text-file extensions

Review on #3936 (codex /review-pr) flagged that treating any dotted
segment as a user-root file shorthand would silently re-route canonical
URIs for valid dotted user ids (e.g. alice.smith) into the current
user space. Shorthand now triggers only when the first segment ends
in a common text-file extension; dotted or email-style user ids keep
resolving as canonical user ids. Adds regression tests pinning both
behaviors.

* fix(mcp): resolve user URIs against current user

* test(mcp): pin plain-file writes directly at the user root

The user-root shorthand exists so an agent can drop viking://user/persona.md
into its workspace, but every new test went through an intermediate directory
(viking://user/project/zeus-persona.md), leaving the no-directory shape — the
one that anchors the write coordinator's refresh at the user root itself —
uncovered. Add the missing case.

Also correct the write tool docstring: the create-extension allowlist applies
to any newly created file, including one created by mode="replace" falling
back to create, not only to an explicit mode="create".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 11:11:27 +08:00
agent 00f3738edb feat(usage): emit resource-scoped experience usage records (#3921)
* feat(usage): expand experience tracking and log schema

* fix(usage): preserve experience count event names

* refactor(agent-evolution): use generic OpenViking tools

* fix(usage): capture generic OpenViking tool events

* feat(skills): guide cross-agent experience retrieval

* fix(usage): address generic tool migration review
2026-08-11 22:20:18 +08:00
zihengli cd55ec89a6 feat(assets): support fixed Git commits, explicit targets, and private repository auth (#3703)
* feat/openviking_assets_support_git_commit_id

* feat/openviking_assets_support_to

* feat/private_git_support_watch

* fix: doc_and_ut

* fix: doc_and_ut

* fix: adapt git token url
2026-08-11 14:07:21 +08:00
Qin Haojieandsponge225 b877ababa5 perf(queue): 流式调度语义向量化任务 (#3636)
* perf(queue): stream semantic vectorization tasks

* fix(queue): isolate semantic work context

* feat(config): make parse concurrency configurable

* test(queue): update semantic vectorization fakes

* fix(queue): correct semantic vectorization conflict resolution

---------

Co-authored-by: sponge225 <1670519171@qq.com>
2026-08-10 21:24:11 +08:00
Qin Haojie 7abd6ab249 refactor(client): remove Python embedded mode (#3712)
* refactor(client): remove Python embedded mode

Consolidate Python consumers on the HTTP SDK while keeping shared server and storage capabilities unchanged.

* refactor(client): remove obsolete embedded leftovers
2026-08-10 18:00:00 +08:00
Jiahui Zhou 1d02a72b2b Remove qdrant and opengauss vector backends (#3872) 2026-08-07 19:57:45 +08:00
baojun-zhang bd5cce09c7 feat(pathlock): adjust pathlock config (#3854) 2026-08-07 13:21:30 +08:00
baojun-zhang 758fc7f0fa feat(queuefs): add bounded redis startup stale-recovery sweeps (#3748)
* feat(queuefs): add bounded redis startup stale-recovery sweeps

* fix(queuefs): decouple redis startup recovery from heartbeat with bounded 0/30/60 sweeps
2026-08-06 13:19:20 +08:00
444cc87bf8 feat: OIDC and LDAP as new auth mode for OpenViking (#3708)
* feat: support oidc and ldap auth

* feat: support oidc and ldap auth

* fix: remove heima partner, clean up auth docs, add web-studio unsupported auth banner

- Remove heima from partner list in README (en/zh/ja)
- Remove unsupported env var references (OPENVIKING_AUTH_MODE, OPENVIKING_USERNAME,
  OPENVIKING_PASSWORD) from LDAP auth docs
- Remove temporary switch bash snippets from auth docs
- Fix ldap_password description
- Add web-studio unsupported-auth-mode banner for oidc/ldap servers

* fix: address OIDC/LDAP review comments on auth plugin design

Key changes driven by PR review:

- **Role mapping**: OIDC and LDAP external identities always resolve to
  USER role. Removed map_role() calls and group_membership-based role
  mapping. Admin access is gated by the root API key mechanism only.

- **LDAP credential extraction**: Removed query-parameter-based username/
  password extraction (security concern — passwords in URLs can leak via
  shell history, proxy logs, and monitoring). Clients must use Basic Auth
  header or form data.

- **OIDC identifier sanitization**: Auth0 and other providers may include
  characters like "|" in the `sub` claim. These are now replaced with "_"
  to produce valid OpenViking user identifiers.

- **Dead code removal**: Removed _extract_groups, memberof_attribute,
  require_root_api_key_for_admin, _initialize_api_key_manager, and
  get_request_context_checks from both plugins since they are no longer
  needed.

- **Docs**: Removed query-parameter curl example, memberof_attribute and
  require_root_api_key_for_admin config references.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* feat: support oidc and ldap auth

* feat: support oidc and ldap auth

* fix(auth): bind lazy OIDC imports at module scope

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-06 12:36:37 +08:00
Jiahui Zhou d2056e971d Feat/session auto commit v2 (#3736) 2026-08-05 16:18:49 +08:00
2cc96e393e feat(retrieval): assemble auto-recall context server-side via /search mode="context" (#3534)
* feat(retrieval): assemble auto-recall context server-side via /search mode="context"

Auto-recall assembly lived in every harness plugin: each one searched per
memory type, read hits back one by one, and stitched a context block with its
own budget and degradation rules. The implementations drifted, and the shared
weaknesses showed up in production injections — roughly half of the entries
degraded to a bare URI plus a score, character budgets distorted up to 6x on
CJK text, and adjacent turns re-injected the same memories.

This moves assembly into the server as one round trip. /find stays an unchanged
stateless primitive. /search gains mode="context" (mode="list" is the default
and byte-identical to before), and /recall becomes a thin preset over the same
kernel with its v1 field names folded onto the new contract.

New assembly kernel under openviking/retrieve/context_assembler/:

- Token budgeting with a CJK-aware estimate replaces the character budget.
- detail="auto" fills breadth-first then deepens: every candidate gets a
  readable floor, then overview, then full for high-scoring entries. An
  oversized tier falls back to the previous one instead of being truncated,
  bounded by max_tokens / candidates * 2 per entry.
- Overview extraction dispatches by source: memory files use their leading
  Summary section, code files reuse code_outline signatures, long documents use
  a heading tree plus first paragraph.
- Directory hits start at overview and read their .overview.md sidecar, since
  directories carry no stored abstract; their full tier stays capped at
  overview. v1 injected the sidecar as if it were a whole file.
- Quotas generalize beyond memory types to resources and skills, with purpose
  presets supplying ratios when quotas are absent.
- dedup_turns keeps a per-session ledger at {session_uri}/.recall_log.json so
  every harness inherits cross-turn dedup; exclude_uris remains as the
  stateless fallback.
- Rendering flattens to one <memory uri=... type=... score=... detail=...>
  element per entry. Every tier carries its URI, so the model can always drill
  down through the MCP read tool.
- Query expansion and digest rewriting are opt-in and fail closed: both have
  timeout fuses, and a failed rewrite still returns the unrewritten block.
  Retrieval failures are counted into stats rather than silently yielding an
  empty block.

Plugins now send one context request, falling back to /recall and then to raw
find on older deployments, and cache that outcome so only the first turn pays
for the probe. The tri-state recallRewrite knob chooses between local host-CLI
compression and the server digest, and client-side settings move to a plugin
section in ovcli.conf.

* refactor(retrieval): give context tiers a per-category default

The tier ladder assumed `abstract` is a cheap summary. For memory files it
is not: the memory writer stores the whole stripped body in that scalar
because it doubles as the embedding text, so `abstract` costs the same as
`full` and the ladder runs `uri < overview < abstract = full`. Two of the
model's properties fell out of that: exempting `abstract` from the per-entry
cap let a single entry eat several times the budget, and `detail` — which
only ever set a ceiling — collapsed to two distinguishable behaviours across
its four values, since `auto` already allowed `full` for memory.

Tiers now come from a per-category constant table that treats the storage
shape as a given: `events` starts at overview (the one memory type whose
`# Summary` extraction is a real compression) and may deepen to full on
leftover budget; every other category is served at `abstract`, which for
memory already is the complete file at zero read cost and for resources and
skills is the generated 256-char summary. The table carries the note to move
`events` back to `abstract` once the writer stores a separate summary scalar.

Falling out of that: prefetch now reads only the candidates whose planned
tier needs a body rather than every candidate, `detail` becomes a real pin
(start and ceiling) and additionally accepts a per-category map, and
`full_score_threshold` is gone — leftover budget is spent in score order
instead of behind an absolute threshold the observed score band cannot
support. `auto` is still accepted on the wire as a synonym for "unset".

Assembly fixes found alongside:

- Removing the abstract cap exemption would turn an oversized abstract into
  a bare URI, so it now falls back to overview first — for memory that is a
  cheaper substitute, not a step up.
- Rewrite timeouts were reported as failures on Python 3.10, where
  `asyncio.TimeoutError` is a separate class from the builtin.
- `stats.rewrite_usage` read `token_tracker` off `VLMConfig`, which has no
  such attribute; usage was structurally always null. It now reads the model
  instance's tracker and reports only when the call count moved by exactly
  one, since that tracker is shared.
- A single malformed ledger record made every deduped recall in that session
  fail, and the file was never rewritten, so it could not heal. Records are
  now coerced on read and dropped on the next write, along with records left
  ahead of the clock by an archive rotation.
- Entries served as a bare URI no longer enter the dedup cooldown: they lost
  to budget pressure, not to the reader having already seen them.
- The render envelope only neutralised a literal `</memory>`, so a body could
  forge a sibling entry with its own uri, type and score.
- Flat-mode gathering re-derived the category from the URI, reading
  `viking://resources/backup/memories/events/log.md` as an event.
- Cooled and excluded URIs are compensated with extra rows, so a fully cooled
  bucket falls through to the next-best hits instead of coming back empty.
- `/recall` quotas overlay the v1 bucket defaults again; `{"events": 5}` had
  started dropping the other three buckets.
- The MCP `recall` signature sent its own defaults as if the caller had, which
  resolved a different profile than `POST /recall`; an unknown `detail` value
  raised `KeyError` through the whole call instead of degrading.

* feat(codex): inject profile context on session start

Reuse the shared profile builder for startup, clear, and resume hooks while preserving archive injection and orphan-session status output.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(retrieval): raise rewrite timeout default to 30s

* docs(agents): document low-latency recall settings

* fix(codex): prefer luna as recall compressor fallback

* refactor(plugins): unify recall compression setting

* feat(plugins): enable recall compression by default

* docs(agents): use absolute links in image docs

* fix(retrieval): address context assembly review feedback

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* test: trim redundant context assembly coverage

* fix(retrieval): address second-round context assembly review

- Drop the backticked `/search` from the deprecated-recall row in both API
  overviews. The reference checker scans the whole row after the method cell
  for backticked paths, so it read the description as a route named
  `POST /search` and Build Docs failed on an unknown, undocumented route.
- Accept ovcli.conf's full field set in both Python readers. The file's schema
  belongs to the Rust CLI, which writes `root_api_key`, `output`,
  `echo_command`, `show_progress` and `verbose` and ignores unknown keys; the
  two Python readers had drifted into stricter subsets, so the shipped example
  already failed to load in both. Adding the new `plugin` section to a working
  ovcli.conf would have broken `ov doctor` and every SDK client the same way.
- Return 400 from `mode="context"` for a request `mode="list"` also rejects.
  Retrieval validates query and image_url before searching, and the gather
  fuse swallowed that rejection along with genuine scope failures, so a body
  of `{"mode":"context"}` came back 200 with an empty block instead of the
  documented parameter error. Runtime failures still degrade into
  `stats.retrieval_errors`.
- Let a context request that asks for a server-side digest outlast the
  server's rewrite fuse. The plugin's ordinary 15s request timeout is shorter
  than the 30s fuse, so a rewrite that finished inside its own budget was
  aborted client-side, discarding the whole response — including the
  uncompressed block the server returns when a rewrite fails — and falling
  back to `/recall`. The deadline is only extended when the body actually
  requests a rewrite, and `OPENVIKING_RECALL_CONTEXT_TIMEOUT_MS` /
  `plugin.recallContextTimeoutMs` pins it.

* chore(plugins): sync shared modules into the zcode snapshot

* fix(retrieval): align context quotas and plugin defaults

Restore cross-domain coding recall, reuse authoritative actor resource
scopes, and make bucket quotas the sole width control in purpose mode.
Keep plugin defaults server-owned while preserving explicit legacy limit
settings through quota conversion.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(retrieval): preserve recall compatibility

Restore the deprecated recall threshold default, distinguish successful empty rewrites from compressor failures, and document legacy quota floors across coding-agent plugins.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-05 12:31:21 +08:00
Kchen 8d1d52fe5d 资源导入:支持解析后不拆分文档 (#3645) 2026-08-05 11:34:10 +08:00
Haoyu ZhangandQin Haojie 3f3554256b feat: 支持基于火山方舟的音视频多模态理解 (#3563)
* feat: add audio and video understanding via VLM

* docs: design media resource guards

* fix: bound media staging concurrency

* fix: cap unknown-size media staging

* test: stage media in routing fake

* test: exercise media staging callbacks

* test: trim media understanding coverage

* chore: 清理实现计划文档

* fix: 修复多凭证切换问题

---------

Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-03 16:44:22 +08:00
Hao Zheandzhiheng.liu c4d2b27c64 fix(ov): harden CLI configuration and command behavior (#3552)
* docs: fix stale commands, paths, and provider claims

Sweep findings: D-03, D-04, D-05, D-06, D-07, D-08, D-09. Align setup and API examples with current configuration and CLI behavior.

(cherry picked from commit 2b21ee513b)

* fix(review): correct crypto output flag

Addresses blocking review finding on #3401.

(cherry picked from commit 85bb502709)

* docs(ov): finish stale CLI path cleanup

Complete the #3401 salvage by updating the API-writing templates and the remaining encryption guide examples to the Rust CLI surface.

* fix(ov): make local content operations race-safe

Salvage the safe-I/O portions of OpenViking#3414: reject unsupported local watch requests before upload, atomically create download targets, and write snapshot output before reporting JSON success. Add focused regression coverage.

* fix(ov): validate timeout and node-limit inputs

Salvage and complete OpenViking#3414 by validating every timeout and node-limit surface consistently while preserving config commands as a repair path for invalid persisted values.

* fix(ov): allow explicit help before language setup

Salvage OpenViking#3416 with a narrower contract: only clap-recognized -h/--help requests bypass first-run language selection. Bare command groups, legacy -help, and option values keep the existing gate.

* docs(ov): align session and snapshot command examples

Salvage OpenViking#3419 by correcting positional session and snapshot examples, documenting the canonical observer filesystem command, and keeping fs as a compatible alias.

* fix(ov): honor configured output defaults safely

Salvage and complete OpenViking#3424 with CLI-over-config precedence, runtime validation for normal commands, and a table fallback that leaves config repair commands usable. Also clarify the Python-client versus Rust-CLI upload-mode controls.

* fix(ov): preserve zero node-limit semantics

* ci: skip embedding-dependent resource test without secrets

* fix(cli): validate compile timeout consistently

---------

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
2026-07-31 21:45:53 +08:00
DuTao 94de606c2b fix(bot): make max_tokens optional and configurable (#3654) 2026-07-31 15:16:42 +08:00
zihengli 36d419aaa6 fix:openviking assets import external connector switch (#3634)
* fix:openviking assets import external connector switch

* fix(tests): update quick-start fake embedder compatibility

* fix:openviking assets import external connector switch

* fix:openviking assets import external connector switch
2026-07-31 13:55:41 +08:00
zgy 49b182045b refactor(parser): Refactor code summaries to fixed skeleton-first routing (#3568)
* Refactor code summary skeleton routing

* Simplify code skeleton routing configuration

* Render C tag skeletons as signatures

* Revert "Render C tag skeletons as signatures"

This reverts commit 8e342055f8.

* Simplify fixed code skeleton summary route

* Inline process skeleton rendering

* Simplify code skeleton routing entrypoints

* Fix code summary review issues

* Address final code summary review feedback

* Route failed tags skeletons to LLM fallback

* Restore CUDA and TS extension routing

* Improve code skeleton query coverage

* Route semantic code detection through skeleton support

* Move process skeleton engine into ast package

* Admit skeleton-supported files during directory scan

* Align code summary docs after main merge

* Reduce code skeleton fallback log verbosity

* chore: require grep-ast 0.9.0
2026-07-31 11:38:57 +08:00
baojun-zhang 7c956f23bc feat(config): add compatible default timeout for ragfs pathlock (#3641)
- add storage.agfs.pathlock.lock_timeout_secs
- use pathlock default timeout instead of hardcoded zero in wrapper
- map legacy storage.transaction.lock_timeout when new config is unset
- remote redolog by using  persistent `session_commit` queue.
2026-07-31 11:23:27 +08:00
Qin Haojie b047becb06 fix(auth): enforce account user role boundaries (#3633)
Keep ROOT as a server-owned identity while allowing account admins to promote users within their own account.
2026-07-30 19:48:15 +08:00
zihengli 47bbf7a66a feat(cli): add Openviking asset manifest mode to add-resource (#3358)
* feat: implement server-resolved OpenViking Assets manifests

Add the openviking-assets/1 declaration flow with server-owned configuration parsing and native Rust CLI execution.

- Resolve one flat Manifest against one Catalog through an authenticated server endpoint with strict schema and Git semantic validation.
- Reject recursive includes and unsafe clone URLs; return a resolved plan without submitting resources or running server-side batches.
- Keep local credential aliases, manifest state, dry-run, failure isolation, and per-asset create/sync orchestration in the CLI.
- Generate normalized stable asset identities on the server and remove the CLI direct SHA-1 dependency.
- Update flat examples and add server resolver/API plus Rust CLI coverage.

* feat: implement server-resolved OpenViking Assets manifests

* feat: implement server-resolved OpenViking Assets manifests

* fix(pathlock): tolerate missing lock token after recursive delete

* feat: implement server-resolved OpenViking Assets manifests

* feat: implement server-resolved OpenViking Assets manifests
2026-07-30 15:29:07 +08:00
agent 0ec2bb0ec5 feat: live-reload Agent Evolution and add file usage sink (#3573)
* feat(agent-evolution): reload global switch at commit time

* feat(agent-evolution): expose configured account in status

* test(agent-evolution): cover account in status response

* fix(agent-evolution): align live config reload semantics

* fix(agent-evolution): tolerate non-object live config

* fix(usage-reporter): use snake case count fields

* feat(usage-reporter): add file log sink

* fix

* fix: address live reload and usage sink review findings

* fix(usage-reporter): complete file sink compatibility

* fix: make experience snapshot source unambiguous

* docs(usage-reporter): align count record implementation plan

* fix: address agent evolution review blockers

* fix(usage-reporter): preserve Windows rollover deadline

* fix(usage-reporter): encode file records as JSON envelopes

* fix(usage-reporter): use snake case unique id
2026-07-29 19:25:35 +08:00
Jiahui Zhou 5d1ba45be4 Feat/add resource processing mode (#3566)
* feat: add resource processing mode

* fix: keep semantic artifacts in vectors-only add resource

* test: support processing mode in api test client

* docs: document add resource processing mode

* fix: align processing mode after resource ingestion refactor

* feat: expose processing mode in TypeScript SDK

* fix: preserve add resource compatibility
2026-07-28 20:09:06 +08:00
zihengli a1e468b982 feat(connector): support more git like platform (#3531)
* feat(connector): support more git like platform

* feat(connector): support more git like platform

* feat(connector): support more git like platform

* feat(connector): support more git like platform

* feat(connector): support more git like platform
2026-07-28 19:17:47 +08:00
Hao Zheandzhiheng.liu f0445e0cce docs(build): repair contributor guidance and maintenance tooling (#3551)
* docs: fix broken links and anchors across READMEs and guides

Sweep findings: D-10, D-11, D-12, D-13, D-14, D-15. Restore valid documentation targets and stable cross-page anchors.

(cherry picked from commit e3504d633d)

* docs: correct contributor and release references

Reconstruct the factual parts of draft #3397 against current upstream: use the supported setup wizard, align the repository tree and workflow names with tracked files, document current release paths, and repair the bug-bounty link. Excludes install-policy and subjective content rewrites.

Based-on: b332e19e40
Based-on: c89afb17f2
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>

* build: propagate recipe failures and align CMake minimum

Keep build failures visible, use isolated temporary extraction paths, and enforce the native build's CMake 3.15 floor across all contributor guides. CMake version parsing accepts prerelease and vendor suffixes.

Based-on: 8943a12285
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>

* fix(scripts): surface backfill enumeration failures

Preserve the safety fix from draft #3415 while retaining legacy no-op arguments for existing operational scripts. Deprecated arguments now remain parse-compatible, advertise their status in help, and emit explicit warnings when used.

Based-on: 5516d96048
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>

---------

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
2026-07-28 18:11:33 +08:00
agent 8391d3a758 feat: add global Agent Evolution switch and HTTP usage sink (#3223)
* feat: add per-user agent evolution settings

* simplify Agent Evolution user settings

* fix: preserve agent evolution client compatibility

* feat(snapshot): add path diff API

* feat(snapshot): expose path diff in clients and CLI

* fix(agent-evolution): gate case memory production

* fix(agent-evolution): preserve configuration compatibility

* feat(usage): add built-in HTTP sink

* fix(agent-evolution): address PR review findings

* fix(usage): isolate HTTP outbox by destination

* docs(usage): define CountRecord HTTP mapping

* docs(usage): plan CountRecord HTTP implementation

* feat(usage): emit CountRecord over HTTP

* docs(agent-evolution): design global switch

* docs(agent-evolution): plan global switch migration

* feat(agent-evolution): make production switch global

* fix(agent-evolution): preserve embedded defaults

* test(agent-evolution): cover failed archive policy replay

* docs(agent-evolution): clarify embedded compatibility

* docs(agent-evolution): expose global switch in example config

* refactor(agent-evolution): align global setting terminology

* fix(agent-evolution): preserve session skill extraction

* fix(usage-reporter): capitalize count record keys
2026-07-27 20:09:38 +08:00
DuTao 3319ffd5b7 feat(bot): support VLM credential failover (#3503)
* bot支持多模型

* fix pr

* fix(bot): isolate console VLM credentials
2026-07-27 12:11:48 +08:00
Rocke Dong b62a4d2cfa docs: list all AST extraction languages (#3511) 2026-07-26 14:31:08 +08:00
Colter Dahlberg e97b74d2e5 feat(embedder): support extra_body passthrough in OpenAI embedder config (#3495)
* feat(embedder): support extra_body passthrough in OpenAI embedder config

Adds optional `extra_body` (dict) to the OpenAI dense embedder, merged
into every embeddings.create call. Motivating use case: OpenRouter
provider routing ({"provider": {"sort": "latency"}}) — default routing
shows p90=35s/max=127s tail latency that kills interactive recall
(A/B: sorted routing is consistently sub-second).

Explicit query_param/document_param keys still take precedence on
conflict.

* feat(config): wire extra_body through embedding config layer

Add optional extra_body field to EmbeddingModelConfig and pass it to
OpenAIDenseEmbedder for the openai/azure providers, including the
multi-credential failover merge (parent-level model-behavior field).

* docs(config): document embedding extra_body with OpenRouter routing example

* docs(config): restore concrete host/cors_origins values in EN full schema
2026-07-24 15:07:44 +08:00
Yuanqing ZHAOandYuanqing Zhao 40dd05271c perf(vectordb): micro-batch compatible cuVS searches (#3382)
* perf(vectordb): micro-batch compatible cuVS searches

* fix(vectordb): serialize micro-batch device admission

* perf(vectordb): pipeline warm cuVS micro-batch admission

* docs(cuvs): align micro-batching guidance

* fix(cuvs): warm-batch empty filters

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-23 11:10:52 +08:00
yufeng 82f5383fd5 docs: add Python installer tabs (#3481)
* docs: add Python installer tabs

* docs: clarify installed commands
2026-07-23 10:41:18 +08:00
Yu Zhangandzhangyu.34 a57adb951f fix(rerank): support DashScope nested request/response envelope (#3463)
* fix(rerank): support DashScope nested request/response envelope

OpenAIRerankClient sent a flat request body ({"model", "query",
"documents"}) and parsed "results" at the top level of the response.
DashScope (qwen3-rerank) requires a nested envelope:

  Request:  {"model", "input": {"query", "documents"}, "parameters": ...}
  Response: {"output": {"results": [...]}, "request_id", "usage"}

This caused DashScope rerank to silently fail — the response had no
top-level "results" key, so the client returned None.

Changes:
- Add _is_dashscope() to detect DashScope endpoints by host marker.
- Add _build_request_body() that produces the nested envelope for
  DashScope and the flat body for standard OpenAI/Cohere services.
- Add _extract_results() that reads output.results for DashScope and
  top-level results for standard services.
- Accept both "relevance_score" (singular, DashScope) and
  "relevance_scores" (plural, some providers) in result items.
- Add 13 tests covering host detection, body construction, response
  parsing, end-to-end mocked flows for both providers, plural key
  handling, empty documents, and sparse results.

Fixes #3459

* fix(rerank): detect DashScope protocol by URL path, not hostname

Reviewer noted the previous hostname-based switch broke the documented
qwen3-rerank compatible-api endpoint (/compatible-api/v1/reranks), which
must use the flat OpenAI-style body and top-level results.

Switch to path-based detection: only /api/v1/services/rerank uses the
native nested input/output envelope; everything else (including the
DashScope compatible-api and generic OpenAI/Cohere gateways) keeps the
flat protocol. Rename _is_dashscope -> _uses_nested_envelope for clarity.

Add regression tests covering the compatible-api flat path and reconcile
the existing native-path fixtures to the nested envelope.

* docs(rerank): use qwen3-rerank for compatible-api example

The compatible-api/v1/reranks endpoint uses the flat OpenAI-compatible
protocol; qwen3-vl-rerank is a native-envelope model served at
/api/v1/services/rerank. Align the example model with the endpoint the
implementation selects by URL path.

---------

Co-authored-by: zhangyu.34 <zhangyu.34@bytedance.com>
2026-07-22 19:30:41 +08:00
blakejia 12b4aab13d docs: add S3-compatible storage pitfalls to multi-write guide (#3390)
- Add directory_marker_mode: none to all S3 config examples
- Add S3-compatible storage notes with required fields table
- Add Docker networking guidance for Linux vs macOS/Windows
- Remove private IP addresses from examples (use localhost)
- Apply changes to both English and Chinese versions
2026-07-22 15:34:21 +08:00
yufeng a949517f27 docs: add OpenViking Helper integration and fix stale links (#3445)
* docs: add OpenViking Helper integration

* docs: use mock data in Helper screenshots
2026-07-22 14:14:57 +08:00
yufeng 085aaf832d docs: correct context paths and refine navigation (#3385) 2026-07-21 11:45:17 +08:00
Yuanqing ZHAOandYuanqing Zhao fa19ac0a75 perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest (#3277)
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest

Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.

- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit

Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.

* fix(vectordb): reject stale index replacements

* fix(vectordb): harden bulk rebuild lifecycle

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-16 18:59:16 +08:00
19ca274a24 fix(retrieve): bound reranker input size (#3289)
* fix(retrieve): bound reranker input size

* fix(retrieve): make rerank input limit opt-in

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-07-16 17:19:38 +08:00
chenjw 1b4534effd docs: 本地 trace 文件路径改名 + 补充产生/上传 trace 排查文档 (#3263)
* docs: 本地 trace 文件路径改名 + 补充产生/上传 trace 排查文档

- 将默认本地 trace 路径从 ~/.openviking/data/traces/offline-traces.jsonl
  改为 ~/.openviking/logs/traces.jsonl,不再建 traces 子目录
- 新增 upload_offline_trace.py 脚本,支持上传 JSONL trace 到远端 OTLP
- 上传脚本上传成功后收集并打印 trace_id 列表
- 在中英文 observability guide 中新增「产生本地 Trace 并提交排查」一节,
  覆盖用户开启 local trace、复现问题、提交 JSONL 给管理员、管理员上传的完整流程
- 更新 test_server_config_loader.py 断言

* fix: return non-zero exit code when main file missing or no batches uploaded

Address review feedback: when --file points to a non-existent path, or
when uploaded_batches == 0 (empty file / all lines invalid), the script
now returns exit code 1 instead of silently returning 0.
2026-07-16 10:16:12 +08:00
Jiahui Zhou 1c46d44fbc Fix/reindex preserve owners (#3096)
* fix: preserve reindex content owners

feat: allow trusted admin role assertion

feat: prune orphan vectors during reindex

fix: harden reindex memory body reads

feat: expose reindex prune options in clients

fix(cli): prefer workspace sdk for compat clients

fix: harden reindex prune orphans

* test: align reindex expectations after rebase
2026-07-14 20:38:30 +08:00
huangruitengandhuangruiteng d0305c2f06 fix(mcp): expose context_type filter (#3181)
* fix(mcp): expose retrieval filters

* fix(mcp): limit retrieval scope to context type

---------

Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-14 11:17:23 +08:00
DuTao 41e002fa0b docs(bot): align documentation and tool guidance with current implementation (#3215)
* 优化bot的doc

* 删除无用的security,增加ov的 guides

* ov下增加bot的综述文章
2026-07-13 16:18:07 +08:00
huangruitengandhuangruiteng 3b516d6414 docs: clarify config reload and hybrid search boundaries (#3206)
Co-authored-by: huangruiteng <huangruiteng@bytedance.com>
2026-07-13 13:51:36 +08:00
yufeng 0a14967f6b docs: align MCP references with implementation (#3146)
* docs: align MCP references with implementation

* docs: fix remaining factual drift

* docs: correct remaining API examples

* docs: fix observer status response type
2026-07-11 15:54:07 +08:00
Yuanqing ZHAOandYuanqing Zhao 7e6a0515f9 perf(cuvs): optimize filters, rebuilds, concurrency, and memory (#3092)
* perf(cuvs): fast-path cached native filter routes

* perf(cuvs): parallelize auto filter preflight

* perf(cuvs): add search route telemetry

* test(cuvs): use a valid telemetry vector dimension

* perf(cuvs): reuse native filter preflight results

* perf(cuvs): allow concurrent snapshot searches

* perf(cuvs): coalesce optional background rebuilds

* perf(cuvs): coordinate per-GPU build admission

* perf(cuvs): add opt-in float16 search

* build(cuvs): support vector benchmark harnesses

* perf(cuvs): bound concurrent GPU searches

* perf(cuvs): avoid partial background rebuilds

* fix(cuvs): address rebuild and telemetry review feedback

* fix(cuvs): defer rebuild until index initialization

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-10 17:22:34 +08:00
ByteDanceLiuYang 0d6b53639c fix(vectordb): only write content field for VikingDB backends (#3114)
* fix(vectordb): only write content field for VikingDB backends

* fix: update doc
2026-07-10 15:16:12 +08:00
黄云龙 49b8e3d56c docs: translate Grafana+Prometheus monitoring guide from Chinese to English (#3105) 2026-07-10 15:11:58 +08:00
Qin Haojie 3003ed61d7 feat(retrieval): support image search (#3093)
Add multimodal image vectorization and image query support across the server, SDKs, and CLI.
2026-07-09 16:42:53 +08:00
Yuanqing ZHAOandYuanqing Zhao 39c778c953 feat: add cuVS vector search backend (#2974)
* feat: add cuVS vector search backend

* docs: add agent memory benchmark strategy

* bench: add cuVS index performance harness

* bench: add public ANN dataset tuning

* docs: record preliminary cuVS index results

* docs: clarify warm index latency

* docs: order cuVS before qdrant

* bench: aggregate independent index runs

* bench: order aggregate variants consistently

* docs: add repeatable index scaling results

* bench: add collection lifecycle benchmark

* docs: add collection lifecycle results

* perf: cache prepared cuvs filters

* docs: report prepared filter cache results

* bench: add async vector concurrency benchmark

* bench: aggregate service concurrency runs

* docs: add async concurrency results

* docs: clarify cuVS dtype behavior

* feat: add memory-aware cuVS auto mode

* feat: reuse native filters for cuVS search

* docs: publish cuVS integration plan as Markdown

* fix: route selective filters before cuVS rebuild

* docs: record selective-first routing results

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-07 12:21:10 +08:00