* fix(bot): only require VKE credentials when the TOS storage path needs them
Sweep findings: C-14. Gate AK/SK validation on an actual TOS deployment and drop the unused cluster ID.
(cherry picked from commit 8790ba509f)
* fix(server): apply configured temp_upload.default_mode to uploads
Sweep findings: B-11, D-01. Apply the documented configured upload mode when requests omit it.
(cherry picked from commit a0dc2498e8)
* fix(bot): make one-click Docker deployment generate a working config and port mapping
Sweep findings: C-12. Generate the active ov.conf and keep gateway and Docker ports aligned.
(cherry picked from commit c22af83ab6)
* fix(server): make --bot work and propagate bot flags to workers
Sweep findings: B-09, B-15. Honor the public Bot alias and replay resolved Bot settings in worker processes.
(cherry picked from commit 32a898ca14)
* fix(docker): derive entrypoint/health port from configured server port
Sweep findings: F-06. Keep server startup and every container health check on the same effective port.
(cherry picked from commit 000795c7e3)
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(bot): never silently replace corrupt config, session, or cron stores
Sweep findings: C-03, C-05, C-06. Fail fast on corrupt persisted state before any overwrite.
* test(bot): drop corrupt-state regression tests per review
The three fixes (loader raise on invalid config, session/cron refuse to
overwrite corrupt state) stay; the accompanying tests are removed as not
worth their weight.
* fix(bot): pass sender_name in all channel adapters
Sweep findings: C-01. Supply display-name fallbacks so inbound messages reach the bus.
* fix(bot): make sender_name optional and wire real WhatsApp pushName
Root-cause guard: _handle_message required sender_name positionally, but
InboundMessage.sender_name is str|None=None and context.py already falls back
to sender_id, so the required-ness was an accidental signature/contract
mismatch. Make it optional (reordered after the still-required chat_id/content;
all 11 call sites use keyword args) so no future adapter can crash on it.
WhatsApp: the previous call read data.get("senderName")/data.get("pushName"),
neither of which the bridge ever sends, so it silently always fell back to the
numeric id. Forward baileys' msg.pushName through the bridge payload and read it
in Python, so WhatsApp group chats show real display names like other channels.
Commit 4d34025c added the required `sender_name` parameter to
`BaseChannel._handle_message()` and updated `feishu.py`, but
`slack.py` and `email.py` were not updated in the same change.
This caused a `TypeError` on every inbound message in both channels.
Because Slack SDK swallows exceptions in asyncio listener callbacks,
the error was silent — the bot received events, added emoji reactions,
but never called the LLM or sent any reply.
Fix: pass `sender_name=sender_id` in SlackChannel and
`sender_name=sender` in EmailChannel.
Co-authored-by: scott.kim <scott@ScottMacBookPro.local>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
`ov chat --with-bot` regressed in 0.4.5: a config with only a `vlm` section
(no `bot.agents`) and a `server` whose effective auth mode is `api_key` fails
with `litellm.InternalServerError: OpenAIException - Missing credentials ...
set OPENAI_API_KEY`, even though the user configured e.g. deepseek. 0.4.3
worked.
Root cause: the 0.4.5 ov_server auth rewrite made
`_merge_current_ov_server_config` call `_fill_user_api_key_from_ovcli()` ->
`load_ovcli_config()` whenever the effective auth mode is `api_key`. When the
user only configured `vlm` and has no valid ovcli identity, that raises
`ValueError`, which is re-raised. `load_config()`'s broad
`except (json.JSONDecodeError, ValueError)` then swallows it and silently falls
back to a default `Config()` — model `openai/doubao-seed-2-0-pro-260215`, empty
provider, empty api_key. `_make_provider` takes the legacy LiteLLMProvider
branch on that `openai/*` model with no key -> the OpenAI missing-credentials
error. In 0.4.3 the ovcli lookup only ran in `remote` mode (api_key present),
so a vlm-only config never hit it.
Fix: a missing/invalid ovcli (OpenViking user) identity only affects OpenViking
memory/file tools — it must not abort loading the rest of the bot config. On
`load_ovcli_config()` failure, warn and continue with the user api_key unset
instead of re-raising. Degraded OpenViking auth is still surfaced separately by
validate_openviking_auth(). This restores 0.4.3 behavior for vlm-only configs
while keeping the user-key auto-fill when ovcli is configured.
Adds a regression test covering both the ovcli-fails and ovcli-succeeds paths.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
PR #2944 added `bot.agents.thinking` (default `true`) which mutates
provider reasoning params for VolcEngine (`thinking={"type":"enabled"}`),
DashScope (`extra_body.enable_thinking`), and OpenAI reasoning models
(`reasoning_effort`), but the canonical Vikingbot config tables and
samples in bot/README.md and bot/README_CN.md were not updated.
Document the new default-on option and its per-provider behavior in both
the EN and ZH config reference and sample config so operators can
discover the off-switch (latency/cost/compatibility tuning).
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* fix: stabilize studio identity and streaming chat
* fix: hide unsupported studio terminal commands
* fix: remove unsupported terminal command copy
* fix: run selected terminal suggestion on enter
* fix: group supported terminal commands
* fix: add terminal quick start and history
* fix: scope session visibility by user
* fix: harden bot user scoping
* fix: forward request scoped bot identity
* fix: add terminal quick start translations
* fix: add terminal command group translations
* fix: simplify studio identity scoping
* fix: support api key copy on dev urls
* fix: stop passing agent id to ov http client
* fix: search follow-up memory questions
Add MiniMax-M3 as the new flagship/default recommended model alongside
the existing MiniMax-M2.7 and MiniMax-M2.7-highspeed options.
- Update provider registry comment to list M3 as default with M2.7 as alternative.
- Update bot README recommended models block to suggest M3.
- Add unit tests for M3 keyword match, prefix resolution, and system message merging.
Co-authored-by: octo-patch <octo-patch@github.com>
* feat(tau2/vikingbot): config-driven experience recall + per-domain isolation
Switch tau2 self-improvement behaviour from core-code patches to three
ov.conf flags (recall_exp_first_round_only, exp_recall_limit,
exp_recall_max_chars), so the VikingBot core is unchanged for non-tau2
users.
- context.py: when recall_exp_first_round_only=true, skip per-turn
user+agent memory retrieval and inject experience once on the first
user-turn; accepts explicit agent_id to scope retrieval per domain
- memory.py: exp_recall_limit and exp_recall_max_chars read from config
instead of hardcoded values
- schema.py: add three new OpenVikingConfig fields (all default to
existing behaviour so existing deployments are unaffected)
- ov_server.py: extract _is_session_key() helper to unify the two
places that distinguish session keys from per-domain agent ids;
local mode now respects agent_id for namespace isolation (remote mode
already supported this)
- tau2 runner: pass agent_id= instead of memory_users= to build_messages
- README: document Python >=3.12 prerequisite, correct pip install
extra, clarify that isolation works in both local and remote modes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ov_server): remove redundant underscore check in search_experiences
The "_" in self.agent_id guard was a leftover heuristic to distinguish
domain ids from session keys. Now that _is_session_key() handles that
check via "__", the extra "_" condition is unnecessary and actively
breaks agent ids without underscores (e.g. "airline", "retail").
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: embedding images directly
* fix: skip default values in CLI config serialization, relax ovcli.conf validation
- Rust: add skip_serializing_if to Config/UploadConfig fields to avoid
writing null/default values into ovcli.conf
- Python: change OVCLIConfig/OVCLIUploadConfig model_config from
extra: "forbid" to extra: "ignore" for forward compatibility
- Simplify handle_extra_headers_aliases now that extra fields are ignored
- Add VLMProviderAdapter and integrate VLMFactory into _make_provider
* fix(bot): remove unused OpenAI provider and refresh config docs
Drop the unused OpenAI-compatible provider and its stale exports/tests now that explicit provider configs go through the VLM adapter path. Refresh the bot configuration docs with redacted remote OpenViking examples and clarify gateway/chat usage alongside Feishu configuration.
* revert: drop unrelated embedding changes from PR
Restore the embedder and queuefs files to match the upstream volcengine/OpenViking main branch so this PR only carries the bot provider cleanup and documentation updates.
* revert: align remaining embedding helpers with upstream
Restore the context, embedder base, embedding utils, and local index files to match volcengine/OpenViking main so the PR stays focused on bot-only changes.
* docs: sync cn