Read replicas load the API key store once at startup and never rewrite
it, so a user registered/rotated/removed on the writer stays invisible
(new key -> "Invalid API Key"; removed key -> still accepted).
Add an optional background watcher that polls the shared key store and
reloads the in-memory index only when it actually changes:
- APIKeyManager.reload(): strictly read-only refresh that rebuilds state
and swaps it in atomically, never writing or migrating plaintext keys.
- compute_store_signature(): cheap (path, size, modTime) signature over
accounts.json + every users.json so the watcher skips unchanged polls.
- ApiKeyAuthPlugin starts/stops the watcher behind api_key_watch_enabled
(default off) with api_key_watch_interval_seconds; AuthPlugin.shutdown()
is wired into app shutdown to cancel it cleanly.
Add coverage for reload convergence, read-only/no-migrate guarantees,
uninitialized-store tolerance, signature change detection, and watcher
reload/skip/shutdown behavior.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix: remove heima partner, clean up auth docs, add web-studio unsupported auth banner
- Remove heima from partner list in README (en/zh/ja)
- Remove unsupported env var references (OPENVIKING_AUTH_MODE, OPENVIKING_USERNAME,
OPENVIKING_PASSWORD) from LDAP auth docs
- Remove temporary switch bash snippets from auth docs
- Fix ldap_password description
- Add web-studio unsupported-auth-mode banner for oidc/ldap servers
* fix: address OIDC/LDAP review comments on auth plugin design
Key changes driven by PR review:
- **Role mapping**: OIDC and LDAP external identities always resolve to
USER role. Removed map_role() calls and group_membership-based role
mapping. Admin access is gated by the root API key mechanism only.
- **LDAP credential extraction**: Removed query-parameter-based username/
password extraction (security concern — passwords in URLs can leak via
shell history, proxy logs, and monitoring). Clients must use Basic Auth
header or form data.
- **OIDC identifier sanitization**: Auth0 and other providers may include
characters like "|" in the `sub` claim. These are now replaced with "_"
to produce valid OpenViking user identifiers.
- **Dead code removal**: Removed _extract_groups, memberof_attribute,
require_root_api_key_for_admin, _initialize_api_key_manager, and
get_request_context_checks from both plugins since they are no longer
needed.
- **Docs**: Removed query-parameter curl example, memberof_attribute and
require_root_api_key_for_admin config references.
Co-authored-by: TRAE CLI <noreply@bytedance.com>
* feat: support oidc and ldap auth
* feat: support oidc and ldap auth
* fix(auth): bind lazy OIDC imports at module scope
---------
Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
* feat(agent-evolution): reload global switch at commit time
* feat(agent-evolution): expose configured account in status
* test(agent-evolution): cover account in status response
* fix(agent-evolution): align live config reload semantics
* fix(agent-evolution): tolerate non-object live config
* fix(usage-reporter): use snake case count fields
* feat(usage-reporter): add file log sink
* fix
* fix: address live reload and usage sink review findings
* fix(usage-reporter): complete file sink compatibility
* fix: make experience snapshot source unambiguous
* docs(usage-reporter): align count record implementation plan
* fix: address agent evolution review blockers
* fix(usage-reporter): preserve Windows rollover deadline
* fix(usage-reporter): encode file records as JSON envelopes
* fix(usage-reporter): use snake case unique id
* fix(bot): only require VKE credentials when the TOS storage path needs them
Sweep findings: C-14. Gate AK/SK validation on an actual TOS deployment and drop the unused cluster ID.
(cherry picked from commit 8790ba509f)
* fix(server): apply configured temp_upload.default_mode to uploads
Sweep findings: B-11, D-01. Apply the documented configured upload mode when requests omit it.
(cherry picked from commit a0dc2498e8)
* fix(bot): make one-click Docker deployment generate a working config and port mapping
Sweep findings: C-12. Generate the active ov.conf and keep gateway and Docker ports aligned.
(cherry picked from commit c22af83ab6)
* fix(server): make --bot work and propagate bot flags to workers
Sweep findings: B-09, B-15. Honor the public Bot alias and replay resolved Bot settings in worker processes.
(cherry picked from commit 32a898ca14)
* fix(docker): derive entrypoint/health port from configured server port
Sweep findings: F-06. Keep server startup and every container health check on the same effective port.
(cherry picked from commit 000795c7e3)
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
server.host="0.0.0.0" is a bind/listen address, not a connectable
destination. It was used verbatim as the URL the web-studio bot dials
back into ov-server, producing http://0.0.0.0:1933 — so every bot tool
call failed to connect. Users had to work around it by switching host
to 127.0.0.1, which gives up binding on all interfaces.
Add map_bind_host_to_loopback() and use it when deriving the
client-facing server URL, so wildcard binds resolve to loopback while
real hosts pass through unchanged:
- "0.0.0.0", "", "*" -> 127.0.0.1
- "::", "::0", "[::]" -> [::1]
- bare IPv6 literals (e.g. "::1") stay bracketed for URL syntax
Applied in get_server_url_from_server_data (covers the bot proxy, the
bot config loader, and doctor) and the mcp_endpoint public-base-URL
"listen" fallback. server.host="0.0.0.0" now works for the bot without
the 127.0.0.1 workaround.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(grep): integrate VikingDB bm25 keyword search for grep engine
* fix(grep): address CI review feedback: max-size eviction to _count_cache, use Literal, Split regex alternation into individual keywords for bm25 (max 10)
* fix(schema): use dynamic __version__ for schema_version and handle dev suffixes in version comparison
* fix(schema): upsert data to vikingdb lack of content
* chore: add benchmark for retrieval
* fix(grep): vikingdb return 200 and no results means no matching content, not necessary to fallback to local fs
* fix(benchmark): sub uri args; add report
* refactor: code format by ruff
* optimize: move grep config (engine and switch_to_remote_threshold) to ov.conf
* optimize: auto adapt remote_return_limit by agg API; rm unnecessary params in keywords search
* fix: adjust benchmark scripts
* fix(grep): store full content for BM25; use PathScope depth; reduce redundant API calls
* refactor: new benchmark
* fix: step1 add resource by real code data
* feat(benchmark): split grep benchmark into effectiveness/performance suites with async reindex
* optimize (benchmark): adjust keywords and ground truth for testing
* fix: truncate 64KB for content field
* optimize: effectiveness add resource plainly
* optimize: change param use of SearchByKeywords from "keywords" to "query"
* optimize(benchmark): refactor effectiveness scripts
* optimize: ensure raw data for content field
* optimize: fulltext analyzer's stop-words only use symbols
* fix: adapt to new ov cli for benchmark
* optimize: reuse file content to avoid re-read AGFS file
* optimize: tune grep vikingdb defaults and refresh bm25 benchmark scripts
* optimize: benchmark client timeout
* update README
* fix: rm unused param
* fix: default values in docs
* optimize: increase truncate byte size to 1MB for content field for VikingDB
* fix(logger): harden queued stream logging (#2786)
* fix(logger): replace StreamHandler with QueueHandler+QueueListener to prevent thread deadlock
When log.output='stdout' (default) and the server is managed by systemd,
concurrent log writes can deadlock because logging.StreamHandler holds a
thread lock across stream.flush() which blocks on systemd-piped file I/O.
During session.commit() phase 2, multiple async coroutines (memory
extraction, summarization) concurrently call logger.info()/warning()
with large payloads. The first thread's flush() blocks on the pipe,
while all subsequent threads block on handler.acquire() forever.
This permanently silences the server log and prevents _write_done_file()
from executing, leaving phase 2 hanging without .done.
Fix: use QueueHandler + QueueListener from stdlib logging.handlers
(Python 3.2+). QueueHandler.emit() does queue.put(record) with no lock
or I/O, returning immediately. QueueListener has a dedicated single
thread as the sole consumer touching the real StreamHandler, making
lock contention impossible.
Changes in _create_log_handler(): stdout/stderr branches now create
a shared QueueListener with unbounded queue, returning QueueHandler
instances to callers. _build_standard_handler() delegates formatter
and filter setup to the real handler in the listener thread.
Closes: #2752
* fix(logger): harden queued stream logging
---------
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
---------
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Co-authored-by: njuboy11 <njuboy11@users.noreply.github.com>
* chore: clear unused files
* fix(tests): fix unit test
* refactor(auth): introduce plugin-based authentication architecture
Replace the monolithic `openviking/server/auth.py` with an extensible
plugin-based auth system. This refactor extracts the three built-in modes
(`dev`, `api_key`, `trusted`) into separate `AuthPlugin` implementations,
adds a registry for third-party plugins, and preserves all existing behavior
while enabling custom authentication backends (e.g. LDAP, OIDC, mTLS).
Key changes:
- **New public API**: `AuthPlugin` (ABC) and `register_auth_plugin` decorator.
- **New registry**: `AuthPluginRegistry` supports runtime registration.
- **Built-in plugins**: `DevAuthPlugin`, `ApiKeyAuthPlugin`, `TrustedAuthPlugin`.
- **Config change**: `auth_mode` widened from `Literal` to `str` for custom modes.
- **Validation delegated**: `validate_server_config()` now delegates to the active
plugin's `validate_config()`, preserving existing validation semantics.
- **Router compatibility**: All existing `require_*` decorators and `resolve_identity`
/ `get_request_context` dependencies remain unchanged. Routers import the same
symbols from `openviking.server.auth`.
- **Tests**: `conftest.py` manually wires the DevAuthPlugin in ASGI tests (lifespan
not triggered). `test_auth.py` expanded with plugin registration and validation tests.
- **Docs**: `04-authentication.md` (en/zh) updated with plugin registration examples.
Co-Authored-By: claude-sonnet-4-6 <noreply@anthropic.com>
* fix(tests): fix trusted mode test
* fix(tests): fix unit test
---------
Co-authored-by: claude-sonnet-4-6 <noreply@anthropic.com>
* feat(mcp): progressive single-entrypoint upload for local files
Extends `add_resource` MCP tool to handle local-file paths via a server-orchestrated
two-step flow, eliminating the need for `ov` CLI in sandboxed agent environments
(Claude web, Manus) where local FS is unavailable and CLI install is blocked.
Behavior:
- Remote URL → unchanged.
- Local path → server mints a 6-char base62 token, returns prose Step 1 / Step 2
instructions pointing at /api/v1/resources/temp_upload_signed.
Agent uploads, then re-calls add_resource(temp_file_id=...).
- temp_file_id → resolved against per-tenant subdir, ingested via existing pipeline.
Token: in-memory dict, 10-min TTL, dict.pop doubles as replay protection.
Per-tenant temp-dir isolation ({root}/{aid}/{uid}/{tfid}); legacy CLI uploads
keep flat layout via dual-lookup in resolve_uploaded_temp_file_id.
Public base URL resolves env > config > listen-host fallback (12-factor: runtime
env trumps image-baked config; production deployments behind MCP proxy + nginx
must set OPENVIKING_PUBLIC_BASE_URL since the agent-facing URL is not derivable
from the server's request scope).
* feat(mcp): infer public base URL from request headers + emit fallback hint
Adds a third fallback layer between explicit operator config and listen-host
fallback: capture X-Forwarded-Host / X-Forwarded-Proto / Host headers in the
MCP identity middleware and use them when neither OPENVIKING_PUBLIC_BASE_URL
nor ServerConfig.public_base_url is set.
Resolution order is now: env > config > X-Forwarded-* > Host > listen-host.
The first two are explicit; the rest are inferred. When an inferred source is
used, the add_resource prose response appends a troubleshooting hint asking
the user to set OPENVIKING_PUBLIC_BASE_URL on the server if upload fails —
because inferred URLs can be wrong if the reverse-proxy chain doesn't forward
X-Forwarded headers, or if the server listens on 0.0.0.0.
Documents the variable in docker-compose.yml (commented-out env block) and
in the MCP integration guides (zh + en) — covers when it's required and the
full resolution chain.
* fix(mcp): address Copilot review on PR #1847
- Relax temp_file_id regex from `[a-zA-Z0-9]+` extension to any non-separator
chars, and dedupe to a single TEMP_FILE_ID_RE in local_input_guard. The old
pattern rejected `Path("report.my-file").suffix == ".my-file"` and similar
legitimate filenames, breaking the progressive upload flow.
- Hoist `_resolve_temp_or_path` import to module level in mcp_endpoint
(verified no circular import).
- Add `_is_safe_namespace_component` defense-in-depth at the signed-upload
route so a future code path that mints tokens from less-trusted input
still cannot escape the per-tenant directory.
- Broaden partial-file cleanup to any exception via try/finally + flag,
not just HTTPException — prevents OSError/IO failures from leaving
half-written files behind.
- Scope `_cleanup_temp_files` to the tenant subdir at the signed-upload
route to bound the rglob scan; the legacy `/temp_upload` route still
cleans the root level.
- Add round-trip test for unusual filename extensions (.my-file, .bak~, .中文).
* docs(mcp): reflect server-minted temp_file_id in progressive-upload flow
Post-rebase onto TempUploadStore, the agent no longer learns the temp_file_id
from the MCP prose — the server mints it at upload time and returns it in the
JSON response body. Update both en + zh docs accordingly. Also note that the
signed endpoint shares the same persistence layer as /temp_upload, so
local/shared modes (and multi-worker via shared) apply uniformly.
* fix(mcp): address Copilot review on rebased PR #1847
- Drop `upload_signed_max_bytes` config field. The signed endpoint now relies on
TempUploadStore's streaming `temp_upload.shared_max_size_bytes` check (single
source of truth, fires even when Content-Length is missing/chunked). Map
oversize from InvalidArgumentError back to 413.
- Normalize `X-Forwarded-Host` / `X-Forwarded-Proto` to the first comma-separated
value in `_resolve_public_base_url`, matching the OAuth issuer resolver. Fixes
malformed upload URLs under multi-hop proxy chains.
- Complete the `public_base_url` field comment to reflect all five fallback layers
in the resolver, not just env > field > listen.
- Add `watch_interval` / `to` parameters to the MCP tool tables in both en + zh
integration guides — they were merged in from main's Watch Management API
during the rebase but the table wasn't updated.
Add an opt-in middleware that attaches the request and response bodies
(truncated, content-type filtered) onto the active OpenTelemetry root span,
and surface the URL query string as `url.query`. Off by default — bodies
may contain secrets and high-cardinality content; enable via
`server.observability.dump_body.enabled` and bound payload size with
`max_bytes`.
The dump middleware is registered before the HTTP observability middleware
so it nests inside the trace span (Starlette executes later-registered
middleware first). Streaming, multipart, and binary content types are
skipped, and any capture failure is swallowed so the request path is never
affected.
Co-authored-by: chenpengfei <chenpengfei@bytedance.com>
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
* lisence: change the main lisence from Apache-2.0 to AGPL-v3
---------
Co-authored-by: openviking <openviking@example.com>
Reject unknown ov.conf and ovcli.conf fields with friendly suggestions,
and fail fast during server startup instead of silently ignoring typos.
Co-Authored-By: Claude Opus 4.6
* feat(telemetry): add Prometheus metrics exporter via observer pattern
Adds PrometheusObserver implementing BaseObserver with thread-safe
counters and histograms for retrieval, embedding, VLM, and cache
metrics. Exposes /metrics endpoint in Prometheus text exposition
format. Opt-in via server.telemetry.prometheus.enabled config.
No new dependencies - generates Prometheus text format manually.
* style: use dict.fromkeys per ruff C420
* fix(telemetry): wire PrometheusObserver into data collection and address review feedback
- Hook observer into RetrievalStatsCollector and other data paths
- Register metrics router statically in create_app()
- Remove unrelated with_bot/bot_api_url config changes
* fix(telemetry): measure VLM call duration at call sites
Time each VLM API call using time.perf_counter() and pass the
measured duration through to update_token_usage(), which records
it in the Prometheus histogram.
Previously duration_seconds always defaulted to 0.0 because no
backend passed actual timing data. Now all three backends (OpenAI,
VolcEngine, LiteLLM) measure wall-clock time around the API call
in get_completion, get_completion_async, get_vision_completion,
and get_vision_completion_async.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
Add --workers CLI flag and server.workers config option to allow running
uvicorn with multiple worker processes. This prevents a single slow or
blocking request from stalling the entire HTTP server, including
lightweight endpoints like /health.
When workers > 1, the server uses uvicorn's factory mode with an import
string so each worker process can independently initialize the
application.
Configuration:
- CLI: openviking-server --workers 4
- ov.conf: { "server": { "workers": 4 } }
- Default: 1 (preserves existing behavior)
Closes#464
Co-authored-by: r266-tech <r266-tech@users.noreply.github.com>
* feat: add ov chat command and refactor channel architecture
- Add ChatChannel for interactive chat with User:/Bot: labels and thinking display
- Add SingleTurnChannel for one-off -m mode with minimal output
- Add StdioChannel for JSON-based IPC with Rust TUI
- Rename 'vikingbot agent' to 'vikingbot chat'
- Add Python 'ov chat' command that proxies to vikingbot chat
- Add Rust 'ov chat' command that proxies to vikingbot chat
- Refactor ChannelManager to support both config and direct channel addition
- Update event types for better thinking/tool_call/tool_result display
- Default session key: cli__chat__default
* feishu channel opt
* feishu channel opt
* fix: IM channels only process RESPONSE messages
- Update feishu, dingtalk, discord, email, qq, slack, telegram, whatsapp
- Add filter in send() to skip thinking/tool_call/tool_result messages
- Only process is_normal_message (RESPONSE type)
* feat(tracing): add abstract trace decorator for session-aware observability
- Add vikingbot/utils/tracing.py with backend-agnostic @trace decorator
- Use ContextVar for session_id propagation through nested calls
- Implement lazy binding to Langfuse via propagate_attributes
- Update AgentLoop._process_message() to use @trace decorator
- Simplify langfuse initialization logging in commands.py
- Add session_id parameter to litellm_provider.chat()
- Clean up redundant code in utils/helpers.py
The trace decorator abstracts observability concerns, allowing future
switching between Langfuse, OpenTelemetry, or other backends without
modifying business logic.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
* feat: add tracing base on langfuse
* 1. feishu channel opt
2. support multi users
* 1. feishu channel opt
2. support multi users
* 1. feishu channel opt
2. support multi users
* fix(langfuse): use module-level propagate_attributes from SDK v3
The propagate_attributes function is a module-level export in Langfuse
Python SDK v3, not a method of the Langfuse client instance.
- Import propagate_attributes from langfuse module
- Remove misleading warning when propagate_kwargs is empty
- Reduce log noise by changing info logs to debug
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor: unify workspace_id naming and improve tracing integration
Standardize terminology and clean up tracing/Langfuse integration:
- Rename sandbox_key to workspace_id across agent, memory, and tools
- Delete deprecated langfuse_decorator.py (superseded by tracing.py)
- Fix Langfuse v3 SDK propagate_attributes usage (module-level function)
- Improve session_id extraction with better signature inspection
- Reduce log noise in Langfuse attribute propagation
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* 1. feishu channel opt
2. support multi users
* feat(tracing): add user_id extraction support for Langfuse
Add extract_user_id parameter to @trace decorator to enable user
tracking in Langfuse. This allows grouping traces by user in the UI.
- Add extract_user_id parameter to @trace decorator
- Extract user_id from InboundMessage.sender_id
- Pass user_id to Langfuse propagate_attributes
- Update loop.py to use new lambda style for session_id extraction
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix http server
* fix http server
* fix http server
* fix(langfuse): change propagate_attributes log level to info
* feat(tracing): add @observe decorator to create Langfuse traces
* fix(tracing): apply @observe at decoration time, not runtime
* fix http server
* fix(tracing): add detailed diagnostics for Langfuse client status
* fix(langfuse): add diagnostic logging for client initialization
* fix(langfuse): add diagnostic logging for config check
* 飞书chat
* opt http client
* docs(readme): add Langfuse observability configuration guide
* opt http client
* opt http client
* opt http client
* opt http client
* opt http client
* fix(langfuse): fix token reporting to use usage_details format
- Change usage to usage_details for Langfuse v3 SDK compatibility
- Add support for cache_read_input_tokens (OpenAI/Anthropic prompt caching)
- Add logger import and debug logging for token reporting
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* opt http client
* eval command
* eval command
* eval command
* md
* cleanup(tests): remove obsolete test suite and related docs
Remove the entire legacy test suite including:
- Unit tests (test_agent, test_bus, test_channels, test_config)
- Integration tests (test_agent_e2e)
- Test fixtures, utilities, and OpenSpec config
- Test runner tools (tester/)
These tests were outdated and no longer maintained. Future testing
should use a modern testing framework.
* docs(readme): update configuration paths and chat examples
- Update default config path to ~/.openviking/ov.conf
- Add interactive chat mode examples (--no-markdown, --logs flags)
- Remove VKE deployment guide section
- Update Docker volume mount paths
* docs(agent): add comprehensive docstrings to core classes
Add detailed Google-style docstrings to:
- AgentLoop.__init__() - parameters and examples
- AgentLoop._publish_thinking_event() - event publishing
- ToolContext - all attributes documented
- Tool base class - complete usage example
Improves code maintainability and IDE support.
* feat(server): add bot API proxy support and CLI integration
Server changes:
- Add --with-bot flag to enable Bot API proxy
- Register bot_router at /bot/v1 prefix
- Add bot_api_url configuration option
- Initialize bot proxy in bootstrap process
CLI changes:
- Update ov chat endpoint to /bot/v1/chat
- Fix UTF-8 input handling
- Add endpoint configuration via env var
* feat(core): improve agent tools, tracing and session management
- Enhance tool registry with better error handling
- Update OpenAPI channel configuration
- Improve session manager with better state handling
- Enhance Langfuse tracing integration with diagnostic logging
* feat(cli): add agent tools and improve CLI commands
- Enhance CLI commands with new agent tool integration
- Move plugin analysis doc to docs directory
- Add RFC for OpenViking CLI ov-chat command
- Add server restart script
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: DuTao <dutao.1786@bytedance.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: fujiajie.1030 <fujiajie.1030@bytedance.com>
* fix: remove await asyncio and call agfs directly
* feat: mv cli out of openviking
* refactor: mv cli out of openviking
* refactor: mv cli out of openviking