Commit Graph
98 Commits
Author SHA1 Message Date
996128abcc fix(session): split JSONL on newline only, not Unicode line boundaries (#3984) (#3988)
* fix(session): split JSONL on newline only, not Unicode line boundaries (#3984)

* test(session): consolidate unicode JSONL regression coverage

---------

Co-authored-by: mac <bishopapril850965@yahoo.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-14 14:46:43 +08:00
chenjwandqin-ctx ed1bd4b897 refactor(memory): 统一 V3 提取并提升会话提交与评测稳定性 (#3346)
* fix(memory): disable unsupported tool and skill extraction

* refactor(memory): retire SessionCompressorV2

* docs: design service import cycle fix

* fix(import): break QueueFS service import cycle

* update

* docs: design memory overview lock coverage fix

* fix(memory): cover overview files in update leases

* docs: design session commit default concurrency 50

* perf(queue): raise session commit concurrency to 50

* docs: revise session commit concurrency design

* docs: plan session commit default 8

* perf(queue): default session commit concurrency to 8

* fix(bot): disable cron during eval chat

* docs: design memory link lock stabilization

* docs: plan memory link lock stabilization

* fix(memory): stabilize link update lock coverage

* docs: cover remapped post-group link locks

* docs: design plain-content patch validation

* docs: design first failing patch diagnostics

* fix: report actual failing patch block

* fix(memory): remap replacement links before locking

* fix(bot): include trusted identity in health probe

* test: consolidate memory contract coverage

---------

Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-13 21:48:49 +08:00
Qin Haojie c96fbcb85f fix(session): 避免后序归档阻塞队列 Worker (#3944)
* fix(session): 避免归档任务阻塞队列 Worker

后序归档不再占用 Worker 等待前序任务,并根据 QueueFS work 识别和跳过无法恢复的孤儿归档。

* fix(session): 仅调度队首归档任务

同一 Session 只将最早的未完成归档放入 QueueFS,后续归档在前序结束后再依次入队,并兼容升级前已入队任务。

* fix(session): 降低队首归档调度的存储读取

用 QueueFS 运行时索引判断 Session 是否已有归档任务,正常完成后直接调度相邻 Archive,避免每次 Commit 和任务结束都扫描完整历史目录。

* fix(session): 恢复每个归档任务独立入队
2026-08-12 17:35:46 +08:00
Qin Haojie 7abd6ab249 refactor(client): remove Python embedded mode (#3712)
* refactor(client): remove Python embedded mode

Consolidate Python consumers on the HTTP SDK while keeping shared server and storage capabilities unchanged.

* refactor(client): remove obsolete embedded leftovers
2026-08-10 18:00:00 +08:00
Jiajie - He/him/his 6a252eba18 fix(session): validate archive names when loading history (#3922) 2026-08-10 17:48:49 +08:00
7f6085a2f9 feat(memory): support event tag filtering (#3850)
* feat(memory): support event tag filtering

Add session-level default event tags, commit-time overrides, durable queue propagation, and first-write vector index tagging. Include config update APIs and coverage for serialization, concurrency, extraction, and HTTP behavior.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* feat(memory): expose event tags in SDKs and CLI

Add session default tag configuration, config updates, and commit-time event tag overrides across embedded Python, standalone Python, TypeScript, Go, and the Rust CLI. Preserve explicit empty-tag semantics and document each public interface.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(sdk): align legacy session tag APIs

Forward commit-time event tags through the legacy Python HTTP shims and align BaseClient session signatures without adding a new abstract-method requirement for existing subclasses.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* feat(session): allow updating auto-commit policy

Extend PATCH session config to atomically update event tags and auto-commit settings. Merge policy objects by field, use explicit null to disable automatic commits, preserve omitted fields, and expose the contract across SDKs and CLI.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* fix(session): align session config interfaces

Replace the generic session create config JSON flag with explicit event-tag and auto-commit options. Preserve omitted, object, and null auto-commit semantics across HTTP, embedded clients, SDKs, and CLI, reject ambiguous null policy fields, and handle nullable event configuration consistently.

Co-authored-by: TRAE CLI <noreply@bytedance.com>

* test(session): trim redundant event tag tests

---------

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: qin-ctx <qinhaojie.exe@bytedance.com>
2026-08-10 11:58:02 +08:00
t0sakiandTRAE CLI 0ab48f96fc fix(session): recover partial capture sessions (#3820)
Treat messages.jsonl as the materialization boundary for session-aware recall,
repair partial session roots during the existing authoritative append path,
and preserve Claude capture cursors when writes never reach the server.
Also replay explicitly retryable storage conflicts across memory plugins.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
2026-08-06 14:48:42 +08:00
Jiahui Zhou d2056e971d Feat/session auto commit v2 (#3736) 2026-08-05 16:18:49 +08:00
agent 1494dc3052 feat(agent-evolution): improve provenance, history, and account settings (#3695)
* feat(agent-evolution): record trajectory mapping in snapshot commits

* feat(agent-evolution): apply global switch immediately

* fix(snapshot): hide memory fields in visible history

* fix(agent-evolution): preserve batch snapshot provenance

* feat(agent-evolution): scope runtime settings by account

* fix(agent-evolution): serialize experience snapshot finalization
2026-08-03 20:52:46 +08:00
Hao Zheandzhiheng.liu af914ca27b fix(core): harden privacy and background failure handling (#3548)
* fix(server): stop exporting raw query strings and buffering zip responses in observability

Sweep findings: B-03, B-13. Prevent query secrets from reaching traces and keep ZIP responses streaming.

(cherry picked from commit d8ac3dc33b)

* fix(session): tolerate missing/corrupt archive in Phase-2 replay

(NotFoundError / _ArchiveMessagesCorruptError) on a missing or corrupt
archive messages.jsonl instead of returning []. That PR added skip-on-
missing tolerance to the read path (_get_uncovered_archive_messages) and to
resume_queued_commit, but not to the Phase-2 commit replay path
(_prepare_phase2_archive_messages), which calls _read_archive_messages
unguarded while rolling earlier failed archives into the current commit.

Consequence: a terminally-failed earlier archive whose messages.jsonl is
missing/corrupt (legacy "no messages" terminal data, or produced by #3417's
own archive_read terminal path) makes every subsequent commit's Phase-2
extraction raise -> caught by _run_memory_extraction's except -> the current
archive is terminal-failed too. Because the poisoned archive is only removed
from replay once "covered" (which requires a later archive to complete), and
no later archive can ever complete, the session's memory extraction is
permanently poisoned. Raw messages are safe, but extraction is stuck.

Fix: wrap the replay-loop _read_archive_messages call in the same tolerance
_get_uncovered_archive_messages already uses -- skip + warn on not-found
(_is_storage_not_found) and on _ArchiveMessagesCorruptError, re-raise real
storage failures. The skipped archive stays in covered_failed so the current
archive's .done marks it covered, clearing the poison permanently.

Adds a regression test asserting the replay skips a failed archive with a
missing messages.jsonl (and marks it covered) instead of raising, and that a
real storage failure still propagates.

Follow-up to #3417.

(cherry picked from commit 5b8ec9e68a)

* fix(client): align client surfaces without leaking memory metadata

Reconstructs the client-parity work from upstream PR #3439 on current main and strips reserved memory metadata before line slicing in both embedded and HTTP reads.

Based-on: 48b411d58c

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>

* fix(index): propagate semantic vectorization failures safely

Reconstructs upstream PR #3437 on current main, carries enqueue failures through SemanticDagExecutor, and drains the attempt's embedding tracker before retry-visible failure propagation.

Based-on: 02387deb09

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>

* fix(core): close privacy and embedding failure gaps

* fix(memory): strip repeated metadata trailers

* fix(core): close public memory visibility gaps

* ci: skip embedding-dependent resource test without secrets

---------

Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
2026-08-03 16:07:35 +08:00
Qin Haojie 9295a3b955 fix(session): remove actor scope from session lifecycle (#3661)
Keep sessions user-scoped and remove legacy agent fallback that could leak an actor view into commit memory writes.
2026-07-31 19:52:19 +08:00
DuTao 6362681410 fix(session): preserve cumulative checkpoints across archives (#3647) 2026-07-31 16:10:20 +08:00
Qin Haojie fd42b1ad92 feat(tasks): support task cancellation (#3577)
* feat(tasks): support task cancellation

* refactor(tasks): scope cancellation to current user

* feat(cli): support task cancellation

* refactor(tasks): make cancellation queue-aware

* refactor(tasks): simplify cancellation bookkeeping

* test: remove task cancellation coverage

* refactor(tasks): trim cancellation coordination

* fix(tasks): contain cancellation to owned work

* feat(tasks): persist resource source metadata

* fix(tasks): handle cancelled work consistently

* refactor(tasks): make completion queue-aware

* fix(tasks): persist terminal state before queue ack

* test(tasks): remove added lifecycle tests

* docs(tasks): document task cancellation
2026-07-30 20:34:27 +08:00
Eurakaxun 44c6df2622 perf: retrieval, import, LangChain, and session-context optimizations (#3569) 2026-07-30 10:10:11 +08:00
baojun-zhang 2f9451231e refactor(pathlock):using rust implement instead python (#3602)
* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):optimize unit test code

* refactor(pathlock):optimize encryption create func

* refactor(pathlock):avoid releasing handoffed pathlock on enqueue errors

* fix(pathlock): use owned lease capability and handle S3 create-new 409 as conflict

* fix(ragfs): keep original FsContext for multi-write metadata

* fix(pathlock): resolve lease coverage and CAS handling issues

- detect S3 conditional conflicts from structured service errors
- pass transaction leases when deleting skill roots
- let temp cleanup acquire locks for temp paths
- disambiguate cache and pathlock providers in cache tests
- update temp cleanup lease assertions

* fix(ragfs): bypass pathlock for multi-write metadata

* fix(ragfs): revert pathlock fail-fast design

* fix(ragfs):fix(ragfs): use non-blocking fcntl locks for localfs CAS

* fix(ragfs): serialize heartbeat lease refresh with release and report real conflict kind

* fix(ragfs): preserve conflict kind snapshot and drop unused test scaffolding

* fix(ragfs): preserve conflict kind snapshot and drop unused test scaffolding
2026-07-29 19:45:34 +08:00
agent 0ec2bb0ec5 feat: live-reload Agent Evolution and add file usage sink (#3573)
* feat(agent-evolution): reload global switch at commit time

* feat(agent-evolution): expose configured account in status

* test(agent-evolution): cover account in status response

* fix(agent-evolution): align live config reload semantics

* fix(agent-evolution): tolerate non-object live config

* fix(usage-reporter): use snake case count fields

* feat(usage-reporter): add file log sink

* fix

* fix: address live reload and usage sink review findings

* fix(usage-reporter): complete file sink compatibility

* fix: make experience snapshot source unambiguous

* docs(usage-reporter): align count record implementation plan

* fix: address agent evolution review blockers

* fix(usage-reporter): preserve Windows rollover deadline

* fix(usage-reporter): encode file records as JSON envelopes

* fix(usage-reporter): use snake case unique id
2026-07-29 19:25:35 +08:00
baojun-zhang 1841dfed81 Revert "refactor(pathlock):using rust implement instead python (#3557)" (#3597)
This reverts commit 6b538db569.
2026-07-29 11:31:41 +08:00
baojun-zhang 6b538db569 refactor(pathlock):using rust implement instead python (#3557)
* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):using rust implement instead python

* refactor(pathlock):optimize unit test code

* refactor(pathlock):optimize encryption create func

* refactor(pathlock):avoid releasing handoffed pathlock on enqueue errors

* fix(pathlock): use owned lease capability and handle S3 create-new 409 as conflict

* fix(ragfs): keep original FsContext for multi-write metadata

* fix(pathlock): resolve lease coverage and CAS handling issues

- detect S3 conditional conflicts from structured service errors
- pass transaction leases when deleting skill roots
- let temp cleanup acquire locks for temp paths
- disambiguate cache and pathlock providers in cache tests
- update temp cleanup lease assertions

* fix(ragfs): bypass pathlock for multi-write metadata

* fix(ragfs): revert pathlock fail-fast design
2026-07-29 11:08:42 +08:00
agent 8391d3a758 feat: add global Agent Evolution switch and HTTP usage sink (#3223)
* feat: add per-user agent evolution settings

* simplify Agent Evolution user settings

* fix: preserve agent evolution client compatibility

* feat(snapshot): add path diff API

* feat(snapshot): expose path diff in clients and CLI

* fix(agent-evolution): gate case memory production

* fix(agent-evolution): preserve configuration compatibility

* feat(usage): add built-in HTTP sink

* fix(agent-evolution): address PR review findings

* fix(usage): isolate HTTP outbox by destination

* docs(usage): define CountRecord HTTP mapping

* docs(usage): plan CountRecord HTTP implementation

* feat(usage): emit CountRecord over HTTP

* docs(agent-evolution): design global switch

* docs(agent-evolution): plan global switch migration

* feat(agent-evolution): make production switch global

* fix(agent-evolution): preserve embedded defaults

* test(agent-evolution): cover failed archive policy replay

* docs(agent-evolution): clarify embedded compatibility

* docs(agent-evolution): expose global switch in example config

* refactor(agent-evolution): align global setting terminology

* fix(agent-evolution): preserve session skill extraction

* fix(usage-reporter): capitalize count record keys
2026-07-27 20:09:38 +08:00
Zayn Jarvis 2c80d44349 fix(kernel): stop conflating storage failures with not-found (#3417)
* fix(kernel): stop conflating storage failures with not-found

Sweep findings: A-03, A-07, A-11, A-12, B-07. Preserve storage and parse failures instead of reporting missing or empty state.

* fix(review): restore archive failure handling

Addresses blocking review finding on #3417.

* fix(review): terminalize corrupt archive records

Addresses blocking review finding on #3417.

* test: adapt pending-archive-skip test to refactored archive scan

Rebase onto main (#3380 turn-aware retention) changed archive refs to carry
an archive_id; update the test mock's _list_archive_refs return so the missing
pending archive still routes through _get_uncovered_archive_messages and is
skipped (not raised).
2026-07-24 17:57:13 +08:00
DuTao 0ab85f450a feat(session): add turn-aware retention and reliable archive recovery (#3380)
* 优化OpenViking的 session compact逻辑,active message 改为turn,压缩 assistant,保留完整user。
详见RFC:https://github.com/volcengine/OpenViking/discussions/3330

* Vikingbot 使用 ov turn session

* fix pr comment

* 更新文档

* fix pr issue
2026-07-24 14:26:46 +08:00
黄云龙 c8fd73aaca fix(session): return in-flight archive messages in get_session_context (#3129) (#3149)
* fix(session): return in-flight archive messages in get_session_context (#3129)

Seed latest_completed_index from 0 instead of commit_count so that
archives whose Phase 1 has completed (messages written, commit_count
advanced) but whose Phase 2 is still running (.done not yet written)
are treated as pending rather than already completed.

The release/0.3.x implementation seeded from 0 and did not have
this bug; the regression was introduced when commit_count was
adopted as the seed value.

* test(session): deterministic regression for pending archive context (#3129)

Replace the monkey-patched commit_async test with a direct
filesystem-state test that sets up the post-Phase-1 archive
(messages.jsonl present, commit_count advanced, no .done marker)
and asserts get_session_context still surfaces the archived
messages.

This is deterministic regardless of the queue-worker architecture
because it creates the archive state directly via the mock AGFS
and loads a fresh session from that state, never calling
commit_async or touching the session compressor.
2026-07-20 08:18:25 +08:00
agent ccc271ff27 feat: add extensible usage reporting (#3222)
* feat: add extensible usage reporting

* fix: harden usage reporter lifecycle

* fix: scope experience usage events correctly

* refactor: generalize usage event schema

* docs: design Codex experience memory tools

* docs: plan Codex experience memory tools

* feat: add Codex experience memory tools

* docs: remove temporary Codex implementation plans

* fix: capture Codex MCP tool parts

* fix: harden experience usage reporting

* fix: reload credentials for local MCP tools

* fix: preserve MCP tool-level errors

* fix: bound synchronous sink shutdown

* fix: enforce sink shutdown timeout

* fix: enforce experience tool contracts

* fix(usage-reporter): keep tool schemas and replay ids stable

* fix(usage-reporter): reject unidentifiable tool events
2026-07-16 15:33:27 +08:00
Qin Haojie b04342798a feat(session): make generated IDs readable (#3287) 2026-07-16 11:30:24 +08:00
Qin Haojie d47f2106ee refactor: remove unused and deprecated APIs (#3272)
Delete dead compatibility paths and test-only helpers so unsupported APIs do not remain as accidental contracts.
2026-07-16 10:49:56 +08:00
Qin Haojie 4a54dbeccf fix(session): resume commits after restart (#3254) 2026-07-15 16:01:17 +08:00
chenjwandClaude fd73dcf23a Feat/自进化(经验记忆)框架重构 (#2503)
* Add trajectory experience learning redesign doc

* auto-commit before eval 20260607_043406

* auto-commit before eval 20260607_044129

* auto-commit before eval 20260607_123706

* auto-commit before eval 20260607_125514

* auto-commit before eval 20260607_133737

* auto-commit before eval 20260607_144649

* auto-commit before eval 20260607_154631

* Refine streaming memory train merge pipeline

* Refine session train policy optimization architecture

* Add VikingMem ARA paper analysis

* Force merge for mixed extraction memory patches

* auto-commit before eval 20260608_134426

* auto-commit before eval 20260608_142108

* auto-commit before eval 20260608_153909

* auto-commit before eval 20260608_154845

* auto-commit before eval 20260608_170143

* update

* auto-commit before eval 20260611_150946

* auto-commit before eval 20260611_153933

* auto-commit before eval 20260611_154251

* Fix tau2 reward wrapper call

* auto-commit before eval 20260611_193803

* auto-commit before eval 20260611_194939

* update

* auto-commit before eval 20260612_111029

* auto-commit before eval 20260612_112104

* auto-commit before eval 20260612_122603

* auto-commit before eval 20260612_123359

* auto-commit before eval 20260612_124303

* auto-commit before eval 20260612_130257

* Fallback peer routing to first conversation peer

* Route self memory through self peer sentinel

* Keep self sentinel out of peer memory paths

* auto-commit before eval 20260612_154051

* auto-commit before eval 20260612_154850

* auto-commit before eval 20260612_161633

* auto-commit before eval 20260612_184022

* auto-commit before eval 20260612_201845

* auto-commit before eval 20260612_202637

* auto-commit before eval 20260612_204040

* auto-commit before eval 20260612_224621

* Fix locomo progress column initialization

* Add memory field versioning

* auto-commit before eval 20260612_232318

* Simplify locomo progress display

* Remove locomo progress elapsed time

* Batch streaming memory merges by group

* Derive patch merge language from patches

* Detect patch merge language from updated files

* auto-commit before eval 20260613_004339

* auto-commit before eval 20260613_005835

* Persist memory update trace id

* auto-commit before eval 20260613_012722

* auto-commit before eval 20260613_013923

* auto-commit before eval 20260613_014708

* Enforce peer scope after memory merge

* auto-commit before eval 20260613_033402

* auto-commit before eval 20260613_151931

* auto-commit before eval 20260613_164217

* chore: raise vikingbot eval parallelism

* chore: tune vikingbot parallelism to 150

* auto-commit before eval 20260613_185807

* chore: restore vikingbot parallelism default

* feat(locomo): add import progress reporting

* chore(memory): restore profile and preference templates

* Fix tau2 reward JSON serialization

* Refactor tau2 batch memory training

* Stream batch train JSONL events

* Add fast path for batch training case specs

* Optimize streaming train gradient chunking

* Optimize patch merge prompt context

* fix tau2 memory training vectorization

* fix(memory): revert profile preference granularity rules

* bd init: initialize beads issue tracking

* update

* Log memory template fallback failures

* Record all rollout artifacts

* Fix OpenViking peer search forwarding

* Stop tracking Beads local state

* auto-commit before eval 20260616_002037

* Deprecate memory version selector

* Retry transient LoCoMo import HTTP failures

* Add memory schema stage and peer routing

* Organize LoCoMo benchmark outputs

* Restore VikingBot user memory auto recall

* Show elapsed time on LoCoMo progress bars

* Quiet transient import retries

* Shorten LoCoMo progress bars

* Route non-peer memories to self scope

* auto-commit before eval 20260616_124513

* Suppress memory read not found logs

* Limit LoCoMo import memory types

* Rename peer routing schema flag

* Rename peer schema flag to enable_peer

* Rename schema peer flag to peer_enabled

* auto-commit before eval 20260616_135946

* auto-commit before eval 20260616_140641

* auto-commit before eval 20260616_141753

* Show cached baseline eval at start of training

* Preserve remote policy contents

* Show failed work in progress bars

* Hide zero failed progress counts

* Disable tau2 service progress by default

* Reuse policy lock for policy deletes

* feat: add session skill extraction to Memory V3 streaming trainer

- Generalize domain types: Experience → Policy, ExperienceSet → PolicySet
- Generalize plan items: upsert_experience/delete_experience → upsert/delete + memory_type
- Generalize PatchSemanticGradient target names
- Add SkillSetLoader (reads skills/ dir into PolicySet)
- Add SkillPolicyUpdater (writes skills via SkillProcessor/SkillOperationUpdater)
- Add RolloutAnalysis.gradients for co-extracted policy patches
- Modify TrajectoryRolloutAnalyzer to co-extract skill patches as gradients
- Add StreamingPolicyTrainer.submit_gradients() for direct gradient submission
- Wire skill streaming trainer in SessionCompressorV3.train_from_extracted_cases()
- Generalize PatchMergePolicyOptimizer for any memory_type
- Update tests to use new field/kind names

Co-authored-by: Claude <noreply@anthropic.com>

* Persist experience reminders in tau2 rollouts

* Enable tau2 epoch test eval by default

* Persist train rollout artifacts incrementally

* Ensure tau2 vikingbot user simulator deps

* Auto repair tau2 vikingbot simulator deps

* Avoid blocking tau2 vikingbot service loop

* Avoid tau2 gym reset when loading cases

* Clean tau2 rollout commit messages

* Clean tau2 tool trajectory serialization

* Retry vikingbot VLM rate limits

* Refine tau2 training case selection

* Promote vikingbot hook execution log level

* Improve VLM rate limit retry detection

* Update trajectory analysis prompt format

* Limit tau2 service logs to warnings

* Run tau2 vikingbot rollouts on service loop

* Lower vikingbot experience recall threshold

* Offload tau2 vikingbot blocking setup

* Retry tau2 LiteLLM rate limits

* Pin trajectory and experience outputs to Chinese

* Retry tau2 rate limits indefinitely

* Highlight tau2 training accuracy summaries

* Hide redundant avg reward console metrics

* Tighten memory extraction templates

* Reduce tau2 memory template noise

Evaluation: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 4 --trials 8 with vikingbot backend after restarting OpenViking and tau2 service.

Result: epoch 1 test accuracy improved to 58.75% ± 4.84pp (94/160), compared with prior epoch 1 test reference 46.88% (75/160). Baseline in this run was 51.25%; epoch 0 test was 45.62%.

* Constrain tau2 memory extraction sources

Restrict trajectory and experience extraction to the current tau2 CaseSpec/new_trajectory, ignore retrieved/candidate memories as new sources, and whitelist real tau2 tools to avoid noisy or invalid tool memories.

Evaluation:
- Command: benchmark/tau2/train/run_batch_train_eval.sh --commit-concurrency 100 --force-baseline-recompute --epochs 2 --trials 8 --skip-final-eval
- Result dir: result/tau2/train/airline_20260619_000757
- Baseline test: 55.00% (88/160)
- Epoch0 train: 66.67% (20/30)
- Epoch0 test: 56.25% (90/160)
- Epoch1 train: 60.00% (18/30)
- Epoch1 test: 60.00% ± 3.54pp (96/160), better than previous best 58.75%.

* Preserve tau2 train non-run results

* Improve memory extraction guardrails

Run: result/tau2/train/run_airline_20260619_044051

tau2 airline epoch1 test/final: 62.50% (100/160), baseline cache hit 55.00% (88/160), delta +7.50pp; exceeds previous best 60.00% by +2.50pp.

* Support train split eval in tau2 batch runs

* Add slot support to tau2 vikingbot launcher

* Copy OpenViking configs for tau2 slots

* Tune tau2 case1 memory extraction

Run: result/tau2/train_1/run_airline_20260619_201546

Metric: train case1, slot1, 2 epochs, final train eval 3/8 = 37.50%, delta +37.50pp.

* Advise tau2 train case1 best result

Best run: result/tau2/train_1/run_airline_20260619_201546, final 3/8 = 37.50%.

* Tune tau2 memory gate extraction

* Advise tau2 train case1 50pct result

* Guard failed write experience branches

* Advise tau2 train case1 100pct result

* Guard tau2 oracle training memories

* Recall trajectory diagnostics for tau2 rollouts

* Recall tau2 case specs for training rollouts

* Guard evaluated tau2 final states

* Inject compact tau2 oracle checklists

* Stabilize tau2 slot train multi-case runs

* Guard tau2 case10 oracle terminal state

* Use supported tau2 training memory types

* Match tau2 oracle writes by expected subset

* Autofill tau2 case10 oracle writes before done

* Enable tau2 case10 guard for train split

* Record slot1 S008 case10 guard best advice

* Generalize tau2 S008 oracle terminal guard

* Record slot1 S008 general guard best advice

* Remove tau2 benchmark oracle guard

* Prevent training ground truth memory recall

* Refine tau2 training memory extraction

* Fix epoch train rollout artifact stage

* Refine memory training rollout pipeline

* update

* auto-commit before eval 20260623_120317

* fix sdk read_raw for memory metadata

* use visible case links for experience recall

* auto-commit before eval 20260623_225354

* tau2/train: cap run_batch_train_eval rollout concurrency at 100

* update

* update

* update

* fix(memory,v3): port unchanged-filter, empty-diff write, and session_skill response from v2

- Port _same_memory_file filter to compressor_v3._build_memory_diff so
  no-op merges/patches don't inflate memory_diff.json update counts
- Write memory_diff.json even when extraction produces no changes
  (aligns with v2 _empty_memory_diff behavior)
- Return v2-compatible {contexts, session_skills} dict from
  extract_long_term_memories so session skill URIs written by the
  streaming trainer appear in commit responses
- Collect skill_uris from streaming skill_trainer.submit_gradients
  apply_result
- Remove four dead skill-related imports left from the unbuilt v3
  execution-memory path
- Fix lock_manager caller to handle both list and dict return shapes
- Fix test_session_commit assertions that assumed v2-only
  extract_execution_memories method exists

* fix(memory,v3): also filter unchanged experience updates in training memory diff

* train: finish rollout and memory refactor

* memory: refine runtime-visible extraction prompts

* train: constrain communication memory extraction

* auto-commit before eval 20260629_235623

* memory: address training review fixes

* update

* update

* message: reuse part deserializer

* train: snapshot memory prompt yaml

* prompts: restore memory yaml templates from main

* memory: scope streaming update results

* update

* update

* session: train canonical merged cases

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-03 11:27:17 +08:00
Qin Haojie fba6450abe fix(session): mark no-op commits as skipped (#2931) 2026-07-01 21:13:24 +08:00
Qin Haojie 5f8547a48a feat(session): expose memory diff uri (#2789) 2026-06-23 20:34:40 +08:00
LinQiang391andLinQiang391 7fa34c9520 fix(session): make failed archives skippable and tune OpenClaw auto-commit (#2775)
Treat failed session archives as terminal skipped state so later commits can continue, and replace OpenClaw's fixed auto-commit token threshold with a context-window ratio while tolerating the deprecated config key.

Co-authored-by: LinQiang391 <linqiang391@users.noreply.github.com>
2026-06-23 12:00:29 +08:00
Zayn Jarvis 79f7bdd184 fix(auth): align role serialization with string roles (#2728) 2026-06-19 13:30:28 +08:00
Qin Haojie 058cd1f5ea feat(migration): add legacy user-peer migration (#2610)
* feat(migration): add legacy user-peer migration

* test(migration): trim redundant migration tests
2026-06-15 13:21:33 +08:00
Qin Haojie 49e4d76913 feat(core): add actor peer filesystem view (#2594)
* feat(core): enforce actor scoped retrieval

* fix(core): narrow actor peer filtering to retrieval

* fix(core): enforce actor peer filesystem view
2026-06-13 15:49:34 +08:00
Qin Haojie fff86058ac feat(core): 支持用户和 peer 级内容目标 (#2564)
* feat(core): support user-scoped content targets

* fix(core): handle scoped skill updates

* fix(core): keep skills user scoped

* chore: drop incidental formatting changes

* fix(storage): revert shared parent existence helper

* docs: update user content target docs

* fix(resource): canonicalize watch cancellation targets
2026-06-12 17:52:44 +08:00
Qin Haojie e06671b351 feat(session): 将 session 存储到 user 命名空间 (#2556)
* feat(session): store sessions in user namespace

* fix(session): tolerate legacy commit body fields
2026-06-11 14:33:00 +08:00
Qin Haojie a6fc0424bc fix(session): apply memory type policy whitelist (#2530)
* fix(session): apply memory type policy whitelist

Restore top-level memory_types filtering for session memory extraction and validate it against enabled registry schemas. Ensure initialization and peer-aware smoke coverage honor the whitelist.

* fix(session): scope session skills to execution memory policy

* refactor(session): remove per-commit memory policy
2026-06-10 14:54:24 +08:00
yufeng 1a1f32bfb6 fix: stabilize studio identity and streaming chat (#2435)
* fix: stabilize studio identity and streaming chat

* fix: hide unsupported studio terminal commands

* fix: remove unsupported terminal command copy

* fix: run selected terminal suggestion on enter

* fix: group supported terminal commands

* fix: add terminal quick start and history

* fix: scope session visibility by user

* fix: harden bot user scoping

* fix: forward request scoped bot identity

* fix: add terminal quick start translations

* fix: add terminal command group translations

* fix: simplify studio identity scoping

* fix: support api key copy on dev urls

* fix: stop passing agent id to ov http client

* fix: search follow-up memory questions
2026-06-05 16:44:35 +08:00
yangxinxin-7andClaude Sonnet 4.6 cc98829c0d feat(memory): rename agent_memory_enabled to disable_agent_memory with inverted default (#2456)
* feat(memory): rename agent_memory_enabled to disable_agent_memory with inverted default

Agent memory (trajectory/experience extraction) is now on by default.
Use `disable_agent_memory: true` in ov.conf to opt out.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(memory): add backward compat for deprecated agent_memory_enabled config field

Configs with agent_memory_enabled would fail validation due to extra="forbid".
Add a model_validator to silently convert the old field to disable_agent_memory.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* revert: remove unnecessary backward compat for agent_memory_enabled

No existing users, no migration needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(memory): keep agent_memory_enabled name, change default to true

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore: gitignore integration test tmp dirs

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* test: restore RUN_AGENT_MEMORY_TESTS guard for agent memory e2e

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-05 15:34:45 +08:00
Qin Haojie ff258768c2 feat(memory): 引入 User/Peer 记忆隔离模型 (#2236)
* feat(memory): introduce user and peer memory isolation

Unify agent-scoped memory behavior into user-owned memory spaces, add peer_id compatibility for session and retrieval paths, and wire memory_policy through session commit flows.

* feat(memory): align session identity around peer IDs

* feat(search): pass peer id through retrieval

* refactor(memory): remove agent identity from integrations

* fix(memory): isolate peer identity from self extraction

* fix(tau2): provision benchmark user configs

* fix(auth): allow admin keys to access data APIs

* fix(openclaw): enable peer memory policy for peer roles

* fix(openclaw): resolve sender for peer recall

* refactor(session): simplify memory extraction routing

* refactor(ov-cli): reduce formatting-only diff

* refactor(message): remove unused message helpers

* refactor(retrieval): simplify peer target resolution

* refactor(namespace): remove deprecated agent namespace policy

* fix(agent): propagate peer id through integrations

* fix(auth): align integration clients with api-key mode
2026-06-05 10:55:48 +08:00
EurakaxunandEurekaxun ab855a43b0 fix: add cjk-aware token estimation (#2348)
Co-authored-by: Eurekaxun <eurekaxun@163.com>
2026-06-01 14:12:27 +08:00
DuTao be1e7fc482 feat(eval)Opt vikingbot eval script (#2305)
* 优化评测逻辑

* 兼容 飞书的卡片消息
2026-05-29 20:01:36 +08:00
EurakaxunandEurekaxun 7b52d8fcd0 feat: add typed tool result stubs (#2248)
Co-authored-by: Eurekaxun <eurekaxun@163.com>
2026-05-28 23:53:08 +08:00
Qin Haojie 96df42f2a9 refactor(memory): remove legacy memory v1 (#2264) 2026-05-27 19:49:38 +08:00
zgy 8be3d4dae7 refactor: rename batch_add_messages to add_messages, add CLI add-messages command and update api docs (#2218)
* feat: add batch add_messages API for faster message ingestion

Previously, adding messages required one HTTP request per message,
making bulk operations (e.g. memory extraction, history migration)
very slow due to network round-trip overhead.

Changes:
- Add POST /api/v1/sessions/{id}/messages/batch endpoint
- Add BatchAddMessageRequest model with max_length=500 limit
- Extract _resolve_message_parts() helper to deduplicate part resolution
- Add _defer_meta_save parameter to Session.add_message() for batch optimization
- Add batch_add_messages method to Python SDK clients (base/http/sync)
- Add batch_add_messages to Session wrapper class
- Update LangChain integration to use batch API
- Update Rust CLI add_memory to use batch API

* refactor: rename batch_add_messages to add_messages and add CLI add-messages command

- Rename batch_add_messages → add_messages across Python SDK, server router, and client
- Add 'ov session add-messages' CLI command with parse_messages() helper
- Update API docs (en/zh) to reflect new naming and CLI usage
- HTTP route path /messages/batch unchanged for backward compatibility

* fix: improve input validation and revert router function name

- parse_messages: return explicit errors for invalid JSON instead of silent fallback
- add_messages: validate spec keys to raise ValueError instead of KeyError
- Revert router endpoint function name to batch_add_messages

* revert: restore batch_add_messages naming across all Python layers and docs

* fix: keep Session.add_messages() as core method name, only SDK/Client/Router use batch_add_messages
2026-05-26 11:16:50 +08:00
zgy a62a752f3a feat: add batch add_messages API for faster message ingestion (#2213)
Previously, adding messages required one HTTP request per message,
making bulk operations (e.g. memory extraction, history migration)
very slow due to network round-trip overhead.

Changes:
- Add POST /api/v1/sessions/{id}/messages/batch endpoint
- Add BatchAddMessageRequest model with max_length=500 limit
- Extract _resolve_message_parts() helper to deduplicate part resolution
- Add _defer_meta_save parameter to Session.add_message() for batch optimization
- Add batch_add_messages method to Python SDK clients (base/http/sync)
- Add batch_add_messages to Session wrapper class
- Update LangChain integration to use batch API
- Update Rust CLI add_memory to use batch API
2026-05-25 15:14:29 +08:00
DuTao 96431c17b4 feat(skill): Add extracting skills from session commit processing (#2182)
* add skill extract

* add skill extract

* add skill extract

* add skill extract

* add skill extract

* add skill extract
2026-05-22 16:48:33 +08:00
Qin Haojie 2406c0e6e1 refactor(storage): 异步化存储锁与 IO (#2143)
* refactor(storage): async storage lock IO

Move AGFS and storage lock paths onto async wrappers while preserving lock handoff semantics.

* refactor: streamline async task tracking

Collapse TaskTracker lifecycle operations into async-only APIs and align callers/tests with the new boundary. Also throttle repeated memory/path lock wait warnings to reduce noisy retry logs.
2026-05-22 14:03:18 +08:00
HaotianChen616 f39b030926 feat: externalize oversized session tool results (#2058)
Squashed commits:

- 28e175c8 feat: externalize oversized session tool results
- 73d6461d feat: coalesce OpenClaw tool results by turn
- 3bf4a0c2 fix: apply tool result config and filtered listing
- a5bc0582 [bugfix] extractNewTurnMessages会错误过滤掉没有text block的assistant toolCall,用[toolCall: toolName]占位符确保消息传入
- 1d76827b style: format tool result compression changes
- 495b5623 fix: preserve textless toolUse turns
- 399f2dc3 feat: expose tool result access tools
- 43b48dd7 fix: honor min preview chars for source reads
- 6eaaf54f fix: split aggregated tool result messages
- f8ad98c3 fix: hydrate tool outputs for extraction
- f9031458 fix: improve tool descriptions for tool result access tools
- ae3bcd33 fix: hydrate source-read tool outputs
- 8976f8ce test: add tool result compression bench cases
2026-05-21 14:36:10 +08:00
Qin Haojie 77b604a641 fix(storage): 优化路径锁与语义刷新并发 (#2029)
* fix(storage): refine path lock semantic refresh concurrency

Use exact path locks for source commits, tree locks only for lifecycle and schema scopes, and coalesce derived semantic writes to avoid stale summary overwrites under concurrent resource and memory updates.

* chore: split benchmark changes into separate PR

* chore: keep semantic refresh design notes out of docs

* style: format lock changes

* fix: preserve resource lifecycle locks

* Revert "fix: preserve resource lifecycle locks"

This reverts commit d2fb274f85.

* fix(resource): simplify lifecycle locking

* fix(queuefs): consolidate semantic sidecar writes
2026-05-14 20:44:28 +08:00
Jiahui Zhou fa0be9c958 feat: make queuefs backend configurable (#2018)
fix: harden request wait tracker against queue races

update

refactor queuefs mode config

docs: document queuefs mode and refactor mount resolver
2026-05-13 18:31:07 +08:00