* fix(server): stop exporting raw query strings and buffering zip responses in observability
Sweep findings: B-03, B-13. Prevent query secrets from reaching traces and keep ZIP responses streaming.
(cherry picked from commit d8ac3dc33b)
* fix(session): tolerate missing/corrupt archive in Phase-2 replay
(NotFoundError / _ArchiveMessagesCorruptError) on a missing or corrupt
archive messages.jsonl instead of returning []. That PR added skip-on-
missing tolerance to the read path (_get_uncovered_archive_messages) and to
resume_queued_commit, but not to the Phase-2 commit replay path
(_prepare_phase2_archive_messages), which calls _read_archive_messages
unguarded while rolling earlier failed archives into the current commit.
Consequence: a terminally-failed earlier archive whose messages.jsonl is
missing/corrupt (legacy "no messages" terminal data, or produced by #3417's
own archive_read terminal path) makes every subsequent commit's Phase-2
extraction raise -> caught by _run_memory_extraction's except -> the current
archive is terminal-failed too. Because the poisoned archive is only removed
from replay once "covered" (which requires a later archive to complete), and
no later archive can ever complete, the session's memory extraction is
permanently poisoned. Raw messages are safe, but extraction is stuck.
Fix: wrap the replay-loop _read_archive_messages call in the same tolerance
_get_uncovered_archive_messages already uses -- skip + warn on not-found
(_is_storage_not_found) and on _ArchiveMessagesCorruptError, re-raise real
storage failures. The skipped archive stays in covered_failed so the current
archive's .done marks it covered, clearing the poison permanently.
Adds a regression test asserting the replay skips a failed archive with a
missing messages.jsonl (and marks it covered) instead of raising, and that a
real storage failure still propagates.
Follow-up to #3417.
(cherry picked from commit 5b8ec9e68a)
* fix(client): align client surfaces without leaking memory metadata
Reconstructs the client-parity work from upstream PR #3439 on current main and strips reserved memory metadata before line slicing in both embedded and HTTP reads.
Based-on: 48b411d58c
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(index): propagate semantic vectorization failures safely
Reconstructs upstream PR #3437 on current main, carries enqueue failures through SemanticDagExecutor, and drains the attempt's embedding tracker before retry-visible failure propagation.
Based-on: 02387deb09
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* fix(core): close privacy and embedding failure gaps
* fix(memory): strip repeated metadata trailers
* fix(core): close public memory visibility gaps
* ci: skip embedding-dependent resource test without secrets
---------
Co-authored-by: zhiheng.liu <zhiheng.liu@bytedance.com>
* Add search retrieval telemetry breakdown
* Fix pyagfs helper annotation imports
* docs: document search relation controls and telemetry fields
Document the new include_relations request parameter and the search telemetry summary fields so the API docs stay aligned with the latest retrieval changes.
* refactor(search): drop relation enrichment and trim telemetry
Remove relation fetching from the retrieval path and delete low-value search telemetry fields so retrieval stays simpler and the telemetry summary focuses on actionable diagnostics.
Track semantic and embedding re-enqueues as first-class queue metrics so
observer output, wait_processed payloads, and telemetry summaries make
retry loops visible before they escalate into hard errors.