Commit Graph
7 Commits
Author SHA1 Message Date
9eac8a6d3d feat(sdk): sync go/ts/python SDKs with server find/search, recall, an… (#3737)
* feat(sdk): sync go/ts/python SDKs with server find/search, recall, and admin changes

Server-side changes recently landed that the language SDKs had drifted from:

- find/search results now return `tags` and no longer return
  `category`/`match_reason`/`relations`/`overview` (#3730). Go's strict
  struct was the only one broken; update MatchedContext accordingly.
- new admin endpoints for agent-evolution and per-account settings (#3695).
- public `search/recall` endpoint was missing from all SDKs.

Changes:
- python: add `level`/`since`/`until`/`time_field` to find/search; add an
  `extra` escape hatch to find/search/add_resource/write/batch_write so new
  server fields can be passed without an SDK bump (only forwarded when set,
  preserving `level=0`); add `recall` and the four admin methods.
- go: fix MatchedContext (add Tags, drop removed fields), add Recall and the
  four admin methods.
- typescript: type MatchedContext/FindResult, add RecallOptions, add `recall`
  and the four admin methods.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): unify options APIs and sync latest server interfaces

- migrate complex Python SDK calls to typed options dictionaries
- add dedicated context search and consistent extra-field handling
- align Go and TypeScript options with omission-aware serialization
- support session config, event tags, Agent Evolution date filters,
  OpenViking Assets, batch write, downloads, and create_parent
- refresh SDK tests and examples across all three languages

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): address options API review findings

- fix Go session extra merging and Python message precedence
- adapt LangChain calls to the Python options API
- migrate repository examples, tests, and documentation

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): complete options migration and message parity

- migrate remaining Python SDK benchmarks to options dictionaries
- normalize empty parts consistently for single and batch messages
- add regression guards for repository SDK call sites

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): align reindex options after main rebase

- preserve reindex tags in Python typed options
- add reindex extra support for Go and TypeScript
- reject official fields passed through extra across SDKs

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): support legacy keyword options

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

docs(sdk): use explicit Python SDK arguments

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): support set tags extra options

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): expose Go add resource options

Expose AddType and ProcessingMode through Go AddResourceOptions and serialize them to the resources API. Add a regression test covering the resulting request payload.

Co-authored-by: TRAE CLI <noreply@bytedance.com>
Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(sdk): flatten core Python client options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

docs(sdk): align Python examples with flattened options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): preserve core API compatibility

Co-authored-by: TRAE CLI <traecli@bytedance.com>

refactor(python-sdk): move resource hints to options

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): align resource option callers

Co-authored-by: TRAE CLI <traecli@bytedance.com>

test(sdk): cover recursive reindex forwarding

Co-authored-by: TRAE CLI <traecli@bytedance.com>

fix(sdk): preserve Go options compatibility

Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(python-sdk): expose message peer id

Co-authored-by: TRAE CLI <traecli@bytedance.com>

test(python-sdk): consolidate options coverage

Co-authored-by: TRAE CLI <traecli@bytedance.com>

feat(python-sdk): add parts and flatten image search

Co-authored-by: TRAE CLI <traecli@bytedance.com>

* docs(sdk): align Python call examples

Co-authored-by: TRAE CLI <traecli@bytedance.com>

---------

Co-authored-by: TRAE CLI <traecli@bytedance.com>
Co-authored-by: Qin Haojie <qinhaojie.exe@bytedance.com>
2026-08-24 14:09:11 +08:00
Qin Haojie 7abd6ab249 refactor(client): remove Python embedded mode (#3712)
* refactor(client): remove Python embedded mode

Consolidate Python consumers on the HTTP SDK while keeping shared server and storage capabilities unchanged.

* refactor(client): remove obsolete embedded leftovers
2026-08-10 18:00:00 +08:00
Zayn Jarvis 6ae4a9398b feat: replace controversial examples with neutral alternatives (#1844) 2026-05-04 12:06:58 +08:00
Qin Haojie 38c324bc97 fix(security): clean up code scanning and runtime findings (#1596)
* fix(security): clean up code scanning and runtime findings

Harden path and logging boundaries, remove noisy cleanup issues,
and keep observability failures from breaking runtime flows.

* fix(security): close werewolf and feishu validation gaps

Block the remaining path traversal bypass in the werewolf demo,
and validate Feishu hosts on the main parse() entry point.
2026-04-21 10:06:46 +08:00
Qin Haojie 3c2f095c4f fix(volcengine): update default doubao embedding model (#1438)
Align docs, examples, the setup wizard, and telemetry tests with the
current Doubao multimodal embedding model so generated configs keep
referencing the supported default consistently.
2026-04-14 15:51:06 +08:00
MaojiaShengandopenviking a7e5417ef2 reorg: remove golang depends (#1339)
* docs: fix docker deployment

* reorg: remove third_party/agfs

* feat(s3fs): add disable_batch_delete option for OSS compatibility

Port of PR #1333 from Go version to Rust:

- Add disable_batch_delete config option to S3Client
- When enabled, use sequential single-object deletes instead of DeleteObjects
- This is for S3-compatible services like Alibaba Cloud OSS that require
  Content-MD5 for DeleteObjects but AWS SDK v2 does not send it by default
- Add documentation and config example for OSS

* fix(s3fs): pass disable_batch_delete config from Python to Rust

Add disable_batch_delete to the s3_plugin_config dict in _generate_plugin_config
so that the Python config can properly control the Rust S3FS plugin's behavior.

* reorg: remove third_party/agfs

* reorg: remove third_party/agfs

* change some docs

* change some docs

---------

Co-authored-by: openviking <openviking@example.com>
2026-04-10 15:16:29 +08:00
zgy 1b3a8f20a0 Feat(benchmark): Add benchmark/RAG : RAG system evaluation framework (#825)
* Add RAGbenchmark: RAG system evaluation framework

* Update README.md

* Update README.md

* Update README.md

* Code structure refactoring

* feat: improve RAG benchmark with dataset sampling and configuration updates

- Add complete dataset sampling scripts with document-level sampling
- Implement filtering logic consistent with adapters (exclude category 5 for Locomo, no answer for SyllabusQA, unanswerable for Qasper)
- Update configuration from raw_data/dataset_dir to dataset_path for clarity
- Enhance adapters with improved path handling and data loading
- Add gitignore for data and output directories
- Add dependencies (datasets, pandas, tavily-python)
- Add test files and documentation

* feat: add stratified sampling support to all datasets

- Implement stratified sampling for Locomo (by category 1-4)
- Implement stratified sampling for SyllabusQA (by question_type)
- Implement stratified sampling for Qasper (by answer type: extractive/free_form/yes_no)
- Implement stratified sampling for FinanceBench (by question_type)
- Add proper handling when sample size cannot be evenly split:
  - Display warning message
  - Distribute remaining QAs to first N categories
  - Fall back to random sampling if sample size too small
- Update prepare_dataset.py to support both 'random' and 'stratified' modes
- Set default sampling mode to 'random'

* Update locomo adapter to support image attachments and other improvements

* Update dataset documentation with actual document counts

* Add benchmark results reference and reproduction steps

* Improve sampling scripts for benchmark reproducibility

* Refactor sample_dataset.py: extract common sampling logic

- Fix two bugs:
  1. num_docs + sample_size + random path: use int indices instead of dict tuples
  2. pure stratified path: use len() for list length calculation

- Extract common sampling utilities:
  - calculate_category_targets()
  - stratified_sample_with_reallocation()
  - random_sample_qas()
  - sample_docs_stratified()
  - sample_docs_random()

- Reduce code duplication by ~60-70%
- Improve maintainability and readability
- Keep full backward compatibility

* Update config.yaml: improve configuration structure

- Add FinanceBench to supported datasets list
- Change to template configuration format
- Add execution: section for better organization

* Fix bug: duplicate worker_end() call in generation failure path

- Remove duplicate monitor.worker_end(success=False) call in run_generation()
- The _process_generation_task() already calls worker_end() in its exception handler
- This prevents double-counting of failed tasks and distorted statistics

* Fix bug: _get_required_syllabi() doesn't support JSON input

- Add JSON file support to _get_required_syllabi()
- Extract syllabus names from JSON keys (same format as _load_from_json())
- This ensures data_prepare() processes correct docx files when using JSON input

* Improve exception re-raising: use bare raise to preserve traceback

- Replace 'raise e' with bare 'raise' to preserve original traceback
- Also remove unused 'e' variable since we don't need it
- This makes debugging easier by showing where the exception actually occurred

* Fix bug: Locomo prompt uses raw gold_answer instead of gold_answer_str

- In Locomo prompt, use gold_answer_str instead of gold_answer
- This ensures consistent formatting when gold_answer is a list
- Both Locomo and Generic prompts now use the same ' | ' separated format

* Improve directory ingest: use os.path.commonpath() for robustness

- Replace manual common ancestor calculation with os.path.commonpath()
- os.path.commonpath() handles all OS path separators correctly
- Add try-except to handle ValueError when no common path exists
- More robust than manual split(os.sep) approach

* benchmark: honor skip_ingestion and fail on LLM retry exhaustion
2026-04-01 14:39:53 +08:00