Align docs, examples, the setup wizard, and telemetry tests with the
current Doubao multimodal embedding model so generated configs keep
referencing the supported default consistently.
* docs: fix docker deployment
* reorg: remove third_party/agfs
* feat(s3fs): add disable_batch_delete option for OSS compatibility
Port of PR #1333 from Go version to Rust:
- Add disable_batch_delete config option to S3Client
- When enabled, use sequential single-object deletes instead of DeleteObjects
- This is for S3-compatible services like Alibaba Cloud OSS that require
Content-MD5 for DeleteObjects but AWS SDK v2 does not send it by default
- Add documentation and config example for OSS
* fix(s3fs): pass disable_batch_delete config from Python to Rust
Add disable_batch_delete to the s3_plugin_config dict in _generate_plugin_config
so that the Python config can properly control the Rust S3FS plugin's behavior.
* reorg: remove third_party/agfs
* reorg: remove third_party/agfs
* change some docs
* change some docs
---------
Co-authored-by: openviking <openviking@example.com>
* Add RAGbenchmark: RAG system evaluation framework
* Update README.md
* Update README.md
* Update README.md
* Code structure refactoring
* feat: improve RAG benchmark with dataset sampling and configuration updates
- Add complete dataset sampling scripts with document-level sampling
- Implement filtering logic consistent with adapters (exclude category 5 for Locomo, no answer for SyllabusQA, unanswerable for Qasper)
- Update configuration from raw_data/dataset_dir to dataset_path for clarity
- Enhance adapters with improved path handling and data loading
- Add gitignore for data and output directories
- Add dependencies (datasets, pandas, tavily-python)
- Add test files and documentation
* feat: add stratified sampling support to all datasets
- Implement stratified sampling for Locomo (by category 1-4)
- Implement stratified sampling for SyllabusQA (by question_type)
- Implement stratified sampling for Qasper (by answer type: extractive/free_form/yes_no)
- Implement stratified sampling for FinanceBench (by question_type)
- Add proper handling when sample size cannot be evenly split:
- Display warning message
- Distribute remaining QAs to first N categories
- Fall back to random sampling if sample size too small
- Update prepare_dataset.py to support both 'random' and 'stratified' modes
- Set default sampling mode to 'random'
* Update locomo adapter to support image attachments and other improvements
* Update dataset documentation with actual document counts
* Add benchmark results reference and reproduction steps
* Improve sampling scripts for benchmark reproducibility
* Refactor sample_dataset.py: extract common sampling logic
- Fix two bugs:
1. num_docs + sample_size + random path: use int indices instead of dict tuples
2. pure stratified path: use len() for list length calculation
- Extract common sampling utilities:
- calculate_category_targets()
- stratified_sample_with_reallocation()
- random_sample_qas()
- sample_docs_stratified()
- sample_docs_random()
- Reduce code duplication by ~60-70%
- Improve maintainability and readability
- Keep full backward compatibility
* Update config.yaml: improve configuration structure
- Add FinanceBench to supported datasets list
- Change to template configuration format
- Add execution: section for better organization
* Fix bug: duplicate worker_end() call in generation failure path
- Remove duplicate monitor.worker_end(success=False) call in run_generation()
- The _process_generation_task() already calls worker_end() in its exception handler
- This prevents double-counting of failed tasks and distorted statistics
* Fix bug: _get_required_syllabi() doesn't support JSON input
- Add JSON file support to _get_required_syllabi()
- Extract syllabus names from JSON keys (same format as _load_from_json())
- This ensures data_prepare() processes correct docx files when using JSON input
* Improve exception re-raising: use bare raise to preserve traceback
- Replace 'raise e' with bare 'raise' to preserve original traceback
- Also remove unused 'e' variable since we don't need it
- This makes debugging easier by showing where the exception actually occurred
* Fix bug: Locomo prompt uses raw gold_answer instead of gold_answer_str
- In Locomo prompt, use gold_answer_str instead of gold_answer
- This ensures consistent formatting when gold_answer is a list
- Both Locomo and Generic prompts now use the same ' | ' separated format
* Improve directory ingest: use os.path.commonpath() for robustness
- Replace manual common ancestor calculation with os.path.commonpath()
- os.path.commonpath() handles all OS path separators correctly
- Add try-except to handle ValueError when no common path exists
- More robust than manual split(os.sep) approach
* benchmark: honor skip_ingestion and fail on LLM retry exhaustion