Commit Graph
6 Commits
Author SHA1 Message Date
Yuanqing ZHAOandYuanqing Zhao fa19ac0a75 perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest (#3277)
* perf(vectordb): coalesce auto cuVS rebuilds during bulk ingest

Add an opt-in bulk-ingest maintenance scope that coalesces Auto cuVS background rebuilds across multiple write batches.

- defer derived GPU maintenance until the outermost bulk scope exits while keeping native writes and persistence visible per call
- harden the background worker against debounce, generation, shutdown, and stale-candidate races
- preserve suspension across index replacement and retire replaced workers
- wait for the final Auto GPU snapshot before vectordb_perf records search QPS
- document that the scope is non-transactional and only schedules readiness on exit

Auto cuVS and background rebuild remain disabled by default. Native CPU and remote backends use no-op hooks, so their existing behavior and dtype are unchanged.

* fix(vectordb): reject stale index replacements

* fix(vectordb): harden bulk rebuild lifecycle

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-16 18:59:16 +08:00
Yuanqing ZHAOandYuanqing Zhao 546da35cc8 fix(benchmark): load WIKI-Dir path mapping (#3279)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-16 11:11:15 +08:00
Yuanqing ZHAOandYuanqing Zhao 0080f94bdc perf(vectordb): batch benchmark upserts (#3264)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-15 19:49:06 +08:00
Yuanqing ZHAOandYuanqing Zhao 7e6a0515f9 perf(cuvs): optimize filters, rebuilds, concurrency, and memory (#3092)
* perf(cuvs): fast-path cached native filter routes

* perf(cuvs): parallelize auto filter preflight

* perf(cuvs): add search route telemetry

* test(cuvs): use a valid telemetry vector dimension

* perf(cuvs): reuse native filter preflight results

* perf(cuvs): allow concurrent snapshot searches

* perf(cuvs): coalesce optional background rebuilds

* perf(cuvs): coordinate per-GPU build admission

* perf(cuvs): add opt-in float16 search

* build(cuvs): support vector benchmark harnesses

* perf(cuvs): bound concurrent GPU searches

* perf(cuvs): avoid partial background rebuilds

* fix(cuvs): address rebuild and telemetry review feedback

* fix(cuvs): defer rebuild until index initialization

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-10 17:22:34 +08:00
Yuanqing ZHAOandYuanqing Zhao d61d802fd0 fix(benchmark): enforce configured search concurrency (#3091)
Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-09 11:06:58 +08:00
Yuanqing ZHAOandYuanqing Zhao 39c778c953 feat: add cuVS vector search backend (#2974)
* feat: add cuVS vector search backend

* docs: add agent memory benchmark strategy

* bench: add cuVS index performance harness

* bench: add public ANN dataset tuning

* docs: record preliminary cuVS index results

* docs: clarify warm index latency

* docs: order cuVS before qdrant

* bench: aggregate independent index runs

* bench: order aggregate variants consistently

* docs: add repeatable index scaling results

* bench: add collection lifecycle benchmark

* docs: add collection lifecycle results

* perf: cache prepared cuvs filters

* docs: report prepared filter cache results

* bench: add async vector concurrency benchmark

* bench: aggregate service concurrency runs

* docs: add async concurrency results

* docs: clarify cuVS dtype behavior

* feat: add memory-aware cuVS auto mode

* feat: reuse native filters for cuVS search

* docs: publish cuVS integration plan as Markdown

* fix: route selective filters before cuVS rebuild

* docs: record selective-first routing results

---------

Co-authored-by: Yuanqing Zhao <2604121+yuanqingz@users.noreply.github.com>
2026-07-07 12:21:10 +08:00