mirror of
https://github.com/NandhaKishorM/laya.git
synced 2026-09-28 16:02:56 +08:00
Add --quantize to scripts/export_onnx.py: dynamic per-channel INT8 weight-only quantization of the MatMul layers, written as a sidecar (<output>.int8.onnx) next to the fp32 export so A/B stays possible. Measured on the English checkpoint (CPU, 20 ticket states x choice/noul/score): 1.6 GB -> 581 MB, p50 ~340 ms -> ~250 ms, zero decision changes (max probability drift 0.09); per-tensor scales flipped 3 of 20 states, so per-channel is what ships. Weight-free suite on a synthetic MatMul graph, run in a new CI lane that installs onnx.