Commit Graph
2 Commits
Author SHA1 Message Date
Aashish 21cb46074d feat(onnx): opt-in INT8 quantized copy in the ONNX export
Add --quantize to scripts/export_onnx.py: dynamic per-channel INT8
weight-only quantization of the MatMul layers, written as a sidecar
(<output>.int8.onnx) next to the fp32 export so A/B stays possible.
Measured on the English checkpoint (CPU, 20 ticket states x
choice/noul/score): 1.6 GB -> 581 MB, p50 ~340 ms -> ~250 ms, zero
decision changes (max probability drift 0.09); per-tensor scales
flipped 3 of 20 states, so per-channel is what ships. Weight-free
suite on a synthetic MatMul graph, run in a new CI lane that installs
onnx.
2026-09-25 20:34:56 +05:45
Swayam PatilandNandakishor 60965402db feat: add torch.compile support and ONNX Runtime inference engine (#215)
Co-authored-by: Nandakishor <48623612+NandhaKishorM@users.noreply.github.com>
2026-09-23 19:30:16 +05:30