fix(docker): keep torch's native Triton JIT off so a GPU image can infer without a C compiler

torch 2.14 imports torch._native, which swaps eager bmm/topk/sum/norm for Triton kernels
compiled on the first CUDA call. The slim runtime image carries no compiler, so /health is
ok and every request fails (#365). TORCH_DISABLE_NATIVE_JIT=1 is torch's own switch for
those overrides: stock kernels, same answers, same latency, zero bytes added to the image.

Before/after on an RTX 2000 Ada (driver 590.48.01), images built with TORCH_INDEX=cu130,
no compiler, cold Triton cache: main -> "Failed to find C compiler"; this branch -> answers.
p50 19.7 vs 20.0 ms for three questions with the overrides on and off.
This commit is contained in:
Bhushan Kinge
2026-09-24 12:49:34 -07:00
parent 970dc8c5f6
commit 846a75d3fe
4 changed files with 20 additions and 0 deletions
+3
View File
@@ -34,6 +34,9 @@ jobs:
print(torch.__version__, cusparselt.version())
'
docker run --rm --network none laya:spark pip check
# torch 2.14 compiles Triton kernels on the first CUDA inference unless this is set (#365).
- name: A CUDA image must keep torch's native Triton JIT off
run: docker run --rm --network none laya:spark sh -c 'test "$TORCH_DISABLE_NATIVE_JIT" = 1'
smoke:
# Native runners for both architectures; torch wheels install for the target, so
+5
View File
@@ -33,7 +33,12 @@ LABEL org.opencontainers.image.title="Laya Docker quickstart" \
org.opencontainers.image.source="https://github.com/NandhaKishorM/laya" \
org.opencontainers.image.licenses="Apache-2.0"
# torch 2.14 swaps some eager CUDA ops (bmm, topk, sum, norms) for Triton kernels that it
# compiles on the first inference, which needs a C compiler this image does not carry: the
# container reports healthy, then every request fails (#365). The stock kernels give the same
# answers at the same latency.
ENV PATH="/opt/venv/bin:$PATH" \
TORCH_DISABLE_NATIVE_JIT=1 \
PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
USE_TF=0 \
+6
View File
@@ -48,6 +48,12 @@ The sample rejects unavailable CUDA before loading a checkpoint. Laya can still
fall back to CPU after a memory or inference error, so inspect its warnings.
Rebuild when switching between CPU and CUDA configurations.
The image sets `TORCH_DISABLE_NATIVE_JIT=1`. PyTorch 2.14 otherwise replaces some eager
CUDA ops with Triton kernels that it compiles on the first inference, which needs a C
compiler the slim image does not carry: the container reports healthy and then fails every
request (#365). The stock kernels give the same answers at the same latency. Set the same
variable on a bare-metal install if `predict` fails with `Failed to find C compiler`.
This uses [Compose GPU reservations](https://docs.docker.com/compose/how-tos/gpu-support/).
Windows requires Docker Desktop's supported WSL2 GPU setup. Apple MPS,
AMD/ROCm and Intel GPU containers are outside this quickstart; use CPU unless
+6
View File
@@ -184,6 +184,12 @@ check_true("compose.http/leaves the quickstart service alone",
check_true("Dockerfile/installs the serve extra", '".[serve]"' in dockerfile, dockerfile[:400])
check_true("Dockerfile/still runs pip check", "pip check" in dockerfile)
# torch 2.14's eager Triton kernels compile on the first CUDA inference and need a C compiler the
# slim runtime image does not have (#365). The kill switch keeps the stock kernels.
check_true("Dockerfile/runtime stage disables torch's native Triton JIT (#365)",
re.search(r"^\s*TORCH_DISABLE_NATIVE_JIT=1", dockerfile.partition("AS runtime")[2], re.M) is not None,
"without TORCH_DISABLE_NATIVE_JIT=1 a GPU image serves 500s while /health stays green")
# Overrides for `laya` never reach `laya-serve`, a separate service. If the CUDA override does
# not repeat the args for laya-serve, that service silently serves on CPU.
check_true("compose.cuda/covers laya-serve too",