fix(compose): forward the autocast dtype overrides the runtime reads

laya/agent.py selects the forward's dtype from LAYA_CUDA_AMP and LAYA_CPU_AMP, and no
compose file that sets LAYA_DEVICE passed either name, so the published deployment always
served the checkpoint's amp_dtype whatever the operator exported. README measures the
difference as decisions, not rounding: bf16 flips 3 of 864 argmaxes on the parity_fast set
where fp16 flips none, at the same latency.

Both GPU services get the pair, because an override on `laya` never reaches `laya-serve`.
An empty value means "the checkpoint's own amp_dtype", which is what these images shipped
before, so the passthrough default changes nothing until it is set. The MPS row gate is
deliberately left out: no container here can select MPS.
This commit is contained in:
Aashish
2026-09-27 04:15:46 +05:45
parent 4066d5d5fb
commit d724a099bb
5 changed files with 31 additions and 1 deletions
+8
View File
@@ -8,6 +8,10 @@ services:
TORCH_INDEX: "${LAYA_TORCH_INDEX:-cu128}"
environment:
LAYA_DEVICE: "${LAYA_DEVICE:-cuda}"
# The runtime reads these when it picks the autocast dtype; empty keeps the checkpoint's
# own `amp_dtype`. README's threshold section measures what fp16 vs bf16 decides differently.
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
deploy:
resources:
reservations:
@@ -23,6 +27,10 @@ services:
TORCH_INDEX: "${LAYA_TORCH_INDEX:-cu128}"
environment:
LAYA_DEVICE: "${LAYA_DEVICE:-cuda}"
# The runtime reads these when it picks the autocast dtype; empty keeps the checkpoint's
# own `amp_dtype`. README's threshold section measures what fp16 vs bf16 decides differently.
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
deploy:
resources:
reservations:
+4
View File
@@ -37,6 +37,10 @@ services:
LAYA_HOST: "${LAYA_HOST:-0.0.0.0}"
LAYA_PORT: "${LAYA_PORT:-8000}"
LAYA_DEVICE: "${LAYA_DEVICE:-cpu}"
# Autocast dtype the runtime reads when building the model; empty keeps the checkpoint's own
# `amp_dtype`. Set LAYA_CUDA_AMP=fp16 to serve fp16 on a GPU host.
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
LAYA_PRELOAD: "${LAYA_PRELOAD:-0}"
LAYA_MODELS: "${LAYA_MODELS:-}"
LAYA_THREADS: "${LAYA_THREADS:-${OMP_NUM_THREADS:-4}}"
+8
View File
@@ -13,6 +13,10 @@ services:
TORCH_INDEX: "${LAYA_TORCH_INDEX:-cu130}"
environment:
LAYA_DEVICE: "${LAYA_DEVICE:-cuda}"
# Autocast dtype the runtime reads; empty keeps the checkpoint's own `amp_dtype`. Both
# services need it, because an override on `laya` never reaches `laya-serve`.
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
deploy:
resources:
reservations:
@@ -29,6 +33,10 @@ services:
TORCH_INDEX: "${LAYA_TORCH_INDEX:-cu130}"
environment:
LAYA_DEVICE: "${LAYA_DEVICE:-cuda}"
# Autocast dtype the runtime reads; empty keeps the checkpoint's own `amp_dtype`. Both
# services need it, because an override on `laya` never reaches `laya-serve`.
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
deploy:
resources:
reservations:
+6
View File
@@ -7,6 +7,12 @@ services:
TORCH_VERSION: "${LAYA_TORCH_VERSION:-2.14.0}"
environment:
LAYA_DEVICE: "${LAYA_DEVICE:-cpu}"
# The autocast dtype the runtime picks when it builds the model on the selected device
# (`laya/agent.py`). An empty value means "use the checkpoint's own `amp_dtype`", which is
# what these images have always served. README's threshold section measures what fp16 vs
# bf16 decides differently.
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
LAYA_MODEL: "${LAYA_MODEL:-auto}"
LAYA_MODEL_PATH: "${LAYA_MODEL_PATH:-}"
LAYA_REQUEST_FILE: "${LAYA_REQUEST_FILE:-/opt/laya/examples/request.json}"
+5 -1
View File
@@ -68,6 +68,8 @@ work with `docker run -e`; Compose-only settings are identified below.
| Variable | Default | Purpose |
| --- | --- | --- |
| `LAYA_DEVICE` | `cpu` / `cuda` | Device selected by the base / GPU configuration |
| `LAYA_CUDA_AMP` | unset (checkpoint's `amp_dtype`) | `fp16` or `bf16` for the CUDA forward. Not cosmetic: the README's threshold section measures bf16 flipping 3 of 864 argmaxes on the parity set where fp16 flips none |
| `LAYA_CPU_AMP` | unset | `bf16` opts the CPU forward into bf16; anything else leaves it fp32 |
| `LAYA_MODEL` | `auto` | Router alias: `auto`, `english`, `multilingual`, `typed-decisions` |
| `LAYA_MODEL_PATH` | unset | Compatible checkpoint path inside the container |
| `LAYA_REQUEST_FILE` | bundled request | JSON request path inside the container |
@@ -83,7 +85,9 @@ work with `docker run -e`; Compose-only settings are identified below.
| `LAYA_TORCH_VERSION` | `2.14.0` | **Compose build:** pinned PyTorch version |
Compose forwards the runtime variables except `HF_HOME`, which stays aligned
with its fixed cache mount. If overriding `HF_HOME` in `docker run` or your own
with its fixed cache mount, and except `LAYA_MPS_AMP_MIN_ROWS`, the MPS row gate,
which no image here can reach because no container here can select MPS.
If overriding `HF_HOME` in `docker run` or your own
Compose file, provide a matching mount writable by UID 10001. Direct Docker
builds select PyTorch with `--build-arg TORCH_INDEX=cu128`; runtime `-e` cannot
change the installed wheel.