mirror of
https://github.com/NandhaKishorM/laya.git
synced 2026-09-28 07:52:57 +08:00
fix(compose): forward the autocast dtype overrides the runtime reads
laya/agent.py selects the forward's dtype from LAYA_CUDA_AMP and LAYA_CPU_AMP, and no compose file that sets LAYA_DEVICE passed either name, so the published deployment always served the checkpoint's amp_dtype whatever the operator exported. README measures the difference as decisions, not rounding: bf16 flips 3 of 864 argmaxes on the parity_fast set where fp16 flips none, at the same latency. Both GPU services get the pair, because an override on `laya` never reaches `laya-serve`. An empty value means "the checkpoint's own amp_dtype", which is what these images shipped before, so the passthrough default changes nothing until it is set. The MPS row gate is deliberately left out: no container here can select MPS.
This commit is contained in:
@@ -8,6 +8,10 @@ services:
|
||||
TORCH_INDEX: "${LAYA_TORCH_INDEX:-cu128}"
|
||||
environment:
|
||||
LAYA_DEVICE: "${LAYA_DEVICE:-cuda}"
|
||||
# The runtime reads these when it picks the autocast dtype; empty keeps the checkpoint's
|
||||
# own `amp_dtype`. README's threshold section measures what fp16 vs bf16 decides differently.
|
||||
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
|
||||
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
@@ -23,6 +27,10 @@ services:
|
||||
TORCH_INDEX: "${LAYA_TORCH_INDEX:-cu128}"
|
||||
environment:
|
||||
LAYA_DEVICE: "${LAYA_DEVICE:-cuda}"
|
||||
# The runtime reads these when it picks the autocast dtype; empty keeps the checkpoint's
|
||||
# own `amp_dtype`. README's threshold section measures what fp16 vs bf16 decides differently.
|
||||
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
|
||||
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
|
||||
@@ -37,6 +37,10 @@ services:
|
||||
LAYA_HOST: "${LAYA_HOST:-0.0.0.0}"
|
||||
LAYA_PORT: "${LAYA_PORT:-8000}"
|
||||
LAYA_DEVICE: "${LAYA_DEVICE:-cpu}"
|
||||
# Autocast dtype the runtime reads when building the model; empty keeps the checkpoint's own
|
||||
# `amp_dtype`. Set LAYA_CUDA_AMP=fp16 to serve fp16 on a GPU host.
|
||||
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
|
||||
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
|
||||
LAYA_PRELOAD: "${LAYA_PRELOAD:-0}"
|
||||
LAYA_MODELS: "${LAYA_MODELS:-}"
|
||||
LAYA_THREADS: "${LAYA_THREADS:-${OMP_NUM_THREADS:-4}}"
|
||||
|
||||
@@ -13,6 +13,10 @@ services:
|
||||
TORCH_INDEX: "${LAYA_TORCH_INDEX:-cu130}"
|
||||
environment:
|
||||
LAYA_DEVICE: "${LAYA_DEVICE:-cuda}"
|
||||
# Autocast dtype the runtime reads; empty keeps the checkpoint's own `amp_dtype`. Both
|
||||
# services need it, because an override on `laya` never reaches `laya-serve`.
|
||||
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
|
||||
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
@@ -29,6 +33,10 @@ services:
|
||||
TORCH_INDEX: "${LAYA_TORCH_INDEX:-cu130}"
|
||||
environment:
|
||||
LAYA_DEVICE: "${LAYA_DEVICE:-cuda}"
|
||||
# Autocast dtype the runtime reads; empty keeps the checkpoint's own `amp_dtype`. Both
|
||||
# services need it, because an override on `laya` never reaches `laya-serve`.
|
||||
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
|
||||
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
|
||||
@@ -7,6 +7,12 @@ services:
|
||||
TORCH_VERSION: "${LAYA_TORCH_VERSION:-2.14.0}"
|
||||
environment:
|
||||
LAYA_DEVICE: "${LAYA_DEVICE:-cpu}"
|
||||
# The autocast dtype the runtime picks when it builds the model on the selected device
|
||||
# (`laya/agent.py`). An empty value means "use the checkpoint's own `amp_dtype`", which is
|
||||
# what these images have always served. README's threshold section measures what fp16 vs
|
||||
# bf16 decides differently.
|
||||
LAYA_CUDA_AMP: "${LAYA_CUDA_AMP:-}"
|
||||
LAYA_CPU_AMP: "${LAYA_CPU_AMP:-}"
|
||||
LAYA_MODEL: "${LAYA_MODEL:-auto}"
|
||||
LAYA_MODEL_PATH: "${LAYA_MODEL_PATH:-}"
|
||||
LAYA_REQUEST_FILE: "${LAYA_REQUEST_FILE:-/opt/laya/examples/request.json}"
|
||||
|
||||
+5
-1
@@ -68,6 +68,8 @@ work with `docker run -e`; Compose-only settings are identified below.
|
||||
| Variable | Default | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `LAYA_DEVICE` | `cpu` / `cuda` | Device selected by the base / GPU configuration |
|
||||
| `LAYA_CUDA_AMP` | unset (checkpoint's `amp_dtype`) | `fp16` or `bf16` for the CUDA forward. Not cosmetic: the README's threshold section measures bf16 flipping 3 of 864 argmaxes on the parity set where fp16 flips none |
|
||||
| `LAYA_CPU_AMP` | unset | `bf16` opts the CPU forward into bf16; anything else leaves it fp32 |
|
||||
| `LAYA_MODEL` | `auto` | Router alias: `auto`, `english`, `multilingual`, `typed-decisions` |
|
||||
| `LAYA_MODEL_PATH` | unset | Compatible checkpoint path inside the container |
|
||||
| `LAYA_REQUEST_FILE` | bundled request | JSON request path inside the container |
|
||||
@@ -83,7 +85,9 @@ work with `docker run -e`; Compose-only settings are identified below.
|
||||
| `LAYA_TORCH_VERSION` | `2.14.0` | **Compose build:** pinned PyTorch version |
|
||||
|
||||
Compose forwards the runtime variables except `HF_HOME`, which stays aligned
|
||||
with its fixed cache mount. If overriding `HF_HOME` in `docker run` or your own
|
||||
with its fixed cache mount, and except `LAYA_MPS_AMP_MIN_ROWS`, the MPS row gate,
|
||||
which no image here can reach because no container here can select MPS.
|
||||
If overriding `HF_HOME` in `docker run` or your own
|
||||
Compose file, provide a matching mount writable by UID 10001. Direct Docker
|
||||
builds select PyTorch with `--build-arg TORCH_INDEX=cu128`; runtime `-e` cannot
|
||||
change the installed wheel.
|
||||
|
||||
Reference in New Issue
Block a user