- nix/laya-serve.nix: maxLoaded, maxTokenBudget and revision options for LAYA_MAX_LOADED (#603), LAYA_MAX_TOKEN_BUDGET (#566) and LAYA_REVISION (#605); test_packaging counts mcp/device.py's _ENV_KEY as the LAYA_DEVICE read serve now goes through (#574). - docs/docker.md documents LAYA_REVISION. - test_audit_scope knows the llamaindex/crewai extras and onnxscript. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
13 KiB
Docker quickstart
Run the SDK without installing Python or PyTorch on your host. For the CPU quickstart, allow 8 GB of RAM and 10 GB of free disk, with Docker Engine or Docker Desktop and Compose v2 or newer.
From the repository root:
docker compose run --build --rm laya
This builds the checkout, runs the sample request
on CPU and prints JSON covering choice, score and noul. The first request
downloads the selected public Hugging Face checkpoint; no account is needed.
Allow several minutes for its first download.
Weights stay in a named volume. Subsequent runs use docker compose run --rm laya.
Predictions and confidence still need evaluation on your workload. See the benchmark limits.
For ARM64 hosts, DGX Spark and Apple Silicon, see ARM64 and DGX Spark containers.
NVIDIA GPU / CUDA
Install a compatible NVIDIA driver and configure Docker with the NVIDIA Container Toolkit. The GPU image uses PyTorch CUDA 12.8 wheels. Check your GPU's compute capability and driver against PyTorch's supported builds; older cards may require a different build. Allow additional disk space for CUDA layers. VRAM needs depend on the checkpoint, batch size and input length.
docker compose -f compose.yaml -f compose.cuda.yaml run --build --rm laya
The override selects GPU 0 and defaults to LAYA_DEVICE=cuda. Set
LAYA_GPU_ID to another host index or UUID. That GPU appears as device 0
inside the container. Check access without downloading weights:
docker compose -f compose.yaml -f compose.cuda.yaml run --rm laya python -c \
'import torch; assert torch.cuda.is_available(); print(torch.cuda.get_device_name(0)); print(torch.ones(1, device="cuda").cpu())'
The sample rejects unavailable CUDA before loading a checkpoint. Laya can still fall back to CPU after a memory or inference error, so inspect its warnings. Rebuild when switching between CPU and CUDA configurations.
The image sets TORCH_DISABLE_NATIVE_JIT=1. PyTorch 2.14 otherwise replaces some eager
CUDA ops with Triton kernels that it compiles on the first inference, which needs a C
compiler the slim image does not carry: the container reports healthy and then fails every
request (#365). The stock kernels give the same answers at the same latency. Set the same
variable on a bare-metal install if predict fails with Failed to find C compiler.
This uses Compose GPU reservations. Windows requires Docker Desktop's supported WSL2 GPU setup. Apple MPS, AMD/ROCm and Intel GPU containers are outside this quickstart; use CPU unless you configure and validate another backend.
Configuration
Set Compose variables in your shell, a local .env file, or the service's
environment block. Don't commit secrets in .env. Runtime variables also
work with docker run -e; Compose-only settings are identified below.
| Variable | Default | Purpose |
|---|---|---|
LAYA_DEVICE |
cpu / cuda |
Device selected by the base / GPU configuration |
LAYA_CUDA_AMP |
unset (checkpoint's amp_dtype) |
fp16 or bf16 for the CUDA forward. Not cosmetic: the README's threshold section measures bf16 flipping 3 of 864 argmaxes on the parity set where fp16 flips none |
LAYA_CPU_AMP |
unset | bf16 opts the CPU forward into bf16; anything else leaves it fp32 |
LAYA_MODEL |
auto |
Router alias: auto, english, multilingual, typed-decisions |
LAYA_MODEL_PATH |
unset | Compatible checkpoint path inside the container |
LAYA_REVISION |
unset | Hub commit, branch or tag used for every checkpoint download, or reviewed for the reviewed SHAs in laya/revisions.py; a revision= argument still wins |
LAYA_REQUEST_FILE |
bundled request | JSON request path inside the container |
OMP_NUM_THREADS |
4 |
CPU threads; keep within available cores |
HF_TOKEN / HF_TOKEN_FILE |
unset | Optional Hugging Face credential |
LAYA_API_KEY / LAYA_API_KEY_FILE |
unset | laya-serve only: require Authorization: Bearer <key> |
LAYA_PORT |
8000 |
laya-serve only: container port, and the host port published for it |
HF_HUB_OFFLINE |
0 |
1 uses only cached checkpoints |
HF_HOME |
/home/laya/.cache/huggingface |
Cache path; see mount requirement below |
LAYA_CACHE_VOLUME |
project model cache | Compose only: named cache volume |
LAYA_GPU_ID |
0 |
Compose only: NVIDIA device index or UUID |
LAYA_TORCH_INDEX |
cpu / cu128 / cu130 |
Compose build: PyTorch wheel index |
LAYA_TORCH_VERSION |
2.14.0 |
Compose build: pinned PyTorch version |
Compose forwards the runtime variables except HF_HOME, which stays aligned
with its fixed cache mount, and except LAYA_MPS_AMP_MIN_ROWS, the MPS row gate,
which no image here can reach because no container here can select MPS.
If overriding HF_HOME in docker run or your own
Compose file, provide a matching mount writable by UID 10001. Direct Docker
builds select PyTorch with --build-arg TORCH_INDEX=cu128; runtime -e cannot
change the installed wheel.
LAYA_MODEL=english OMP_NUM_THREADS=2 docker compose run --build --rm laya
docker build -t laya:local .
docker run --rm -e LAYA_MODEL=english -e OMP_NUM_THREADS=2 \
-v laya-model-cache:/home/laya/.cache/huggingface laya:local
For your own request:
docker compose run --rm --volume "$PWD/request.json:/inputs/request.json:ro" \
--env LAYA_REQUEST_FILE=/inputs/request.json laya
For a commented configuration with request, checkpoint and secret-file mounts,
see compose.example.yml:
docker compose -f compose.yaml -f compose.example.yml run --build --rm laya
Add -f compose.cuda.yaml before run for NVIDIA GPUs. The example is an
override of compose.yaml, so cache and image settings stay in one place.
Secret files
HF_TOKEN_FILE reads a mounted UTF-8 file at startup, trims surrounding
whitespace and takes precedence over HF_TOKEN. Unreadable, empty or invalid
files stop startup without printing their contents. The file must be readable
by UID 10001. _FILE applies only to supported secrets, not every setting.
With HF_TOKEN_PATH pointing to an existing host file outside the checkout:
docker compose run --rm --volume "$HF_TOKEN_PATH:/run/secrets/hf_token:ro" \
--env HF_TOKEN_FILE=/run/secrets/hf_token laya
Docker secrets or Kubernetes Secret volumes can supply the same file. Values are loaded into the process environment at startup; restart after changing a file. Never use tokens as build arguments or bake them into images. Public checkpoints need no token.
Fine-tuned checkpoints
This image runs inference. Fine-tuning happens outside it — the fine-tuning notebook runs the whole loop on Kaggle's free 2xT4 GPUs and exports a checkpoint this image can serve. Background and open questions about the training interface stay in #4 and #26.
Point LAYA_CHECKPOINT_PATH to an absolute host directory containing
rl_agent_config.json, model.safetensors and matching tokenizer files:
docker compose run --rm --volume "$LAYA_CHECKPOINT_PATH:/models/custom" \
--env LAYA_MODEL_PATH=/models/custom laya
Use a working copy writable by UID 10001 because the loader may update tokenizer
configuration. A LoRA adapter alone is not a complete checkpoint. Leave
LAYA_MODEL=auto when setting LAYA_MODEL_PATH; an explicit alias and local path
are mutually exclusive. The local-path response comes from the Agent and has no
Router routing metadata. These settings also work with the CUDA override.
Evaluate fine-tuned checkpoints on held-out examples before relying on them.
Development and cleanup
Open a Python prompt with docker compose run --rm laya python. To run the
existing routing/criteria checks and secret-file tests against your checkout
without downloading weights:
docker compose run --rm --volume "$PWD:/workspace:ro" --workdir /workspace laya \
sh -ec 'python tests/test_router.py; python tests/test_criteria.py; python tests/test_docker_entrypoint.py'
Rebuild with --build after changing source or the bundled example. The image
runs as UID/GID 10001. New named volumes inherit the image cache directory's
ownership; host directories must be writable by that UID. Keep model caches
writable for tokenizer compatibility updates.
--rm removes completed containers. docker compose down retains the cache.
To delete downloaded weights, run docker compose down --volumes using the
same Compose files and LAYA_CACHE_VOLUME setting. The next request downloads
them again; don't remove a cache shared with another project.
HTTP serving
The image ships laya-serve, so the same build that runs the one-shot quickstart can
serve the Jev-compatible API. compose.http.yaml adds it as a second service and leaves
laya alone:
docker compose -f compose.yaml -f compose.http.yaml up --build laya-serve
curl -s localhost:8000/health
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' \
--data @examples/docker/request.json
For NVIDIA, add the CUDA override. It repeats the build args and the device reservation
for laya-serve, because laya-serve is a separate service and overrides for laya
never reach it:
docker compose -f compose.yaml -f compose.http.yaml -f compose.cuda.yaml up --build laya-serve
up keeps the service running in the foreground; -d detaches. Weights go to the same
named model-cache volume as the quickstart, so serving after a quickstart run starts
with the checkpoints already on disk. Stop with docker compose ... down, using the same
Compose files.
The port is published on 127.0.0.1 only. The API has no authentication until
LAYA_API_KEY is set, so set a key before exposing it with
LAYA_BIND_ADDRESS=0.0.0.0, and put a TLS reverse proxy in front for remote clients.
/health does not require authentication in either case.
The service has a healthcheck on /health. The server preloads before it starts
listening, so with LAYA_PRELOAD=1 a healthy container has its checkpoints loaded.
docker compose ... up -d --wait laya-serve returns once it is healthy.
/health reports device as the device a resident checkpoint actually computes
on, which is not always what LAYA_DEVICE asked for: a checkpoint that wants a
GPU it cannot get falls back to CPU silently and still answers correctly.
checkpoint_devices names each loaded checkpoint, and device_is_preference is
true only while nothing is resident, so a deployment that quietly lost its GPU
says so instead of echoing its own configuration back.
Server configuration
These apply to the laya-serve service only.
| variable | default | effect |
|---|---|---|
LAYA_HOST |
0.0.0.0 |
bind address inside the container |
LAYA_PORT |
8000 |
container port, and the host port published for it |
LAYA_BIND_ADDRESS |
127.0.0.1 |
host address the port is published on |
LAYA_PRELOAD |
0 |
1 builds every checkpoint at startup instead of on first request |
LAYA_MODELS |
(all) | comma list to preload: english,multilingual,typed-decisions |
LAYA_THREADS |
OMP_NUM_THREADS |
caps torch intra-op threads; keep at or below physical cores |
LAYA_AUTO_TASK |
0 |
1 lets the router reach typed-decisions automatically |
LAYA_MAX_LOADED |
2 |
Checkpoints kept resident; LAYA_AUTO_TASK makes a third reachable on demand, and a cap below what routing chooses rebuilds one per switch |
LAYA_MAX_CONCURRENT |
16 |
requests admitted at once; later ones get 503 (a value that does not parse, or is not positive, falls back to 16) |
LAYA_LOG_LEVEL |
info |
uvicorn log level |
LAYA_API_KEY |
(none) | when set, requires Authorization: Bearer <key> |
LAYA_MAX_TOKEN_BUDGET |
8192 |
cap on per-request max_len and head_max_len overrides |
LAYA_PRELOAD defaults to 0 here rather than the package default of 1, because
preloading makes the first boot download all three checkpoints. Set it to 1 for a
long-running deployment so the first request does not pay for the build.
LAYA_PORT sets both the published host port and the port the server binds, so the two
cannot drift. Change one place to move the service:
LAYA_PORT=9000 docker compose -f compose.yaml -f compose.http.yaml up --build laya-serve
Bearer token from a file
LAYA_API_KEY_FILE is read once at startup, moved into LAYA_API_KEY, and the _FILE
variable is removed before the server execs. Prefer this to putting the key in the
environment:
docker compose -f compose.yaml -f compose.http.yaml run --rm \
--volume "$PWD/laya_api_key:/run/secrets/laya_api_key:ro" \
-e LAYA_API_KEY_FILE=/run/secrets/laya_api_key \
--service-ports laya-serve