* feat: self-hostable server + Nix flake (Jev-compatible /v1/systemone)
Laya's predict() output is already schema-compatible with TypeSafe Jev's
decision API, so this adds only the HTTP surface a Jev client needs to talk
to a self-hosted Laya:
- laya/serve.py: FastAPI app exposing POST /v1/systemone (+ /health), calling
Router.predict, with optional LAYA_API_KEY bearer auth. Config via env
(LAYA_HOST/PORT/DEVICE/PRELOAD/MODELS/AUTO_TASK). Heavy imports are deferred
so `import laya.serve` stays GPU- and web-dep-free.
- laya-serve console script + laya[serve] extra (fastapi, uvicorn).
- flake.nix: overlay building laya from torch-bin (prebuilt CUDA), a
laya-serve runner, a dev shell, and nixosModules.default.
- nix/laya-serve.nix: hardened DynamicUser systemd service (services.laya-serve)
with CUDA device access and a HF weight cache.
- tests/test_serve.py: shim tests (injected router, real FastAPI TestClient).
Verified end-to-end on an RTX 3090: all three checkpoints preload on CUDA,
English/Hindi requests route correctly and return Jev-shaped payloads, 6/6
tests pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019NnoUnPhBMoS4S7eYWHYfe
* fix(nix): build the service package from the host pkgs, not an overlay
Hosts that inject `pkgs` via specialArgs (read-only nixpkgs) ignore
module-level `nixpkgs.overlays`, so `pkgs.laya-serve` was missing and the
NixOS toplevel failed to evaluate. Factor the package into nix/package.nix and
have services.laya-serve.package default to `(pkgs.callPackage ./package.nix
{}).laya-serve`, built from the host's own pkgs. The flake overlay/packages now
reuse the same file.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019NnoUnPhBMoS4S7eYWHYfe
* feat(serve): LAYA_THREADS knob + services.laya-serve.threads for CPU tuning
Cap torch intra-op threads for CPU inference. serve.py reads LAYA_THREADS and
calls torch.set_num_threads (torch imported only when a limit is set); the NixOS
module exposes `threads` and wires it to LAYA_THREADS + OMP_NUM_THREADS.
Motivated by a thread sweep on a 32-core/64-thread EPYC: oversubscribing the
logical core count is a ~4x latency regression, and single-request latency is
often best below the physical core count (e.g. 8-16), while batched throughput
peaks at the physical count.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019NnoUnPhBMoS4S7eYWHYfe
* chore(nix): commit flake.lock pinning nixpkgs to ccad53c
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019NnoUnPhBMoS4S7eYWHYfe
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>