Files
PerryLinkandPerryLink 2d2c15ce88 feat(research): add a reproducible per-language evaluation harness (#210)
Adds `research/eval/`: a small harness that points at a Laya checkpoint and
produces a per-language accuracy and calibration report plus machine-readable
per-case output. This addresses the ask in #35 -- "a fixed prompt format plus a
per-language ECE report is exactly what the repo lacks ... per-case JSON would be
very welcome too".

The repository's benchmark scripts are research code: they load every checkpoint,
run every part, and print tables. There was no small harness a third party could
point at a checkpoint, and no per-case record behind the published numbers.

Sampling, prompt and metrics follow research/scripts/bench_local.py so results are
comparable with the published tables: mteb/amazon_massive_intent test split, first
`--per-lang` rows, `random.Random(13)` created fresh per language, gold label plus
`rng.sample` of the rest, the instruction "What is the user asking for in
`utterance`?", and accuracy / macro-F1 / ECE(15 bins) / mean confidence.

`--unclamped` scores with the checkpoint's raw bucket temperatures rather than the
clamped ones Agent applies, which is what reproduces the committed sweep.

Verification against research/results/cpu_51_language_sweep.json, both checkpoints
over all 51 languages:

  english       per-language accuracy 51/51 identical, macro_accuracy 0.2269 = committed
  multilingual  per-language accuracy  6/51,           macro_accuracy 0.3661 -> 0.4008

The english checkpoint reproduces every per-language accuracy, so the sampling,
prompt, option construction and inference path match the committed run. The
english macro_ece gap (0.7331 -> 0.5709) is the #42 temperature clamp: choice:11+
is 0.1006 raw and 0.5 as served.

The multilingual row is NOT reconciled and is documented as such rather than
hidden: 45 of 51 accuracies differ, mostly upward, while macro_ece barely moves --
consistent with that checkpoint having an empty temperature_by_options, so the
clamp cannot explain it. Ruled out: the option sets (same digest as the english
run), the weights (bundled and standalone multilingual are byte-identical, all 170
tensors torch.equal), the dataset (revision 940fd47a, last modified 2026-02-24),
and build_sequence (unchanged since v0.2.0).

Also noted: head_max_len is load-bearing for accuracy, not just option truncation.
The english checkpoint at its shipped head_max_len=192 scores 0.82; forcing 256 or
512 drops it to 0.79.

research/eval/test_laya_eval.py: 47 offline tests, no checkpoint and no network.
Nothing in laya/ changes and no new required dependency is added; `datasets` is
needed to run it, which is why it lives under research/.

Co-authored-by: PerryLink <255665900+PerryLink@users.noreply.github.com>
2026-09-23 12:37:01 +05:30

0 lines
0 B
Python