mirror of
https://github.com/NandhaKishorM/laya.git
synced 2026-09-28 07:52:57 +08:00
Adds `research/eval/`: a small harness that points at a Laya checkpoint and produces a per-language accuracy and calibration report plus machine-readable per-case output. This addresses the ask in #35 -- "a fixed prompt format plus a per-language ECE report is exactly what the repo lacks ... per-case JSON would be very welcome too". The repository's benchmark scripts are research code: they load every checkpoint, run every part, and print tables. There was no small harness a third party could point at a checkpoint, and no per-case record behind the published numbers. Sampling, prompt and metrics follow research/scripts/bench_local.py so results are comparable with the published tables: mteb/amazon_massive_intent test split, first `--per-lang` rows, `random.Random(13)` created fresh per language, gold label plus `rng.sample` of the rest, the instruction "What is the user asking for in `utterance`?", and accuracy / macro-F1 / ECE(15 bins) / mean confidence. `--unclamped` scores with the checkpoint's raw bucket temperatures rather than the clamped ones Agent applies, which is what reproduces the committed sweep. Verification against research/results/cpu_51_language_sweep.json, both checkpoints over all 51 languages: english per-language accuracy 51/51 identical, macro_accuracy 0.2269 = committed multilingual per-language accuracy 6/51, macro_accuracy 0.3661 -> 0.4008 The english checkpoint reproduces every per-language accuracy, so the sampling, prompt, option construction and inference path match the committed run. The english macro_ece gap (0.7331 -> 0.5709) is the #42 temperature clamp: choice:11+ is 0.1006 raw and 0.5 as served. The multilingual row is NOT reconciled and is documented as such rather than hidden: 45 of 51 accuracies differ, mostly upward, while macro_ece barely moves -- consistent with that checkpoint having an empty temperature_by_options, so the clamp cannot explain it. Ruled out: the option sets (same digest as the english run), the weights (bundled and standalone multilingual are byte-identical, all 170 tensors torch.equal), the dataset (revision 940fd47a, last modified 2026-02-24), and build_sequence (unchanged since v0.2.0). Also noted: head_max_len is load-bearing for accuracy, not just option truncation. The english checkpoint at its shipped head_max_len=192 scores 0.82; forcing 256 or 512 drops it to 0.79. research/eval/test_laya_eval.py: 47 offline tests, no checkpoint and no network. Nothing in laya/ changes and no new required dependency is added; `datasets` is needed to run it, which is why it lives under research/. Co-authored-by: PerryLink <255665900+PerryLink@users.noreply.github.com>
0 lines
0 B
Python
0 lines
0 B
Python
The file is empty.