research: re-run the 51-language sweep in both temperature regimes (the #208 ask) (#222)

* research: re-run the 51-language sweep in both temperature regimes (the #208 ask)

The committed `cpu_51_language_sweep.json` stores `ece` and `mean_confidence` from
before #42 clamped temperatures to `[0.5, 5]`, so its calibration columns no longer
reproduce. This adds the re-run #208 asked for, with the old file left in place so
the before/after is visible.

`research/results/cpu_51_language_sweep_clamped.json` holds the committed columns,
an `--unclamped` re-run and a default re-run side by side, per language, from the
same 51 languages and 5,100 cases.

    committed      unclamped re-run   clamped re-run
    0.2269         0.2269             0.2269     macro accuracy
    0.7331         0.7331             0.5709     macro ECE
    0.2053         0.2053             0.2053     macro F1

Macro accuracy reproduces at 0.2269 and the raw-temperature re-run reproduces the
committed macro ECE at 0.7331, so the clamp is the only variable left. Per language
the unclamped run agrees with the committed file on `accuracy` and `macro_f1` in
51/51, on `ece` in 48/51 and on `mean_confidence` in 49/51; the three that differ do
so by 0.0001, the last stored digit, because the committed run used torch 2.8.0 and
this one 2.14.0.

The clamped run differs on `ece` and `mean_confidence` in 51/51, every one lower.
`choice:11+` is the only bucket the clamp moves and every case here is a 20-option
question, so it applies to all 5,100. `accuracy` and `macro_f1` cannot move with T at
all: a temperature-scaled softmax has the same argmax at every positive temperature.

Worth recording because it cuts against the reading that the clamp only distorts the
published number: `acc_at_50_coverage` is the one rank-quality column that consumes
the confidence values, and it goes up, macro 0.3004 -> 0.3020 and `en` 0.94 -> 0.98.

The new file is guarded by `research/eval/test_laya_eval.py`, which fails if the
committed columns stopped matching the unclamped re-run or if the clamped re-run
started matching them.

Deliberate choices tested:

- `bench_local.py` was not modified. It produced the committed file, and rewriting it
  to reproduce its own pre-clamp output would make the old and new runs come from
  different scripts. `laya_eval.py` already has `--unclamped` for this, and the
  three columns it does not compute (`brier`, `nll`, `seconds`, `dropped`) are
  recorded in the new file rather than quietly dropped.
- The clamped re-run is committed alongside rather than replacing anything, so
  `cpu_51_language_sweep.json` keeps the accuracy columns other tables cite.
- `ci.yml` is untouched. #184 is adding the check that every model-free suite is
  wired into both lanes, and this suite is currently in neither; adding it here
  would collide with that PR and with the verbatim `test` job in #212.

Known limitation: only the `english` checkpoint was re-run. The committed file's
`part_b` covers the English checkpoint only, and the calibration columns #208 tracks
are in `part_a`, so the multilingual half is untouched and unmeasured here.

Verified: `python research/eval/test_laya_eval.py` -> 62 passed, 0 failed.

* research: record the multilingual eliminations for the committed-row gap

Three more candidates for the unreconciled multilingual row are now ruled out on
re-measurement rather than by inspection: the shipped head_max_len (256, read from
the checkpoint's own config and recorded in the run), which of the two multilingual
copies was measured (bundled subfolder and standalone repo both give 0.4008 / 6-of-51),
and a checkpoint change since the committed sweep (weights and config are identical by
size and hash at every revision in that window).

Also corrects the test count in the same file, which the new #208 guard changes.

---------

Co-authored-by: PerryLink <255665900+PerryLink@users.noreply.github.com>
Co-authored-by: NandhaKishorM <nandakishor@convaiinnovations.com>
This commit is contained in:
PerryLink
2026-09-23 19:13:27 +05:30
committed by GitHub
co-authored by PerryLink NandhaKishorM
parent 4419a01351
commit fcf1d7fe5d
5 changed files with 1572 additions and 4 deletions
+13 -1
View File
@@ -9,7 +9,19 @@ Every checkpoint answered **byte-identical questions** in each run (fixed seed).
| Applications | the seven workflow themes + the datasets where Jev numbers exist, all three checkpoints (laya 0.2.1, CPU, 400 cases per task, seed 13, 2026-09-19) | `research/results/app_benchmark_results.json` |
**Calibration columns in the CPU sweep predate the temperature clamp.** The 51-language ECE and mean-confidence figures were produced before #42 clamped temperatures to `[0.5, 5]`, so today's package reports different confidence for the affected buckets (`choice:11+` is now served at 0.5, not 0.1006). Accuracy columns are unaffected. A re-run with the current package is tracked in #208.
**Calibration columns in the CPU sweep predate the temperature clamp.** The 51-language ECE and mean-confidence figures were produced before #42 clamped temperatures to `[0.5, 5]`, so today's package reports different confidence for the affected buckets. Accuracy columns are unaffected, because a temperature-scaled softmax has the same argmax at every positive temperature.
The same 51 languages and 5,100 cases have now been re-run with 0.3.7 in both regimes (`research/results/cpu_51_language_sweep_clamped.json`, [#208](https://github.com/NandhaKishorM/laya/issues/208)). Macro accuracy reproduces at **0.2269** exactly, and macro ECE moves **0.7331 → 0.5709**:
| | committed | re-run, raw temperatures | re-run, as served |
|---|---|---|---|
| macro accuracy | 0.2269 | 0.2269 | 0.2269 |
| macro ECE | 0.7331 | **0.7331** | 0.5709 |
| macro F1 | 0.2053 | 0.2053 | 0.2053 |
| mean confidence, `en` | 0.9989 | **0.9989** | 0.9582 |
| ECE, `en` | 0.1789 | **0.1789** | 0.1382 |
The raw-temperature column reproduces the committed file, so the only variable left is the clamp. `choice:11+` is the sole bucket it moves, and every case in this sweep is a 20-option question, so the clamp applies to all 5,100 — and lowers ECE in all 51 languages. `acc_at_50_coverage` is the one rank-quality column that uses the confidence values: macro 0.3004 → 0.3020, and `en` 0.94 → 0.98, so the flatter distribution selects a slightly better half rather than a worse one.
---
+1
View File
@@ -27,6 +27,7 @@ installed its abseil runtime can deadlock model construction on macOS/Python 3.9
|---|---|
| `results/t4_colab_benchmark.json` | 17,416 questions on one T4, both checkpoints, identical questions per model |
| `results/cpu_51_language_sweep.json` | 51 languages x 2 checkpoints, MASSIVE intent, 20 options |
| `results/cpu_51_language_sweep_clamped.json` | the same 51 languages and 5,100 cases re-run with 0.3.7, raw temperatures and served temperatures side by side ([#208](https://github.com/NandhaKishorM/laya/issues/208)) |
## Headline findings
+41 -3
View File
@@ -137,8 +137,24 @@ having an empty `temperature_by_options`, so the clamp cannot explain it.
Ruled out: the option sets (identical digest to the english run), the weights
(bundled and standalone multilingual are byte-identical, all 170 tensors
`torch.equal`), the dataset (revision `940fd47a`, last modified 2026-02-24), and
`build_sequence` (unchanged since `v0.2.0`). It is in the multilingual inference path
between `laya 0.2.0` and `0.3.6` and is **not** reconciled. Flagged rather than hidden.
`build_sequence` (unchanged since `v0.2.0`). Also ruled out, on re-measurement:
* **the shipped `head_max_len`**, which matters here because this checkpoint ships
`256` and english ships `192`. The harness reads it from the checkpoint's own
config and the run's `config` block records `head_max_len: 256, max_len: 1024`, so
the multilingual numbers above were not taken at english's budget. Re-running with
the value read from config gives the same `0.4008`, and `6/51` again.
* **which of the two multilingual copies was measured.** The bundled `multilingual/`
subfolder and the standalone `convaiinnovations/laya-multilingual` repo were each
run end to end over all 51 languages and both give `macro_accuracy 0.4008`,
`macro_ece 0.3911`, `6/51`.
* **a checkpoint change since the committed sweep.** `multilingual/model.safetensors`
is `643835514` bytes at `sha256 b99c8bea…` and `multilingual/rl_agent_config.json`
is `472` bytes at `sha256 00e35f88…` at every revision from the sweep's timestamp to
today; the Hub commits in that window are model-card `docs:`/`assets:` only.
It is in the multilingual inference path between `laya 0.2.0` and `0.3.6` and is
**not** reconciled. Flagged rather than hidden.
Related: **`head_max_len` is load-bearing for accuracy**, not just for option
truncation. The english checkpoint at its shipped `head_max_len=192` scores 0.82;
@@ -150,7 +166,7 @@ forcing 256 or 512 drops it to 0.79.
checkpoint, no network:
```bash
python research/eval/test_laya_eval.py # 49 passed, 0 failed
python research/eval/test_laya_eval.py # 64 passed, 0 failed
```
It pins the upstream constants (seed 13, 20 options, the exact instruction string),
@@ -176,3 +192,25 @@ not just against this harness's own arithmetic.
binned differently. It now matches `laya.common.ece_score`,
`research/scripts/bench_local.py` and `research/scripts/build_benchmark_nb.py`, and
`test_laya_eval.py` asserts that agreement.
### The temperature clamp, measured both ways
`research/results/cpu_51_language_sweep_clamped.json` carries the same re-run twice, once per
regime, against the committed columns. Macro accuracy reproduces the committed file exactly
and macro ECE is the only macro figure that moves:
| | committed | `--unclamped` | default |
|---|---|---|---|
| `macro_accuracy` | 0.2269 | **0.2269** | 0.2269 |
| `macro_ece` | 0.7331 | **0.7331** | 0.5709 |
| `macro_f1` | 0.2053 | **0.2053** | 0.2053 |
Per language, the unclamped run agrees with the committed file on `accuracy` and `macro_f1`
in **51/51**, on `ece` in **48/51** and on `mean_confidence` in **49/51**. The handful that
differ do so by `0.0001`, the last stored digit: the committed run used torch 2.8.0 and this
one 2.14.0. The clamped run differs from the committed file on `ece` and `mean_confidence` in
**51/51**, every one of them lower, because it is the only column the clamp can move.
`accuracy`, `macro_f1` and `n` are identical in all three columns by construction: scaling
logits by any positive temperature does not change the argmax. That is why a re-run can settle
the calibration question without reopening the accuracy numbers.
+84
View File
@@ -160,6 +160,90 @@ check("const/instructions match bench_local.py",
INSTRUCTIONS, "What is the user asking for in `utterance`?")
# ------------------------------------------- the #208 before/after re-run file
# research/results/cpu_51_language_sweep_clamped.json records the committed sweep,
# the pre-clamp re-run and the served-temperature re-run side by side. It is only
# useful if it still agrees with the committed file, so that agreement is a test.
import json # noqa: E402
_RESULTS = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))),
"research", "results")
def _load(name):
with open(os.path.join(_RESULTS, name), encoding="utf-8") as fh:
return json.load(fh)
_rerun = _load("cpu_51_language_sweep_clamped.json")
_sweep = _load("cpu_51_language_sweep.json")["part_a"]["by_model"]["english"]
_langs = _sweep["per_language"]
# Both files round each per-language figure to 4 decimals (bench_local.py:138-144), so
# agreement has to be judged at that resolution: two files can disagree by 1 in the
# last stored digit for reasons that have nothing to do with the temperatures, and
# three languages do (am, el, ro) because the committed run used torch 2.8.0 and this
# one 2.14.0. The macro figures match exactly, which is why that is the headline.
_QUANT = 1.5e-4
def _close(a, b, tol=_QUANT):
return abs(a - b) <= tol
check("208/one entry per language", len(_rerun["per_language"]), len(_langs))
check("208/case count is languages x per_lang",
_rerun["config"]["n_cases"],
_rerun["config"]["languages"] * _rerun["config"]["per_lang"])
check("208/macro accuracy copied from the committed sweep",
_rerun["macro"]["committed"]["accuracy"], _sweep["macro_accuracy"])
check("208/macro ece copied from the committed sweep",
_rerun["macro"]["committed"]["ece"], _sweep["macro_ece"])
# accuracy is argmax of a temperature-scaled softmax, so it cannot move with T
check_true("208/accuracy identical in all three regimes",
all(len({_rerun["per_language"][lg][r]["accuracy"] for r in
("committed", "unclamped_rerun", "clamped_rerun")}) == 1
for lg in _langs))
check_true("208/macro_f1 identical in all three regimes",
all(len({_rerun["per_language"][lg][r]["macro_f1"] for r in
("committed", "unclamped_rerun", "clamped_rerun")}) == 1
for lg in _langs))
# the committed calibration columns must match the unclamped re-run, and only the
# clamped re-run may differ -- that is the whole claim
check_true("208/ece matches the committed file in the unclamped re-run",
all(_close(_rerun["per_language"][lg]["unclamped_rerun"]["ece"],
_langs[lg]["ece"]) for lg in _langs))
check_true("208/ece differs from the committed file in the clamped re-run",
all(not _close(_rerun["per_language"][lg]["clamped_rerun"]["ece"],
_langs[lg]["ece"]) for lg in _langs))
check_true("208/mean_confidence matches in the unclamped re-run",
all(_close(_rerun["per_language"][lg]["unclamped_rerun"]["mean_confidence"],
_langs[lg]["mean_confidence"]) for lg in _langs))
check_true("208/mean_confidence differs in the clamped re-run",
all(not _close(_rerun["per_language"][lg]["clamped_rerun"]["mean_confidence"],
_langs[lg]["mean_confidence"]) for lg in _langs))
# the unclamped re-run is a reproduction, not an approximation: name the tolerance
# so a future change that widens it has to say so
check_true("208/unclamped reproduction agrees in at least 48 of 51 languages",
sum(1 for lg in _langs
if _close(_rerun["per_language"][lg]["unclamped_rerun"]["ece"], _langs[lg]["ece"])) >= 48)
# the clamp can lower ECE everywhere without lowering rank quality, so this is a
# guard against the misleading "the clamp makes it worse" reading
check_true("208/the clamp lowers macro ece",
_rerun["macro"]["clamped_rerun"]["macro_ece"]
< _rerun["macro"]["unclamped_rerun"]["macro_ece"])
check_true("208/every per-language delta is reported",
all("delta_ece" in v and "delta_mean_confidence" in v
for v in _rerun["per_language"].values()))
check("208/only choice:11+ is the clamped bucket",
_rerun["temperature_choice_11_plus"]["unclamped_rerun"], 0.10058280825614929)
check("208/the served bucket is 0.5",
_rerun["temperature_choice_11_plus"]["clamped_rerun"], 0.5)
print("\n%d passed, %d failed" % (len(PASS), len(FAIL)))
for f in FAIL:
print(" FAIL " + f)
File diff suppressed because it is too large Load Diff