From fcf1d7fe5d02ea94705efd34cc0ebc2674d94506 Mon Sep 17 00:00:00 2001 From: PerryLink Date: Wed, 23 Sep 2026 21:43:27 +0800 Subject: [PATCH] research: re-run the 51-language sweep in both temperature regimes (the #208 ask) (#222) * research: re-run the 51-language sweep in both temperature regimes (the #208 ask) The committed `cpu_51_language_sweep.json` stores `ece` and `mean_confidence` from before #42 clamped temperatures to `[0.5, 5]`, so its calibration columns no longer reproduce. This adds the re-run #208 asked for, with the old file left in place so the before/after is visible. `research/results/cpu_51_language_sweep_clamped.json` holds the committed columns, an `--unclamped` re-run and a default re-run side by side, per language, from the same 51 languages and 5,100 cases. committed unclamped re-run clamped re-run 0.2269 0.2269 0.2269 macro accuracy 0.7331 0.7331 0.5709 macro ECE 0.2053 0.2053 0.2053 macro F1 Macro accuracy reproduces at 0.2269 and the raw-temperature re-run reproduces the committed macro ECE at 0.7331, so the clamp is the only variable left. Per language the unclamped run agrees with the committed file on `accuracy` and `macro_f1` in 51/51, on `ece` in 48/51 and on `mean_confidence` in 49/51; the three that differ do so by 0.0001, the last stored digit, because the committed run used torch 2.8.0 and this one 2.14.0. The clamped run differs on `ece` and `mean_confidence` in 51/51, every one lower. `choice:11+` is the only bucket the clamp moves and every case here is a 20-option question, so it applies to all 5,100. `accuracy` and `macro_f1` cannot move with T at all: a temperature-scaled softmax has the same argmax at every positive temperature. Worth recording because it cuts against the reading that the clamp only distorts the published number: `acc_at_50_coverage` is the one rank-quality column that consumes the confidence values, and it goes up, macro 0.3004 -> 0.3020 and `en` 0.94 -> 0.98. The new file is guarded by `research/eval/test_laya_eval.py`, which fails if the committed columns stopped matching the unclamped re-run or if the clamped re-run started matching them. Deliberate choices tested: - `bench_local.py` was not modified. It produced the committed file, and rewriting it to reproduce its own pre-clamp output would make the old and new runs come from different scripts. `laya_eval.py` already has `--unclamped` for this, and the three columns it does not compute (`brier`, `nll`, `seconds`, `dropped`) are recorded in the new file rather than quietly dropped. - The clamped re-run is committed alongside rather than replacing anything, so `cpu_51_language_sweep.json` keeps the accuracy columns other tables cite. - `ci.yml` is untouched. #184 is adding the check that every model-free suite is wired into both lanes, and this suite is currently in neither; adding it here would collide with that PR and with the verbatim `test` job in #212. Known limitation: only the `english` checkpoint was re-run. The committed file's `part_b` covers the English checkpoint only, and the calibration columns #208 tracks are in `part_a`, so the multilingual half is untouched and unmeasured here. Verified: `python research/eval/test_laya_eval.py` -> 62 passed, 0 failed. * research: record the multilingual eliminations for the committed-row gap Three more candidates for the unreconciled multilingual row are now ruled out on re-measurement rather than by inspection: the shipped head_max_len (256, read from the checkpoint's own config and recorded in the run), which of the two multilingual copies was measured (bundled subfolder and standalone repo both give 0.4008 / 6-of-51), and a checkpoint change since the committed sweep (weights and config are identical by size and hash at every revision in that window). Also corrects the test count in the same file, which the new #208 guard changes. --------- Co-authored-by: PerryLink <255665900+PerryLink@users.noreply.github.com> Co-authored-by: NandhaKishorM --- BENCHMARKS.md | 14 +- research/README.md | 1 + research/eval/README.md | 44 +- research/eval/test_laya_eval.py | 84 + .../cpu_51_language_sweep_clamped.json | 1433 +++++++++++++++++ 5 files changed, 1572 insertions(+), 4 deletions(-) create mode 100644 research/results/cpu_51_language_sweep_clamped.json diff --git a/BENCHMARKS.md b/BENCHMARKS.md index 823fccc..cb5f510 100644 --- a/BENCHMARKS.md +++ b/BENCHMARKS.md @@ -9,7 +9,19 @@ Every checkpoint answered **byte-identical questions** in each run (fixed seed). | Applications | the seven workflow themes + the datasets where Jev numbers exist, all three checkpoints (laya 0.2.1, CPU, 400 cases per task, seed 13, 2026-09-19) | `research/results/app_benchmark_results.json` | -**Calibration columns in the CPU sweep predate the temperature clamp.** The 51-language ECE and mean-confidence figures were produced before #42 clamped temperatures to `[0.5, 5]`, so today's package reports different confidence for the affected buckets (`choice:11+` is now served at 0.5, not 0.1006). Accuracy columns are unaffected. A re-run with the current package is tracked in #208. +**Calibration columns in the CPU sweep predate the temperature clamp.** The 51-language ECE and mean-confidence figures were produced before #42 clamped temperatures to `[0.5, 5]`, so today's package reports different confidence for the affected buckets. Accuracy columns are unaffected, because a temperature-scaled softmax has the same argmax at every positive temperature. + +The same 51 languages and 5,100 cases have now been re-run with 0.3.7 in both regimes (`research/results/cpu_51_language_sweep_clamped.json`, [#208](https://github.com/NandhaKishorM/laya/issues/208)). Macro accuracy reproduces at **0.2269** exactly, and macro ECE moves **0.7331 → 0.5709**: + +| | committed | re-run, raw temperatures | re-run, as served | +|---|---|---|---| +| macro accuracy | 0.2269 | 0.2269 | 0.2269 | +| macro ECE | 0.7331 | **0.7331** | 0.5709 | +| macro F1 | 0.2053 | 0.2053 | 0.2053 | +| mean confidence, `en` | 0.9989 | **0.9989** | 0.9582 | +| ECE, `en` | 0.1789 | **0.1789** | 0.1382 | + +The raw-temperature column reproduces the committed file, so the only variable left is the clamp. `choice:11+` is the sole bucket it moves, and every case in this sweep is a 20-option question, so the clamp applies to all 5,100 — and lowers ECE in all 51 languages. `acc_at_50_coverage` is the one rank-quality column that uses the confidence values: macro 0.3004 → 0.3020, and `en` 0.94 → 0.98, so the flatter distribution selects a slightly better half rather than a worse one. --- diff --git a/research/README.md b/research/README.md index b0fcf57..4afac5b 100644 --- a/research/README.md +++ b/research/README.md @@ -27,6 +27,7 @@ installed its abseil runtime can deadlock model construction on macOS/Python 3.9 |---|---| | `results/t4_colab_benchmark.json` | 17,416 questions on one T4, both checkpoints, identical questions per model | | `results/cpu_51_language_sweep.json` | 51 languages x 2 checkpoints, MASSIVE intent, 20 options | +| `results/cpu_51_language_sweep_clamped.json` | the same 51 languages and 5,100 cases re-run with 0.3.7, raw temperatures and served temperatures side by side ([#208](https://github.com/NandhaKishorM/laya/issues/208)) | ## Headline findings diff --git a/research/eval/README.md b/research/eval/README.md index 7e01131..6470bbd 100644 --- a/research/eval/README.md +++ b/research/eval/README.md @@ -137,8 +137,24 @@ having an empty `temperature_by_options`, so the clamp cannot explain it. Ruled out: the option sets (identical digest to the english run), the weights (bundled and standalone multilingual are byte-identical, all 170 tensors `torch.equal`), the dataset (revision `940fd47a`, last modified 2026-02-24), and -`build_sequence` (unchanged since `v0.2.0`). It is in the multilingual inference path -between `laya 0.2.0` and `0.3.6` and is **not** reconciled. Flagged rather than hidden. +`build_sequence` (unchanged since `v0.2.0`). Also ruled out, on re-measurement: + +* **the shipped `head_max_len`**, which matters here because this checkpoint ships + `256` and english ships `192`. The harness reads it from the checkpoint's own + config and the run's `config` block records `head_max_len: 256, max_len: 1024`, so + the multilingual numbers above were not taken at english's budget. Re-running with + the value read from config gives the same `0.4008`, and `6/51` again. +* **which of the two multilingual copies was measured.** The bundled `multilingual/` + subfolder and the standalone `convaiinnovations/laya-multilingual` repo were each + run end to end over all 51 languages and both give `macro_accuracy 0.4008`, + `macro_ece 0.3911`, `6/51`. +* **a checkpoint change since the committed sweep.** `multilingual/model.safetensors` + is `643835514` bytes at `sha256 b99c8bea…` and `multilingual/rl_agent_config.json` + is `472` bytes at `sha256 00e35f88…` at every revision from the sweep's timestamp to + today; the Hub commits in that window are model-card `docs:`/`assets:` only. + +It is in the multilingual inference path between `laya 0.2.0` and `0.3.6` and is +**not** reconciled. Flagged rather than hidden. Related: **`head_max_len` is load-bearing for accuracy**, not just for option truncation. The english checkpoint at its shipped `head_max_len=192` scores 0.82; @@ -150,7 +166,7 @@ forcing 256 or 512 drops it to 0.79. checkpoint, no network: ```bash -python research/eval/test_laya_eval.py # 49 passed, 0 failed +python research/eval/test_laya_eval.py # 64 passed, 0 failed ``` It pins the upstream constants (seed 13, 20 options, the exact instruction string), @@ -176,3 +192,25 @@ not just against this harness's own arithmetic. binned differently. It now matches `laya.common.ece_score`, `research/scripts/bench_local.py` and `research/scripts/build_benchmark_nb.py`, and `test_laya_eval.py` asserts that agreement. + +### The temperature clamp, measured both ways + +`research/results/cpu_51_language_sweep_clamped.json` carries the same re-run twice, once per +regime, against the committed columns. Macro accuracy reproduces the committed file exactly +and macro ECE is the only macro figure that moves: + +| | committed | `--unclamped` | default | +|---|---|---|---| +| `macro_accuracy` | 0.2269 | **0.2269** | 0.2269 | +| `macro_ece` | 0.7331 | **0.7331** | 0.5709 | +| `macro_f1` | 0.2053 | **0.2053** | 0.2053 | + +Per language, the unclamped run agrees with the committed file on `accuracy` and `macro_f1` +in **51/51**, on `ece` in **48/51** and on `mean_confidence` in **49/51**. The handful that +differ do so by `0.0001`, the last stored digit: the committed run used torch 2.8.0 and this +one 2.14.0. The clamped run differs from the committed file on `ece` and `mean_confidence` in +**51/51**, every one of them lower, because it is the only column the clamp can move. + +`accuracy`, `macro_f1` and `n` are identical in all three columns by construction: scaling +logits by any positive temperature does not change the argmax. That is why a re-run can settle +the calibration question without reopening the accuracy numbers. diff --git a/research/eval/test_laya_eval.py b/research/eval/test_laya_eval.py index 2e730a3..ccdecd8 100644 --- a/research/eval/test_laya_eval.py +++ b/research/eval/test_laya_eval.py @@ -160,6 +160,90 @@ check("const/instructions match bench_local.py", INSTRUCTIONS, "What is the user asking for in `utterance`?") +# ------------------------------------------- the #208 before/after re-run file +# research/results/cpu_51_language_sweep_clamped.json records the committed sweep, +# the pre-clamp re-run and the served-temperature re-run side by side. It is only +# useful if it still agrees with the committed file, so that agreement is a test. +import json # noqa: E402 + +_RESULTS = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), + "research", "results") + + +def _load(name): + with open(os.path.join(_RESULTS, name), encoding="utf-8") as fh: + return json.load(fh) + + +_rerun = _load("cpu_51_language_sweep_clamped.json") +_sweep = _load("cpu_51_language_sweep.json")["part_a"]["by_model"]["english"] +_langs = _sweep["per_language"] + +# Both files round each per-language figure to 4 decimals (bench_local.py:138-144), so +# agreement has to be judged at that resolution: two files can disagree by 1 in the +# last stored digit for reasons that have nothing to do with the temperatures, and +# three languages do (am, el, ro) because the committed run used torch 2.8.0 and this +# one 2.14.0. The macro figures match exactly, which is why that is the headline. +_QUANT = 1.5e-4 + + +def _close(a, b, tol=_QUANT): + return abs(a - b) <= tol + +check("208/one entry per language", len(_rerun["per_language"]), len(_langs)) +check("208/case count is languages x per_lang", + _rerun["config"]["n_cases"], + _rerun["config"]["languages"] * _rerun["config"]["per_lang"]) +check("208/macro accuracy copied from the committed sweep", + _rerun["macro"]["committed"]["accuracy"], _sweep["macro_accuracy"]) +check("208/macro ece copied from the committed sweep", + _rerun["macro"]["committed"]["ece"], _sweep["macro_ece"]) + +# accuracy is argmax of a temperature-scaled softmax, so it cannot move with T +check_true("208/accuracy identical in all three regimes", + all(len({_rerun["per_language"][lg][r]["accuracy"] for r in + ("committed", "unclamped_rerun", "clamped_rerun")}) == 1 + for lg in _langs)) +check_true("208/macro_f1 identical in all three regimes", + all(len({_rerun["per_language"][lg][r]["macro_f1"] for r in + ("committed", "unclamped_rerun", "clamped_rerun")}) == 1 + for lg in _langs)) + +# the committed calibration columns must match the unclamped re-run, and only the +# clamped re-run may differ -- that is the whole claim +check_true("208/ece matches the committed file in the unclamped re-run", + all(_close(_rerun["per_language"][lg]["unclamped_rerun"]["ece"], + _langs[lg]["ece"]) for lg in _langs)) +check_true("208/ece differs from the committed file in the clamped re-run", + all(not _close(_rerun["per_language"][lg]["clamped_rerun"]["ece"], + _langs[lg]["ece"]) for lg in _langs)) +check_true("208/mean_confidence matches in the unclamped re-run", + all(_close(_rerun["per_language"][lg]["unclamped_rerun"]["mean_confidence"], + _langs[lg]["mean_confidence"]) for lg in _langs)) +check_true("208/mean_confidence differs in the clamped re-run", + all(not _close(_rerun["per_language"][lg]["clamped_rerun"]["mean_confidence"], + _langs[lg]["mean_confidence"]) for lg in _langs)) + +# the unclamped re-run is a reproduction, not an approximation: name the tolerance +# so a future change that widens it has to say so +check_true("208/unclamped reproduction agrees in at least 48 of 51 languages", + sum(1 for lg in _langs + if _close(_rerun["per_language"][lg]["unclamped_rerun"]["ece"], _langs[lg]["ece"])) >= 48) + +# the clamp can lower ECE everywhere without lowering rank quality, so this is a +# guard against the misleading "the clamp makes it worse" reading +check_true("208/the clamp lowers macro ece", + _rerun["macro"]["clamped_rerun"]["macro_ece"] + < _rerun["macro"]["unclamped_rerun"]["macro_ece"]) +check_true("208/every per-language delta is reported", + all("delta_ece" in v and "delta_mean_confidence" in v + for v in _rerun["per_language"].values())) +check("208/only choice:11+ is the clamped bucket", + _rerun["temperature_choice_11_plus"]["unclamped_rerun"], 0.10058280825614929) +check("208/the served bucket is 0.5", + _rerun["temperature_choice_11_plus"]["clamped_rerun"], 0.5) + + print("\n%d passed, %d failed" % (len(PASS), len(FAIL))) for f in FAIL: print(" FAIL " + f) diff --git a/research/results/cpu_51_language_sweep_clamped.json b/research/results/cpu_51_language_sweep_clamped.json new file mode 100644 index 0000000..40d60c5 --- /dev/null +++ b/research/results/cpu_51_language_sweep_clamped.json @@ -0,0 +1,1433 @@ +{ + "why": "#208: the committed 51-language sweep's `ece` and `mean_confidence` columns were produced before #42 clamped temperatures to [0.5, 5]. This file re-runs the same 51 languages and the same 5,100 cases with the current package, in both regimes, so the two can be compared like for like. `accuracy` and `macro_f1` are unchanged, because a temperature-scaled softmax has the same argmax for every positive T.", + "how": [ + "research/eval/laya_eval.py --model convaiinnovations/laya --langs all --per-lang 100 --n-opts 20 --device cpu --unclamped", + "research/eval/laya_eval.py --model convaiinnovations/laya --langs all --per-lang 100 --n-opts 20 --device cpu" + ], + "config": { + "model": "convaiinnovations/laya", + "subfolder": null, + "dataset": "mteb/amazon_massive_intent", + "split": "test", + "languages": 51, + "per_lang": 100, + "n_opts": 20, + "seed": 13, + "device": "cpu", + "head_max_len": 192, + "max_len": 512, + "instructions": "What is the user asking for in `utterance`?", + "laya_version": "0.3.7", + "n_cases": 5100 + }, + "temperature_choice_11_plus": { + "committed_sweep": "0.2.0", + "unclamped_rerun": 0.10058280825614929, + "clamped_rerun": 0.5, + "note": "a 20-option question is bucket `choice:11+`; it is the only bucket the clamp moves, and every case in this sweep is a 20-option question" + }, + "macro": { + "committed": { + "accuracy": 0.2269, + "ece": 0.7331, + "languages": 51 + }, + "unclamped_rerun": { + "languages": 51, + "macro_accuracy": 0.2269, + "macro_ece": 0.7331, + "macro_f1": 0.2053, + "seconds": 1444.9 + }, + "clamped_rerun": { + "languages": 51, + "macro_accuracy": 0.2269, + "macro_ece": 0.5709, + "macro_f1": 0.2053, + "seconds": 1360.1 + } + }, + "per_language": { + "af": { + "n": 100, + "committed": { + "accuracy": 0.29, + "macro_f1": 0.2874, + "ece": 0.6873, + "mean_confidence": 0.9752, + "acc_at_50_coverage": 0.42 + }, + "unclamped_rerun": { + "accuracy": 0.29, + "macro_f1": 0.2874, + "ece": 0.6873, + "mean_confidence": 0.9752, + "acc_at_50_coverage": 0.42 + }, + "clamped_rerun": { + "accuracy": 0.29, + "macro_f1": 0.2874, + "ece": 0.5588, + "mean_confidence": 0.8488, + "acc_at_50_coverage": 0.42 + }, + "delta_ece": -0.1285, + "delta_mean_confidence": -0.1264 + }, + "am": { + "n": 100, + "committed": { + "accuracy": 0.12, + "macro_f1": 0.1122, + "ece": 0.8252, + "mean_confidence": 0.9452, + "acc_at_50_coverage": 0.14 + }, + "unclamped_rerun": { + "accuracy": 0.12, + "macro_f1": 0.1122, + "ece": 0.8253, + "mean_confidence": 0.9453, + "acc_at_50_coverage": 0.14 + }, + "clamped_rerun": { + "accuracy": 0.12, + "macro_f1": 0.1122, + "ece": 0.6804, + "mean_confidence": 0.7873, + "acc_at_50_coverage": 0.12 + }, + "delta_ece": -0.1448, + "delta_mean_confidence": -0.1579 + }, + "ar": { + "n": 100, + "committed": { + "accuracy": 0.11, + "macro_f1": 0.093, + "ece": 0.8001, + "mean_confidence": 0.8976, + "acc_at_50_coverage": 0.12 + }, + "unclamped_rerun": { + "accuracy": 0.11, + "macro_f1": 0.093, + "ece": 0.8001, + "mean_confidence": 0.8976, + "acc_at_50_coverage": 0.12 + }, + "clamped_rerun": { + "accuracy": 0.11, + "macro_f1": 0.093, + "ece": 0.5795, + "mean_confidence": 0.6833, + "acc_at_50_coverage": 0.12 + }, + "delta_ece": -0.2206, + "delta_mean_confidence": -0.2143 + }, + "az": { + "n": 100, + "committed": { + "accuracy": 0.1, + "macro_f1": 0.0629, + "ece": 0.8247, + "mean_confidence": 0.9247, + "acc_at_50_coverage": 0.14 + }, + "unclamped_rerun": { + "accuracy": 0.1, + "macro_f1": 0.0629, + "ece": 0.8247, + "mean_confidence": 0.9247, + "acc_at_50_coverage": 0.14 + }, + "clamped_rerun": { + "accuracy": 0.1, + "macro_f1": 0.0629, + "ece": 0.5983, + "mean_confidence": 0.6983, + "acc_at_50_coverage": 0.14 + }, + "delta_ece": -0.2264, + "delta_mean_confidence": -0.2264 + }, + "bn": { + "n": 100, + "committed": { + "accuracy": 0.08, + "macro_f1": 0.056, + "ece": 0.8646, + "mean_confidence": 0.9446, + "acc_at_50_coverage": 0.1 + }, + "unclamped_rerun": { + "accuracy": 0.08, + "macro_f1": 0.056, + "ece": 0.8646, + "mean_confidence": 0.9446, + "acc_at_50_coverage": 0.1 + }, + "clamped_rerun": { + "accuracy": 0.08, + "macro_f1": 0.056, + "ece": 0.6267, + "mean_confidence": 0.7067, + "acc_at_50_coverage": 0.12 + }, + "delta_ece": -0.2379, + "delta_mean_confidence": -0.2379 + }, + "cy": { + "n": 100, + "committed": { + "accuracy": 0.12, + "macro_f1": 0.0906, + "ece": 0.8409, + "mean_confidence": 0.9609, + "acc_at_50_coverage": 0.14 + }, + "unclamped_rerun": { + "accuracy": 0.12, + "macro_f1": 0.0906, + "ece": 0.8409, + "mean_confidence": 0.9609, + "acc_at_50_coverage": 0.14 + }, + "clamped_rerun": { + "accuracy": 0.12, + "macro_f1": 0.0906, + "ece": 0.6452, + "mean_confidence": 0.7652, + "acc_at_50_coverage": 0.14 + }, + "delta_ece": -0.1957, + "delta_mean_confidence": -0.1957 + }, + "da": { + "n": 100, + "committed": { + "accuracy": 0.35, + "macro_f1": 0.3082, + "ece": 0.6263, + "mean_confidence": 0.9763, + "acc_at_50_coverage": 0.52 + }, + "unclamped_rerun": { + "accuracy": 0.35, + "macro_f1": 0.3082, + "ece": 0.6263, + "mean_confidence": 0.9763, + "acc_at_50_coverage": 0.52 + }, + "clamped_rerun": { + "accuracy": 0.35, + "macro_f1": 0.3082, + "ece": 0.526, + "mean_confidence": 0.876, + "acc_at_50_coverage": 0.52 + }, + "delta_ece": -0.1003, + "delta_mean_confidence": -0.1003 + }, + "de": { + "n": 100, + "committed": { + "accuracy": 0.42, + "macro_f1": 0.387, + "ece": 0.5584, + "mean_confidence": 0.9784, + "acc_at_50_coverage": 0.66 + }, + "unclamped_rerun": { + "accuracy": 0.42, + "macro_f1": 0.387, + "ece": 0.5584, + "mean_confidence": 0.9784, + "acc_at_50_coverage": 0.66 + }, + "clamped_rerun": { + "accuracy": 0.42, + "macro_f1": 0.387, + "ece": 0.452, + "mean_confidence": 0.872, + "acc_at_50_coverage": 0.68 + }, + "delta_ece": -0.1064, + "delta_mean_confidence": -0.1064 + }, + "el": { + "n": 100, + "committed": { + "accuracy": 0.13, + "macro_f1": 0.1, + "ece": 0.8387, + "mean_confidence": 0.9687, + "acc_at_50_coverage": 0.16 + }, + "unclamped_rerun": { + "accuracy": 0.13, + "macro_f1": 0.1, + "ece": 0.8388, + "mean_confidence": 0.9688, + "acc_at_50_coverage": 0.16 + }, + "clamped_rerun": { + "accuracy": 0.13, + "macro_f1": 0.1, + "ece": 0.6789, + "mean_confidence": 0.8089, + "acc_at_50_coverage": 0.2 + }, + "delta_ece": -0.1598, + "delta_mean_confidence": -0.1598 + }, + "en": { + "n": 100, + "committed": { + "accuracy": 0.82, + "macro_f1": 0.7876, + "ece": 0.1789, + "mean_confidence": 0.9989, + "acc_at_50_coverage": 0.94 + }, + "unclamped_rerun": { + "accuracy": 0.82, + "macro_f1": 0.7876, + "ece": 0.1789, + "mean_confidence": 0.9989, + "acc_at_50_coverage": 0.94 + }, + "clamped_rerun": { + "accuracy": 0.82, + "macro_f1": 0.7876, + "ece": 0.1382, + "mean_confidence": 0.9582, + "acc_at_50_coverage": 0.98 + }, + "delta_ece": -0.0407, + "delta_mean_confidence": -0.0407 + }, + "es": { + "n": 100, + "committed": { + "accuracy": 0.51, + "macro_f1": 0.4767, + "ece": 0.4796, + "mean_confidence": 0.9896, + "acc_at_50_coverage": 0.6 + }, + "unclamped_rerun": { + "accuracy": 0.51, + "macro_f1": 0.4767, + "ece": 0.4796, + "mean_confidence": 0.9896, + "acc_at_50_coverage": 0.6 + }, + "clamped_rerun": { + "accuracy": 0.51, + "macro_f1": 0.4767, + "ece": 0.4225, + "mean_confidence": 0.9245, + "acc_at_50_coverage": 0.62 + }, + "delta_ece": -0.0571, + "delta_mean_confidence": -0.0651 + }, + "fa": { + "n": 100, + "committed": { + "accuracy": 0.14, + "macro_f1": 0.1197, + "ece": 0.8196, + "mean_confidence": 0.9414, + "acc_at_50_coverage": 0.16 + }, + "unclamped_rerun": { + "accuracy": 0.14, + "macro_f1": 0.1197, + "ece": 0.8196, + "mean_confidence": 0.9414, + "acc_at_50_coverage": 0.16 + }, + "clamped_rerun": { + "accuracy": 0.14, + "macro_f1": 0.1197, + "ece": 0.6331, + "mean_confidence": 0.7731, + "acc_at_50_coverage": 0.16 + }, + "delta_ece": -0.1865, + "delta_mean_confidence": -0.1683 + }, + "fi": { + "n": 100, + "committed": { + "accuracy": 0.13, + "macro_f1": 0.1213, + "ece": 0.8486, + "mean_confidence": 0.9786, + "acc_at_50_coverage": 0.16 + }, + "unclamped_rerun": { + "accuracy": 0.13, + "macro_f1": 0.1213, + "ece": 0.8486, + "mean_confidence": 0.9786, + "acc_at_50_coverage": 0.16 + }, + "clamped_rerun": { + "accuracy": 0.13, + "macro_f1": 0.1213, + "ece": 0.7189, + "mean_confidence": 0.8489, + "acc_at_50_coverage": 0.16 + }, + "delta_ece": -0.1297, + "delta_mean_confidence": -0.1297 + }, + "fr": { + "n": 100, + "committed": { + "accuracy": 0.59, + "macro_f1": 0.5822, + "ece": 0.3882, + "mean_confidence": 0.9782, + "acc_at_50_coverage": 0.78 + }, + "unclamped_rerun": { + "accuracy": 0.59, + "macro_f1": 0.5822, + "ece": 0.3882, + "mean_confidence": 0.9782, + "acc_at_50_coverage": 0.78 + }, + "clamped_rerun": { + "accuracy": 0.59, + "macro_f1": 0.5822, + "ece": 0.2995, + "mean_confidence": 0.8895, + "acc_at_50_coverage": 0.78 + }, + "delta_ece": -0.0887, + "delta_mean_confidence": -0.0887 + }, + "he": { + "n": 100, + "committed": { + "accuracy": 0.06, + "macro_f1": 0.039, + "ece": 0.9111, + "mean_confidence": 0.9643, + "acc_at_50_coverage": 0.04 + }, + "unclamped_rerun": { + "accuracy": 0.06, + "macro_f1": 0.039, + "ece": 0.9111, + "mean_confidence": 0.9643, + "acc_at_50_coverage": 0.04 + }, + "clamped_rerun": { + "accuracy": 0.06, + "macro_f1": 0.039, + "ece": 0.7379, + "mean_confidence": 0.7842, + "acc_at_50_coverage": 0.02 + }, + "delta_ece": -0.1732, + "delta_mean_confidence": -0.1801 + }, + "hi": { + "n": 100, + "committed": { + "accuracy": 0.1, + "macro_f1": 0.0758, + "ece": 0.8504, + "mean_confidence": 0.9408, + "acc_at_50_coverage": 0.1 + }, + "unclamped_rerun": { + "accuracy": 0.1, + "macro_f1": 0.0758, + "ece": 0.8504, + "mean_confidence": 0.9408, + "acc_at_50_coverage": 0.1 + }, + "clamped_rerun": { + "accuracy": 0.1, + "macro_f1": 0.0758, + "ece": 0.6416, + "mean_confidence": 0.7416, + "acc_at_50_coverage": 0.08 + }, + "delta_ece": -0.2088, + "delta_mean_confidence": -0.1992 + }, + "hu": { + "n": 100, + "committed": { + "accuracy": 0.09, + "macro_f1": 0.0813, + "ece": 0.8571, + "mean_confidence": 0.9471, + "acc_at_50_coverage": 0.06 + }, + "unclamped_rerun": { + "accuracy": 0.09, + "macro_f1": 0.0813, + "ece": 0.8571, + "mean_confidence": 0.9471, + "acc_at_50_coverage": 0.06 + }, + "clamped_rerun": { + "accuracy": 0.09, + "macro_f1": 0.0813, + "ece": 0.6783, + "mean_confidence": 0.7683, + "acc_at_50_coverage": 0.06 + }, + "delta_ece": -0.1788, + "delta_mean_confidence": -0.1788 + }, + "hy": { + "n": 100, + "committed": { + "accuracy": 0.05, + "macro_f1": 0.0355, + "ece": 0.8348, + "mean_confidence": 0.8848, + "acc_at_50_coverage": 0.08 + }, + "unclamped_rerun": { + "accuracy": 0.05, + "macro_f1": 0.0355, + "ece": 0.8348, + "mean_confidence": 0.8848, + "acc_at_50_coverage": 0.08 + }, + "clamped_rerun": { + "accuracy": 0.05, + "macro_f1": 0.0355, + "ece": 0.5709, + "mean_confidence": 0.6209, + "acc_at_50_coverage": 0.04 + }, + "delta_ece": -0.2639, + "delta_mean_confidence": -0.2639 + }, + "id": { + "n": 100, + "committed": { + "accuracy": 0.36, + "macro_f1": 0.3283, + "ece": 0.6133, + "mean_confidence": 0.9619, + "acc_at_50_coverage": 0.48 + }, + "unclamped_rerun": { + "accuracy": 0.36, + "macro_f1": 0.3283, + "ece": 0.6133, + "mean_confidence": 0.9619, + "acc_at_50_coverage": 0.48 + }, + "clamped_rerun": { + "accuracy": 0.36, + "macro_f1": 0.3283, + "ece": 0.487, + "mean_confidence": 0.8304, + "acc_at_50_coverage": 0.48 + }, + "delta_ece": -0.1263, + "delta_mean_confidence": -0.1315 + }, + "is": { + "n": 100, + "committed": { + "accuracy": 0.11, + "macro_f1": 0.0893, + "ece": 0.8351, + "mean_confidence": 0.9451, + "acc_at_50_coverage": 0.16 + }, + "unclamped_rerun": { + "accuracy": 0.11, + "macro_f1": 0.0893, + "ece": 0.8351, + "mean_confidence": 0.9451, + "acc_at_50_coverage": 0.16 + }, + "clamped_rerun": { + "accuracy": 0.11, + "macro_f1": 0.0893, + "ece": 0.6641, + "mean_confidence": 0.7741, + "acc_at_50_coverage": 0.16 + }, + "delta_ece": -0.171, + "delta_mean_confidence": -0.171 + }, + "it": { + "n": 100, + "committed": { + "accuracy": 0.34, + "macro_f1": 0.3101, + "ece": 0.6474, + "mean_confidence": 0.9735, + "acc_at_50_coverage": 0.52 + }, + "unclamped_rerun": { + "accuracy": 0.34, + "macro_f1": 0.3101, + "ece": 0.6474, + "mean_confidence": 0.9735, + "acc_at_50_coverage": 0.52 + }, + "clamped_rerun": { + "accuracy": 0.34, + "macro_f1": 0.3101, + "ece": 0.5156, + "mean_confidence": 0.8534, + "acc_at_50_coverage": 0.52 + }, + "delta_ece": -0.1318, + "delta_mean_confidence": -0.1201 + }, + "ja": { + "n": 100, + "committed": { + "accuracy": 0.53, + "macro_f1": 0.4942, + "ece": 0.4599, + "mean_confidence": 0.9801, + "acc_at_50_coverage": 0.76 + }, + "unclamped_rerun": { + "accuracy": 0.53, + "macro_f1": 0.4942, + "ece": 0.4599, + "mean_confidence": 0.9801, + "acc_at_50_coverage": 0.76 + }, + "clamped_rerun": { + "accuracy": 0.53, + "macro_f1": 0.4942, + "ece": 0.3657, + "mean_confidence": 0.8957, + "acc_at_50_coverage": 0.76 + }, + "delta_ece": -0.0942, + "delta_mean_confidence": -0.0844 + }, + "jv": { + "n": 100, + "committed": { + "accuracy": 0.16, + "macro_f1": 0.1537, + "ece": 0.8035, + "mean_confidence": 0.9635, + "acc_at_50_coverage": 0.14 + }, + "unclamped_rerun": { + "accuracy": 0.16, + "macro_f1": 0.1537, + "ece": 0.8035, + "mean_confidence": 0.9635, + "acc_at_50_coverage": 0.14 + }, + "clamped_rerun": { + "accuracy": 0.16, + "macro_f1": 0.1537, + "ece": 0.67, + "mean_confidence": 0.83, + "acc_at_50_coverage": 0.16 + }, + "delta_ece": -0.1335, + "delta_mean_confidence": -0.1335 + }, + "ka": { + "n": 100, + "committed": { + "accuracy": 0.09, + "macro_f1": 0.0742, + "ece": 0.8448, + "mean_confidence": 0.9348, + "acc_at_50_coverage": 0.1 + }, + "unclamped_rerun": { + "accuracy": 0.09, + "macro_f1": 0.0742, + "ece": 0.8448, + "mean_confidence": 0.9348, + "acc_at_50_coverage": 0.1 + }, + "clamped_rerun": { + "accuracy": 0.09, + "macro_f1": 0.0742, + "ece": 0.6102, + "mean_confidence": 0.7002, + "acc_at_50_coverage": 0.1 + }, + "delta_ece": -0.2346, + "delta_mean_confidence": -0.2346 + }, + "km": { + "n": 100, + "committed": { + "accuracy": 0.0, + "macro_f1": 0.0, + "ece": 0.9515, + "mean_confidence": 0.9515, + "acc_at_50_coverage": 0.0 + }, + "unclamped_rerun": { + "accuracy": 0.0, + "macro_f1": 0.0, + "ece": 0.9515, + "mean_confidence": 0.9515, + "acc_at_50_coverage": 0.0 + }, + "clamped_rerun": { + "accuracy": 0.0, + "macro_f1": 0.0, + "ece": 0.7051, + "mean_confidence": 0.7051, + "acc_at_50_coverage": 0.0 + }, + "delta_ece": -0.2464, + "delta_mean_confidence": -0.2464 + }, + "kn": { + "n": 100, + "committed": { + "accuracy": 0.11, + "macro_f1": 0.0913, + "ece": 0.8416, + "mean_confidence": 0.9421, + "acc_at_50_coverage": 0.06 + }, + "unclamped_rerun": { + "accuracy": 0.11, + "macro_f1": 0.0913, + "ece": 0.8416, + "mean_confidence": 0.9421, + "acc_at_50_coverage": 0.06 + }, + "clamped_rerun": { + "accuracy": 0.11, + "macro_f1": 0.0913, + "ece": 0.5903, + "mean_confidence": 0.695, + "acc_at_50_coverage": 0.08 + }, + "delta_ece": -0.2513, + "delta_mean_confidence": -0.2471 + }, + "ko": { + "n": 100, + "committed": { + "accuracy": 0.11, + "macro_f1": 0.0943, + "ece": 0.8498, + "mean_confidence": 0.9598, + "acc_at_50_coverage": 0.14 + }, + "unclamped_rerun": { + "accuracy": 0.11, + "macro_f1": 0.0943, + "ece": 0.8498, + "mean_confidence": 0.9598, + "acc_at_50_coverage": 0.14 + }, + "clamped_rerun": { + "accuracy": 0.11, + "macro_f1": 0.0943, + "ece": 0.6804, + "mean_confidence": 0.7904, + "acc_at_50_coverage": 0.14 + }, + "delta_ece": -0.1694, + "delta_mean_confidence": -0.1694 + }, + "lv": { + "n": 100, + "committed": { + "accuracy": 0.1, + "macro_f1": 0.0786, + "ece": 0.8466, + "mean_confidence": 0.9466, + "acc_at_50_coverage": 0.1 + }, + "unclamped_rerun": { + "accuracy": 0.1, + "macro_f1": 0.0786, + "ece": 0.8466, + "mean_confidence": 0.9466, + "acc_at_50_coverage": 0.1 + }, + "clamped_rerun": { + "accuracy": 0.1, + "macro_f1": 0.0786, + "ece": 0.6659, + "mean_confidence": 0.7659, + "acc_at_50_coverage": 0.1 + }, + "delta_ece": -0.1807, + "delta_mean_confidence": -0.1807 + }, + "ml": { + "n": 100, + "committed": { + "accuracy": 0.07, + "macro_f1": 0.0503, + "ece": 0.8572, + "mean_confidence": 0.9272, + "acc_at_50_coverage": 0.1 + }, + "unclamped_rerun": { + "accuracy": 0.07, + "macro_f1": 0.0503, + "ece": 0.8572, + "mean_confidence": 0.9272, + "acc_at_50_coverage": 0.1 + }, + "clamped_rerun": { + "accuracy": 0.07, + "macro_f1": 0.0503, + "ece": 0.601, + "mean_confidence": 0.671, + "acc_at_50_coverage": 0.1 + }, + "delta_ece": -0.2562, + "delta_mean_confidence": -0.2562 + }, + "mn": { + "n": 100, + "committed": { + "accuracy": 0.13, + "macro_f1": 0.1213, + "ece": 0.837, + "mean_confidence": 0.9591, + "acc_at_50_coverage": 0.14 + }, + "unclamped_rerun": { + "accuracy": 0.13, + "macro_f1": 0.1213, + "ece": 0.837, + "mean_confidence": 0.9591, + "acc_at_50_coverage": 0.14 + }, + "clamped_rerun": { + "accuracy": 0.13, + "macro_f1": 0.1213, + "ece": 0.6704, + "mean_confidence": 0.7977, + "acc_at_50_coverage": 0.14 + }, + "delta_ece": -0.1666, + "delta_mean_confidence": -0.1614 + }, + "ms": { + "n": 100, + "committed": { + "accuracy": 0.27, + "macro_f1": 0.2372, + "ece": 0.6877, + "mean_confidence": 0.9577, + "acc_at_50_coverage": 0.42 + }, + "unclamped_rerun": { + "accuracy": 0.27, + "macro_f1": 0.2372, + "ece": 0.6877, + "mean_confidence": 0.9577, + "acc_at_50_coverage": 0.42 + }, + "clamped_rerun": { + "accuracy": 0.27, + "macro_f1": 0.2372, + "ece": 0.5064, + "mean_confidence": 0.7764, + "acc_at_50_coverage": 0.4 + }, + "delta_ece": -0.1813, + "delta_mean_confidence": -0.1813 + }, + "my": { + "n": 100, + "committed": { + "accuracy": 0.06, + "macro_f1": 0.0455, + "ece": 0.8605, + "mean_confidence": 0.9205, + "acc_at_50_coverage": 0.08 + }, + "unclamped_rerun": { + "accuracy": 0.06, + "macro_f1": 0.0455, + "ece": 0.8605, + "mean_confidence": 0.9205, + "acc_at_50_coverage": 0.08 + }, + "clamped_rerun": { + "accuracy": 0.06, + "macro_f1": 0.0455, + "ece": 0.5797, + "mean_confidence": 0.6397, + "acc_at_50_coverage": 0.08 + }, + "delta_ece": -0.2808, + "delta_mean_confidence": -0.2808 + }, + "nb": { + "n": 100, + "committed": { + "accuracy": 0.33, + "macro_f1": 0.3025, + "ece": 0.6484, + "mean_confidence": 0.9692, + "acc_at_50_coverage": 0.5 + }, + "unclamped_rerun": { + "accuracy": 0.33, + "macro_f1": 0.3025, + "ece": 0.6484, + "mean_confidence": 0.9692, + "acc_at_50_coverage": 0.5 + }, + "clamped_rerun": { + "accuracy": 0.33, + "macro_f1": 0.3025, + "ece": 0.5179, + "mean_confidence": 0.8479, + "acc_at_50_coverage": 0.52 + }, + "delta_ece": -0.1305, + "delta_mean_confidence": -0.1213 + }, + "nl": { + "n": 100, + "committed": { + "accuracy": 0.39, + "macro_f1": 0.3968, + "ece": 0.5914, + "mean_confidence": 0.9745, + "acc_at_50_coverage": 0.54 + }, + "unclamped_rerun": { + "accuracy": 0.39, + "macro_f1": 0.3968, + "ece": 0.5914, + "mean_confidence": 0.9745, + "acc_at_50_coverage": 0.54 + }, + "clamped_rerun": { + "accuracy": 0.39, + "macro_f1": 0.3968, + "ece": 0.5025, + "mean_confidence": 0.8831, + "acc_at_50_coverage": 0.56 + }, + "delta_ece": -0.0889, + "delta_mean_confidence": -0.0914 + }, + "pl": { + "n": 100, + "committed": { + "accuracy": 0.24, + "macro_f1": 0.2033, + "ece": 0.7129, + "mean_confidence": 0.9529, + "acc_at_50_coverage": 0.34 + }, + "unclamped_rerun": { + "accuracy": 0.24, + "macro_f1": 0.2033, + "ece": 0.7129, + "mean_confidence": 0.9529, + "acc_at_50_coverage": 0.34 + }, + "clamped_rerun": { + "accuracy": 0.24, + "macro_f1": 0.2033, + "ece": 0.5695, + "mean_confidence": 0.7951, + "acc_at_50_coverage": 0.36 + }, + "delta_ece": -0.1434, + "delta_mean_confidence": -0.1578 + }, + "pt": { + "n": 100, + "committed": { + "accuracy": 0.47, + "macro_f1": 0.4213, + "ece": 0.5116, + "mean_confidence": 0.9718, + "acc_at_50_coverage": 0.66 + }, + "unclamped_rerun": { + "accuracy": 0.47, + "macro_f1": 0.4213, + "ece": 0.5116, + "mean_confidence": 0.9718, + "acc_at_50_coverage": 0.66 + }, + "clamped_rerun": { + "accuracy": 0.47, + "macro_f1": 0.4213, + "ece": 0.4317, + "mean_confidence": 0.8894, + "acc_at_50_coverage": 0.66 + }, + "delta_ece": -0.0799, + "delta_mean_confidence": -0.0824 + }, + "ro": { + "n": 100, + "committed": { + "accuracy": 0.33, + "macro_f1": 0.2976, + "ece": 0.6577, + "mean_confidence": 0.9824, + "acc_at_50_coverage": 0.48 + }, + "unclamped_rerun": { + "accuracy": 0.33, + "macro_f1": 0.2976, + "ece": 0.6576, + "mean_confidence": 0.9824, + "acc_at_50_coverage": 0.48 + }, + "clamped_rerun": { + "accuracy": 0.33, + "macro_f1": 0.2976, + "ece": 0.5499, + "mean_confidence": 0.8668, + "acc_at_50_coverage": 0.48 + }, + "delta_ece": -0.1078, + "delta_mean_confidence": -0.1156 + }, + "ru": { + "n": 100, + "committed": { + "accuracy": 0.31, + "macro_f1": 0.3084, + "ece": 0.668, + "mean_confidence": 0.978, + "acc_at_50_coverage": 0.42 + }, + "unclamped_rerun": { + "accuracy": 0.31, + "macro_f1": 0.3084, + "ece": 0.668, + "mean_confidence": 0.978, + "acc_at_50_coverage": 0.42 + }, + "clamped_rerun": { + "accuracy": 0.31, + "macro_f1": 0.3084, + "ece": 0.5946, + "mean_confidence": 0.9046, + "acc_at_50_coverage": 0.42 + }, + "delta_ece": -0.0734, + "delta_mean_confidence": -0.0734 + }, + "sl": { + "n": 100, + "committed": { + "accuracy": 0.2, + "macro_f1": 0.1679, + "ece": 0.7561, + "mean_confidence": 0.9561, + "acc_at_50_coverage": 0.26 + }, + "unclamped_rerun": { + "accuracy": 0.2, + "macro_f1": 0.1679, + "ece": 0.7561, + "mean_confidence": 0.9561, + "acc_at_50_coverage": 0.26 + }, + "clamped_rerun": { + "accuracy": 0.2, + "macro_f1": 0.1679, + "ece": 0.5994, + "mean_confidence": 0.7994, + "acc_at_50_coverage": 0.26 + }, + "delta_ece": -0.1567, + "delta_mean_confidence": -0.1567 + }, + "sq": { + "n": 100, + "committed": { + "accuracy": 0.21, + "macro_f1": 0.1667, + "ece": 0.7549, + "mean_confidence": 0.9649, + "acc_at_50_coverage": 0.28 + }, + "unclamped_rerun": { + "accuracy": 0.21, + "macro_f1": 0.1667, + "ece": 0.7549, + "mean_confidence": 0.9649, + "acc_at_50_coverage": 0.28 + }, + "clamped_rerun": { + "accuracy": 0.21, + "macro_f1": 0.1667, + "ece": 0.6031, + "mean_confidence": 0.8131, + "acc_at_50_coverage": 0.28 + }, + "delta_ece": -0.1518, + "delta_mean_confidence": -0.1518 + }, + "sv": { + "n": 100, + "committed": { + "accuracy": 0.38, + "macro_f1": 0.3809, + "ece": 0.5956, + "mean_confidence": 0.9706, + "acc_at_50_coverage": 0.5 + }, + "unclamped_rerun": { + "accuracy": 0.38, + "macro_f1": 0.3809, + "ece": 0.5956, + "mean_confidence": 0.9706, + "acc_at_50_coverage": 0.5 + }, + "clamped_rerun": { + "accuracy": 0.38, + "macro_f1": 0.3809, + "ece": 0.4874, + "mean_confidence": 0.8501, + "acc_at_50_coverage": 0.5 + }, + "delta_ece": -0.1082, + "delta_mean_confidence": -0.1205 + }, + "sw": { + "n": 100, + "committed": { + "accuracy": 0.13, + "macro_f1": 0.1149, + "ece": 0.8282, + "mean_confidence": 0.9562, + "acc_at_50_coverage": 0.14 + }, + "unclamped_rerun": { + "accuracy": 0.13, + "macro_f1": 0.1149, + "ece": 0.8282, + "mean_confidence": 0.9562, + "acc_at_50_coverage": 0.14 + }, + "clamped_rerun": { + "accuracy": 0.13, + "macro_f1": 0.1149, + "ece": 0.6318, + "mean_confidence": 0.7511, + "acc_at_50_coverage": 0.14 + }, + "delta_ece": -0.1964, + "delta_mean_confidence": -0.2051 + }, + "ta": { + "n": 100, + "committed": { + "accuracy": 0.12, + "macro_f1": 0.0971, + "ece": 0.8224, + "mean_confidence": 0.9424, + "acc_at_50_coverage": 0.12 + }, + "unclamped_rerun": { + "accuracy": 0.12, + "macro_f1": 0.0971, + "ece": 0.8224, + "mean_confidence": 0.9424, + "acc_at_50_coverage": 0.12 + }, + "clamped_rerun": { + "accuracy": 0.12, + "macro_f1": 0.0971, + "ece": 0.5541, + "mean_confidence": 0.6741, + "acc_at_50_coverage": 0.12 + }, + "delta_ece": -0.2683, + "delta_mean_confidence": -0.2683 + }, + "te": { + "n": 100, + "committed": { + "accuracy": 0.09, + "macro_f1": 0.0722, + "ece": 0.8575, + "mean_confidence": 0.9475, + "acc_at_50_coverage": 0.06 + }, + "unclamped_rerun": { + "accuracy": 0.09, + "macro_f1": 0.0722, + "ece": 0.8575, + "mean_confidence": 0.9475, + "acc_at_50_coverage": 0.06 + }, + "clamped_rerun": { + "accuracy": 0.09, + "macro_f1": 0.0722, + "ece": 0.6689, + "mean_confidence": 0.7568, + "acc_at_50_coverage": 0.06 + }, + "delta_ece": -0.1886, + "delta_mean_confidence": -0.1907 + }, + "th": { + "n": 100, + "committed": { + "accuracy": 0.08, + "macro_f1": 0.048, + "ece": 0.8814, + "mean_confidence": 0.9614, + "acc_at_50_coverage": 0.1 + }, + "unclamped_rerun": { + "accuracy": 0.08, + "macro_f1": 0.048, + "ece": 0.8814, + "mean_confidence": 0.9614, + "acc_at_50_coverage": 0.1 + }, + "clamped_rerun": { + "accuracy": 0.08, + "macro_f1": 0.048, + "ece": 0.7179, + "mean_confidence": 0.7979, + "acc_at_50_coverage": 0.1 + }, + "delta_ece": -0.1635, + "delta_mean_confidence": -0.1635 + }, + "tl": { + "n": 100, + "committed": { + "accuracy": 0.29, + "macro_f1": 0.249, + "ece": 0.6756, + "mean_confidence": 0.9656, + "acc_at_50_coverage": 0.46 + }, + "unclamped_rerun": { + "accuracy": 0.29, + "macro_f1": 0.249, + "ece": 0.6756, + "mean_confidence": 0.9656, + "acc_at_50_coverage": 0.46 + }, + "clamped_rerun": { + "accuracy": 0.29, + "macro_f1": 0.249, + "ece": 0.5121, + "mean_confidence": 0.8021, + "acc_at_50_coverage": 0.46 + }, + "delta_ece": -0.1635, + "delta_mean_confidence": -0.1635 + }, + "tr": { + "n": 100, + "committed": { + "accuracy": 0.14, + "macro_f1": 0.0967, + "ece": 0.788, + "mean_confidence": 0.928, + "acc_at_50_coverage": 0.2 + }, + "unclamped_rerun": { + "accuracy": 0.14, + "macro_f1": 0.0967, + "ece": 0.788, + "mean_confidence": 0.928, + "acc_at_50_coverage": 0.2 + }, + "clamped_rerun": { + "accuracy": 0.14, + "macro_f1": 0.0967, + "ece": 0.5874, + "mean_confidence": 0.7274, + "acc_at_50_coverage": 0.2 + }, + "delta_ece": -0.2006, + "delta_mean_confidence": -0.2006 + }, + "ur": { + "n": 100, + "committed": { + "accuracy": 0.07, + "macro_f1": 0.0645, + "ece": 0.8826, + "mean_confidence": 0.9526, + "acc_at_50_coverage": 0.08 + }, + "unclamped_rerun": { + "accuracy": 0.07, + "macro_f1": 0.0645, + "ece": 0.8826, + "mean_confidence": 0.9526, + "acc_at_50_coverage": 0.08 + }, + "clamped_rerun": { + "accuracy": 0.07, + "macro_f1": 0.0645, + "ece": 0.6428, + "mean_confidence": 0.7128, + "acc_at_50_coverage": 0.08 + }, + "delta_ece": -0.2398, + "delta_mean_confidence": -0.2398 + }, + "vi": { + "n": 100, + "committed": { + "accuracy": 0.06, + "macro_f1": 0.0498, + "ece": 0.8914, + "mean_confidence": 0.9514, + "acc_at_50_coverage": 0.06 + }, + "unclamped_rerun": { + "accuracy": 0.06, + "macro_f1": 0.0498, + "ece": 0.8914, + "mean_confidence": 0.9514, + "acc_at_50_coverage": 0.06 + }, + "clamped_rerun": { + "accuracy": 0.06, + "macro_f1": 0.0498, + "ece": 0.69, + "mean_confidence": 0.75, + "acc_at_50_coverage": 0.06 + }, + "delta_ece": -0.2014, + "delta_mean_confidence": -0.2014 + }, + "zh-CN": { + "n": 100, + "committed": { + "accuracy": 0.62, + "macro_f1": 0.6175, + "ece": 0.376, + "mean_confidence": 0.9885, + "acc_at_50_coverage": 0.86 + }, + "unclamped_rerun": { + "accuracy": 0.62, + "macro_f1": 0.6175, + "ece": 0.376, + "mean_confidence": 0.9885, + "acc_at_50_coverage": 0.86 + }, + "clamped_rerun": { + "accuracy": 0.62, + "macro_f1": 0.6175, + "ece": 0.3198, + "mean_confidence": 0.9244, + "acc_at_50_coverage": 0.84 + }, + "delta_ece": -0.0562, + "delta_mean_confidence": -0.0641 + }, + "zh-TW": { + "n": 100, + "committed": { + "accuracy": 0.46, + "macro_f1": 0.4294, + "ece": 0.5198, + "mean_confidence": 0.9798, + "acc_at_50_coverage": 0.74 + }, + "unclamped_rerun": { + "accuracy": 0.46, + "macro_f1": 0.4294, + "ece": 0.5198, + "mean_confidence": 0.9798, + "acc_at_50_coverage": 0.74 + }, + "clamped_rerun": { + "accuracy": 0.46, + "macro_f1": 0.4294, + "ece": 0.4344, + "mean_confidence": 0.8944, + "acc_at_50_coverage": 0.72 + }, + "delta_ece": -0.0854, + "delta_mean_confidence": -0.0854 + } + }, + "spread": { + "largest_ece_change": [ + { + "lang": "my", + "delta_ece": -0.2808 + }, + { + "lang": "ta", + "delta_ece": -0.2683 + }, + { + "lang": "hy", + "delta_ece": -0.2639 + }, + { + "lang": "ml", + "delta_ece": -0.2562 + }, + { + "lang": "kn", + "delta_ece": -0.2513 + } + ], + "smallest_ece_change": [ + { + "lang": "pt", + "delta_ece": -0.0799 + }, + { + "lang": "ru", + "delta_ece": -0.0734 + }, + { + "lang": "es", + "delta_ece": -0.0571 + }, + { + "lang": "zh-CN", + "delta_ece": -0.0562 + }, + { + "lang": "en", + "delta_ece": -0.0407 + } + ] + }, + "fields_the_harness_does_not_compute": { + "committed_only": [ + "brier", + "nll", + "seconds", + "dropped" + ], + "note": "the committed file stores these and research/eval/laya_eval.py does not, so this re-run covers the six comparable columns rather than replacing the file" + } +}