Same Benchmark Name, Different Results: Why Two MMLU Scores Can't Always Be Compared
Summary
Two MMLU benchmark scores from the same model family are deemed incomparable despite their similar numbers, exposing a critical flaw in AI evaluation: sharing a benchmark name means nothing if the testing conditions differ, as variables like prompt format, grader type, and dataset splits can shift accuracy by several percentage points and even reverse model rankings.
Key Points
- Two MMLU accuracy scores, 0.781 and 0.79, from two builds of the same model family are being examined, and despite sharing the same benchmark name, the verifier returns 'incomparable' because the evaluation frames differ in runner, grader, and dataset split.
- The MMLU benchmark name alone fixes nothing beyond a label, as published research shows that variables like dataset split, implementation, prompt format, grader type, and network access can shift accuracy by several percentage points and even reverse model rankings.
- Comparability is a property of a shared measurement reference, not of a shared number or name, and under the APL AI-Eval profile, two well-formed claims pointing to different frame hashes cannot be subtracted without an applicable bridge that meets strict aspect, scope, and procedure constraints.