LOGPROBS output type) and acc is the fraction of questions where the gold option ranks highest. The two hyperparameters that matter most for matching published numbers are the 5-shot prompt and log-prob (rather than generative) scoring.
Evaluation configuration
Reproduce
Results
Qwen/Qwen3-0.6B-Base, 5-shot acc:
Mill
53.78% ± 1.53
measured
measured
Qwen3 report
52.81%
published baseline
published baseline
Difference
+0.97
within 1σ
within 1σ
Mill reproduces the Qwen3 Technical Report MMLU score for
Qwen/Qwen3-0.6B-Base (52.81%) within one standard error.Per-model results
Other models measured with Mill (5-shot
acc). Published MMLU numbers for these vary by eval protocol across sources, so they are recorded as Mill measurements pending a matched, like-for-like comparison:Per-subtask breakdown (57 tasks)
Per-subtask breakdown (57 tasks)