Skip to main content
MMLU is run as a 5-shot multiple-choice task. Each answer option is scored by its log-probability (LOGPROBS output type) and acc is the fraction of questions where the gold option ranks highest. The two hyperparameters that matter most for matching published numbers are the 5-shot prompt and log-prob (rather than generative) scoring.

Evaluation configuration

Reproduce

Results

Qwen/Qwen3-0.6B-Base, 5-shot acc:

Mill

53.78% ± 1.53
measured

Qwen3 report

52.81%
published baseline

Difference

+0.97
within 1σ
Mill reproduces the Qwen3 Technical Report MMLU score for Qwen/Qwen3-0.6B-Base (52.81%) within one standard error.

Per-model results

Other models measured with Mill (5-shot acc). Published MMLU numbers for these vary by eval protocol across sources, so they are recorded as Mill measurements pending a matched, like-for-like comparison: