Home/evaluation

How We Evaluate

Accuracy at each difficulty level, and their macro average.

easyAccEaccuracy on Easy
mediumAccMaccuracy on Medium
hardAccHaccuracy on Hard
primaryMacro(E + M + H) / 3

Leaderboards

  • Codabench leaderboard - live during the test phase, computed on the hidden test set.
  • Final ranking - Determined based on the reproducibility of the submitted code.
  • Teams are ranked by macro-average accuracy. A baseline is shown for reference.

Verification

After the test phase closes, organisers run each team's code on their own hardware. A run is verified if the reproduced predictions match the submitted scores within a small tolerance defined in the full rules. Unverified runs are listed separately and are not ranked.