Development Phase
from 01 Oct
Build and validate systems on the dev set, and the given baseline.
Given a question written in Sinhala and its answer options, predict the correct option.
The Sinhala MMLU benchmark contains multiple-choice questions written natively in Sinhala across a range of subjects. Each question carries a difficulty label - Easy, Medium or Hard and systems must choose the single correct option.
Use any approach that respects the rules: prompting, fine-tuning, ensembles - as long as the entire inference pipeline totals 8B parameters or less and runs from submitted code with no human in the loop.
The test set is licensed for use in this shared task only. Redistributing test questions, or publishing them online, is strictly prohibited.
from 01 Oct
Build and validate systems on the dev set, and the given baseline.
12 – 22 Oct
Submit predictions for the hidden test set, and participate in the live Codabench leaderboard.
23 Oct – 02 Nov
Organisers reproduce final runs from submitted code. Final rankings are subject to change based on the reproducibility.