Scoring
Scoring turns the answers a model gives to a set of questions into scores for each element, and presents those scores for the part of the benchmark you care about. It is described in the following articles:
| Article | Description |
|---|---|
| Question formats | The ways a question can be put to a model: multiple choice and free text, and the variants of multiple choice in which the model reasons first, with the options shown or hidden. |
| Metrics | What counts as a correct answer, and what is computed from a model's answers: exact-match and normalized accuracy, and the calibration measures. |
| Diagnostics | Finding where a model is weak: slicing scores by what a question is about, worst-case analysis across domains, types and perspectives, and checking whether a model answers from first principles. |
For the questions themselves, see Datasets and the Reference.