Metrics
What is computed from a model's answers
Most questions have one correct answer. The model picks an option and is right or wrong. Some elements, however, test whether a model is consistent instead, such as the certainty effect or the endowment effect. No single option is correct on its own. The model answers two related questions, and the pair is scored: a pair is consistent if a rational agent could have given both answers. Each element page says which kind it is.
A model's overall score on an element is the average over that element's questions.
Accuracy
These use only the answer the model chose.
- Exact match
- The fraction of questions answered correctly.
- Normalized
- One point for a correct answer, minus 1 / (options − 1) for a wrong one, a correction for guessing (Budescu and Bar-Hillel, 1993). Random guessing scores zero on every element, whatever the number of options, so elements can be compared.
Calibration
These use the probability the model gives to each option, so they need a model that reports them. They show whether the model knows when it is unsure.
- Expected calibration error
- The gap between how confident the model is in its top answer and how often that answer is right, averaged over 10 confidence bins (Naeini et al., 2015; Guo et al., 2017). A well-calibrated model that says 70% is right about 70% of the time. We follow the usage in HELM.
- Brier score
- The squared distance between the model's probabilities and the correct answer, averaged over the options (Brier, 1950). Confident wrong answers cost the most.
- Expected probability assignment
- The average probability given to the correct option.
Which probabilities are used
A model's output is a distribution over every token it could produce, and some of that probability can land on text that is not one of the options. The calibration measures are computed on the part of the distribution that falls on the valid options, reshaped into a distribution in one of two ways:
- Conditioning
- Keep only the probabilities on the valid options and renormalize them. If the model put very little weight on any valid option, this can make it look more confident than it was.
- Mixing
- Blend the probabilities on the valid options with a uniform distribution, in proportion to how much probability fell outside them. A model that mostly answered outside the options ends up close to uniform, that is, unsure.
Neither changes which valid option has the highest probability, so accuracy is unaffected. The share of answers that were not a valid option is reported separately.