STSTEER

Scoring

Scoring turns the answers a model gives to a set of questions into scores for each element, and presents those scores for the part of the benchmark you care about. It is described in the following articles:

ArticleDescription
Question formatsThe ways a question can be put to a model: multiple choice and free text, and the variants of multiple choice in which the model reasons first, with the options shown or hidden.
MetricsWhat counts as a correct answer, and what is computed from a model's answers: exact-match and normalized accuracy, and the calibration measures.
DiagnosticsFinding where a model is weak: slicing scores by what a question is about, worst-case analysis across domains, types and perspectives, and checking whether a model answers from first principles.

For the questions themselves, see Datasets and the Reference.