Question formats
The same question can be put to a model in several ways
A format fixes what the model is shown, in what order, and what it gives back. The question and its correct answer stay the same; only the presentation changes, so a difference in score between two formats is a difference in how the model was asked.
There are two baseline formats and three variants of multiple choice. Each diagram reads left to right.
Baseline
The two plain ways of asking: with options to choose from, and without. Scores in the other formats are read against these.
Multiple choice
The model sees the question and its options and returns one option straight away, without reasoning first (no chain of thought, or CoT). The probabilities it gives to the options are what the calibration metrics use.
Free text
The model gets no options and writes out its answer, which is closest to how models are used in practice. The final answer is extracted from the response and compared with the correct value.
Variants of multiple choice
Unlike free text, multiple choice can be posed in more than one way. We use the following three. The first two have the model reason before it answers, known as chain of thought (CoT).
Options shown before CoT
The model writes out its reasoning before it chooses, and sees the options while it reasons.
"No other option is correct" (NOTA)
One option is replaced by this sentence, which we abbreviate as NOTA, as in "none of the above". In one question out of every n, where n is the number of options, it replaces the correct option; otherwise it replaces a wrong one. A model that always picks the sentence does no better than guessing.
- 96
- 140
- 110
- No other option is correct
- 96
- No other option is correct
- 110
- 120