Diagnostics
Finding where a model is weak, or where it matters to you
Accuracy and calibration summarize a model over many questions. Diagnostics look for the places where it is weak, or for how it does on the part of the benchmark that matters for your use. We support three ways of analyzing a model's answers: slicing the questions by what they are about, looking at the worst case, and checking whether the model answers from first principles.
Slicing by what a question is about
Every question carries rich metadata: its element, module and setting, and its type, domain, perspective and difficulty. Any metric can be recomputed on any subset, so results can be cut in whatever way suits the question at hand. Someone deploying a model for medical decisions can score only the medical domain. Someone studying games can score only the strategic elements.
Slicing answers a question you already have. It finds only the weaknesses you thought to look for.
Worst-case analysis
One way to find weaknesses you did not think to look for is to look at the worst case in place of the average. This is what our robustness measures do. For each element, they report the lowest normalized accuracy along one axis:
- Domain
- The lowest score across the element's domains. A model that handles a concept in finance but not in medicine scores badly here.
- Type
- The lowest across the forms the question takes.
- Perspective
- The lowest across points of view.
Read it against the element's average. Where the two are close, the model's grasp of the element does not depend on how the question is dressed. Where the worst case sits far below the average, something that should not matter does, and the axis tells you which slice to open.
Answering from first principles
Another weakness is a model that picks the right option without working the answer out. Two question formats check for this. With free text there are no options to choose from, so the model has to produce the answer itself. "No other option is correct" (NOTA) asks the same thing while staying with multiple choice, by sometimes taking the right answer off the list.
Compare the model's score in either format with its plain multiple-choice score. A model that is working the answer out loses little. One that picks the most plausible option on offer loses a lot.