Tutorial
What this web interface is for
This site does two jobs. It lets you see what the benchmarks test, down to individual questions, and it tells you how to get questions and score a model on them.
Exploring the benchmarks
Each benchmark is a taxonomy: settings contain modules, and modules contain elements. Start from STEER or STEER-ME, whose pages show the settings as colored cards, and click through to a setting, then a module, then an element. The sidebar keeps the whole taxonomy in view while you browse, and the search box in the header finds an element by name.
The element page is where the detail is. Take Law of Demand. At the top is one of its question templates, with the fields that get filled in highlighted. Generate a question fills the template with fresh values and shows the options, and Show answer marks the correct one. Customize lets you fix the type, domain or perspective of the question that is generated. Below that, the page describes what the questions ask and how they vary, and "At a glance" gives how the element is scored and how many of its questions are in the dataset.
If you want to know how the templates were written and checked, see Question creation.
Evaluating a model
There are two sources of questions, and which to use depends on what the results are for. For results that other people can compare with, use the fixed dataset:
from datasets import load_dataset
ds = load_dataset("narunraman/steer_me", "law_of_demand")For questions no model has seen, ask the API for fresh ones. The same seed returns the same questions, so a fresh sample can still be reproduced:
curl "https://steer-benchmark.cs.ubc.ca/api/sample?element_name=law_of_demand&n=5&seed=42"
Then give the questions to your model, collect its answers, and score them. Scoring covers the three choices involved: the format the question is put in, the metrics to report, and the diagnostics for finding where the model fails.
To add an element of your own, or to build a new set of questions, see autoSTEER.