STSTEER

Datasets

A fixed sample of each benchmark, on Hugging Face

Each dataset is one sample from the generator, kept fixed so that results can be compared across papers. Every element contributes 20 questions for each combination of type, domain and difficulty, and at least 200 in total. There is one dataset per benchmark, and each can be loaded whole or one element at a time:

from datasets import load_dataset

steer    = load_dataset("narunraman/steer")       # everything
steer_me = load_dataset("narunraman/steer_me", "law_of_demand")   # one element

Every row has its setting and module, so a setting or module is a filter on the whole dataset. Each element also has a few_shot split of example questions from templates that the test split never uses. For questions outside the fixed sample, use the API. Both datasets are released under CC BY 4.0.