STSTEER

Question creation

From a seed template to a question with a computed answer

Elements

An element is one ability a rational economic agent needs. Elements are grouped into modules, and modules into settings, such as consumption decisions or strategic decisions. Together they form a taxonomy, which is what you browse on this site.

Each question in an element is described by three things: its type, a distinct way of testing the element (for profit maximization, asking for the maximum profit and asking for the labor needed to reach it are two types); its domain, the subject matter, such as medicine, finance or sports; and its perspective, whether it is written in the first, second or third person.

Templates

We do not write questions one at a time. For each type there are a few gold-standard seed templates. A template is a question with labeled fields where the numbers go:

Suppose the price for a basketball is start price, and the quantity demanded is start quantity. If the price of a basketball increases to end price, what is the most likely new quantity demanded?

A language model then rewrites each seed for every other domain and perspective, keeping the same labeled fields. Models do not always keep the economic meaning intact when they change the setting, so every rewritten template is reviewed and edited where needed. A second pass asks the model to rephrase each template while keeping its domain, perspective and fields fixed, which gives many differently worded versions of the same question.

Filling in a question

A question is produced by filling a template's fields with random values. The values are constrained to make economic sense: demand curves slope downward, equilibrium prices are positive, and so on. The correct answer and the wrong options are then computed by code, not written by a model. This is what happens each time you press "Generate a question" on an element page, and no language model is involved at that point.

Checks

Two checks guard against bad templates. First, we fill a template and ask a language model to solve it; disagreements often point to a wording problem or a mistake in the filled values. We use this only to fix errors, never to tune how hard the questions are. Second, the generated templates are spot-checked, and every element was reviewed before release.

Why generate questions

A fixed list of questions ends up in training data, and a model that has seen the test looks better than it is. Because questions here come from templates and random values, new ones can be produced at any time. Rewording matters as well: a large part of the variation in model performance comes from how a question is phrased, so testing many phrasings is a more honest measure than testing one.

The fixed dataset on Hugging Face is one sample from this process, kept stable so that results can be compared. The API gives you fresh samples.