Skip to main content
ScienceBig Five

LLM Evaluation: How Test Setup Moves Scores

|October 5, 2026|6 min read

When you evaluate a language model with a questionnaire, how you put the test to the model is a source of variation that averaging does not remove. In our study of 16 models, the spread across ten test setups at temperature 0 was larger than the spread across three samples at temperature 1.0 in 262 of 280 model-by-scale cells, with a median ratio of 2.89.

The problem

A score for a model is read as a trait: this model is more agreeable than that one. The score also depends on choices the tester makes without noticing, such as the wording of the instruction, the order of the options, and whether items arrive one at a time or together. Sampling noise is the usual worry and can be reduced by averaging. Setup choices are a different source of variation. We call it protocol sensitivity.

Design

Four questionnaires with per-item human answers (Big Five, RIASEC, Dark Triad, Multiple Intelligences), 16 models, 1,200 tasks per model, 11 conditions.

conditionwhat changes
S0one item per request, human item order, human option order, instruction P1, temperature 0 (primary)
S0ridentical to S0, a second independent run
SRoptions shown in reverse order, answer mapped back
SPinstruction P2, a paraphrase of P1
ST1as S0 at temperature 1.0, three independent samples per item
B0, B0rthe whole test in one request, human item order; B0r is a rerun
BS1 to BS4the whole test in one request, item order shuffled with four fixed seeds

The ten temperature-0 administrations (S0, S0r, SR, SP, B0, B0r, BS1 to BS4) define the protocol spread. The three ST1 samples define the sampling noise. Item order can be varied only in whole-test mode, so item-order effects are mixed with the switch from single items to whole tests. Parsing is strict: a reply that is not an in-range option number counts as missing and is never defaulted.

Results for the 13 models reached through paid APIs

groupcells where protocol SD exceeds sampling SDmedian ratio
all 13 models262 of 2802.89
the 3 models from the earlier pilot, now with the fourth test63 of 663.12
the 10 models added in the full run199 of 2142.79
  • The range between the highest and lowest setup has a median of 1.41 human standard deviations (minimum 0.28, maximum 3.84). It is at least 1 in 217 of 280 cells and at least 2 in 47 of 280.
  • By test, the protocol spread is above the sampling spread in 58 of 65 Big Five cells, 67 of 72 RIASEC cells, 38 of 39 Dark Triad cells and 99 of 104 Multiple Intelligences cells.
  • A straight rerun at temperature 0 differs by a median of 0.00 human standard deviations and at most 0.94. Some endpoints do not return identical answers at temperature 0, and we did not verify whether each endpoint honours it.
  • The three models reached through subscription command-line tools, which expose no temperature, show the same direction (58 of 66 cells) with a smaller median range of 0.60 human standard deviations. That is not like for like with the API arm, because the tools add their own context and sample differently.

A checklist for evaluators

Offered as practice that follows from the setups we tried. It is not a finding about which setup is right.

  1. Report the range across setups next to the mean.
  2. Compare the setup spread with the sampling spread, and use the larger one as the error bar.
  3. Run at least a rerun, a reversed-option run and a reworded-instruction run.
  4. Log how many replies could not be read. Under the reworded instruction one model returned something that was not an option on 67 of 168 single-item requests.
  5. Flag constant-answer runs. They produce percentiles that say nothing about the items.
  6. Check whether your endpoint really honours temperature 0.
  7. Pin a dated model snapshot where the provider offers one. Our ids were not dated snapshots.

We do not claim to be the first to show that questionnaire scores for language models are unstable. Related papers, read from their abstracts: Dominguez-Olmedo, Hardt and Mendler-Dünner, "Questioning the Survey Responses of Large Language Models" (arXiv 2306.07951); Röttger et al., "Political Compass or Spinning Arrow?" (arXiv 2402.16786); Meyer, Garcia and Wulff, "Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact" (arXiv 2606.20205); Zierahn, Cachero, Korhonen and Oliver, "Personality Without Persons?" (arXiv 2607.02325). What we add is that each score is placed among human answers to the same items, so the size of the movement can be read in human standard deviations, and that protocol spread is compared with sampling spread like for like.

Limits

One family of instructions. The thresholds in the verdict rule (80% of cells, ratio of at least 2, under 5% for "few") are round numbers fixed after seeing partial data from 10 models. Sampling noise comes from three samples per item, and ten reruns of one temperature-0 setup were not run. Human percentiles come from self-selected site visitors. The 16 models are the ones reachable on our accounts and are not a sample of anything.

Reproduce

Node 20 or newer, no packages to install. Eight scripts recompute every table from the stored answers. The scoring key is a CSV with 168 rows and a separate scorer of about 50 lines. On about 616,000 answer vectors it matched the scorer of the product the human data comes from in every case, with 0 mismatches, and a deliberately wrong weight made 1,233 of 3,992 Big Five human vectors disagree.

FAQ

What is protocol sensitivity?

The amount by which a score changes when the way a test is administered changes and the model and the items stay the same.

Is sampling temperature the main source of variation?

Not in this study. Setup choices moved scores more than temperature-1.0 sampling in 262 of 280 cells.

Can I reproduce this without calling any model?

Yes. The dataset holds every stored answer and the scripts that recompute the results.

Data

huggingface.co/datasets/PeterKol/llm-questionnaire-protocol-sensitivity (commit 21b07690e79af51e489122501538541e75431c76, CC BY 4.0). The human comparison set is huggingface.co/datasets/PeterKol/jobcannon-psychometric-responses.

Ready when you are

Find your Big Five personality profile in 4 minutes.

50 questions. Full result with strengths, blind spots, and careers matched to your type from a database of 2,521 professions.