Skip to main content

Do AI Models Have a Personality? What Tests Show

|October 5, 2026|4 min read

Skip the article — take the Big Five test now

4 min · 50 questions · full breakdown with career matches

Start the Big Five test

A personality test can give an AI model a score, but the score depends on how the test is put to the model. We gave 16 language models four questionnaires under 11 setups. For the 13 models we could call through paid APIs, changing only the setup, with the model and the questions the same, moved scores by a median of 1.41 standard deviations of the human test-takers who took the same test. One test result is not a stable property of a model.

What we did

The 16 models include Claude, Gemini, GPT-family, Llama, DeepSeek, Mistral and Qwen models. Each answered four questionnaires: Big Five (50 items), RIASEC (60), Dark Triad (18) and Multiple Intelligences (40). Each model got the same 1,200 tasks, spread over 11 setups. Every score was then placed among English-speaking people who had answered the same items on JobCannon: 523 to 3,418 people, depending on the test. Those people are visitors who chose to take the test, so the comparison shows where an answer pattern lands, and it is not a population norm.

What we changed

For the 13 models reached through paid APIs, ten of the setups ran at temperature 0, the most repeatable setting, and differed only in how the test was put to the model: a straight rerun, the answer options in reverse order, the instruction reworded, the whole test in one request, and four shuffled item orders. An eleventh setup asked the same items at temperature 1.0 three times, to measure ordinary sampling noise.

what we measured (13 models reached through paid APIs)result
model-by-scale cells where the setup spread beat the sampling spread262 of 280 (median ratio 2.89)
median range between the highest and lowest setup1.41 human standard deviations (largest 3.84)
cells where that range was at least 1 human standard deviation217 of 280
median width of the percentile interval a model spans on one scale41.3 points
cells where the percentile interval was 50 points or wider89 of 280

Read this before you trust a "ChatGPT personality test" screenshot

A screenshot shows one setup. Ours says that the same model, asked the same questions, can land at very different percentiles when the options are reversed or the instruction is reworded. We did not test the ChatGPT app itself. We tested models through their APIs and command-line tools, with the model ids as listed in the dataset.

Habits show up more than content

Some models agree with items whatever the wording. On the Big Five, six models sat at or above the 95th human percentile for agreeing with positively and negatively worded items alike, and two sat at or below the 5th. One model answered every RIASEC item and every Multiple Intelligences item with the same option under the main setup. A percentile from an administration like that says nothing about the items.

How to run a fairer test yourself

These steps come from the setups we tried. Our data show that these choices move scores. They do not show which choice is the right one.

  1. Run the test at least twice.
  2. Reverse the option order and map the answers back.
  3. Reword the instruction once.
  4. Report the range across runs, not a single number.
  5. Look for a constant answer or an agree-with-everything pattern before reading any score.
  6. Read a percentile as "where this answer pattern lands among these people", never as a trait the model has.

What we are not saying

We do not claim that any model has a personality, or that one setup is the correct one. We tried one family of instructions, and a different one, such as role-play framing, could move answers further. Model ids are not dated snapshots, so a later call to the same id may reach a different model. The runs took place between 2026-09-30 and 2026-10-05.

FAQ

Can ChatGPT take a personality test?

Any chat model can be given the items and will return answers that can be scored. We did not test the ChatGPT app, and the answers we collected from other models changed with the setup, as in the table above.

Does a high score mean the AI has that trait?

Our results do not support reading it that way. The same model's percentile on one scale spanned a median of 41.3 points across ten setups.

How is a score calculated?

Plain arithmetic: for each scale, add the points for the chosen answers, divide by the most the scale could give, multiply by 100 and round. The scoring key is public.

Where are the data?

In the open dataset described below. It holds every stored answer, so anyone can recompute the results.

Data

huggingface.co/datasets/PeterKol/llm-questionnaire-protocol-sensitivity (commit 21b07690e79af51e489122501538541e75431c76, CC BY 4.0). The human comparison set is huggingface.co/datasets/PeterKol/jobcannon-psychometric-responses.

Ready when you are

Find your Big Five personality profile in 4 minutes.

50 questions. Full result with strengths, blind spots, and careers matched to your type from a database of 2,521 professions.