Almost every Big Five result you will ever see reports five numbers. That is the level the model is famous at, and for a first look it is the right level. But the five are averages of narrower components, and the averaging is lossy in a specific and consequential way: it can hand two people the same score while describing opposite behaviour. The IPIP-NEO-120 exists to un-average it.
The Averaging Problem
Take Conscientiousness, the domain most often used in work contexts. It contains six facets, and two of them are orderliness — how much structure you impose on your environment — and self-discipline, how reliably you continue with something once the initial motivation has gone. They are related, but they are not the same thing, and plenty of people are high in one and low in the other.
Now consider two people who both score in the middle of Conscientiousness. The first is highly orderly and poorly self-disciplined: immaculate systems, immaculate lists, and a pattern of abandoning things at the two-thirds mark. The second is the reverse: works in visible chaos, and finishes everything they start. Identical domain score. Opposite people. Anyone acting on that score alone is acting on a number that describes neither of them.
This is not an edge case invented to make the point. It is the ordinary situation whenever a domain score lands anywhere near the middle, and middle is where most scores land. A mid-range domain score is the least interpretable output in personality assessment, and it is the one produced most often.
What Thirty Facets Actually Buys You
The Five-Factor Model as elaborated by Costa and McCrae organises each of the five domains into six facets, giving thirty in total. The IPIP-NEO-120 measures all thirty using public-domain items, four per facet, which is where the 120 comes from. Johnson (2014) developed and validated the 120-item form specifically as a shorter alternative to longer IPIP-NEO versions while keeping facet-level resolution intact.
The gain is not more precision on the same information — it is different information. A domain score answers “how much of this trait”. A facet profile answers “which kind”. High Extraversion built on gregariousness and warmth is a sociable person. High Extraversion built on assertiveness and excitement-seeking is a very different colleague, and both read as “extravert” at the domain level.
The same holds inside Neuroticism, where anxiety and anger are separate facets, and inside Agreeableness, where trust and modesty routinely diverge. In each case the domain label is accurate and nearly useless, and the facet pattern is the thing you would actually want to know before working with someone — or before drawing a conclusion about yourself.
Why Short Tests Cannot Do This
It is a straightforward arithmetic constraint rather than a quality judgement. Estimating a score reliably needs several items; thirty facets therefore need somewhere around a hundred items minimum. A 50-item Big Five test spends ten items per domain, which is enough to place the domain and nowhere near enough to split it six ways. A ten-item instrument spends two per domain and is measuring the domain roughly on purpose.
That is a legitimate trade, and the short instruments are not defective — they are answering a different question at a different cost. The 50-item Big Five test gives you a solid domain profile in a few minutes, and the TIPI-10 gives you a rough one in about two. Both are the right choice when the question is “roughly where do I sit”.
What no short instrument can do is tell you which half of a middling domain score you are. If that is your question — and it usually becomes your question the second time you take a Big Five test — the item count is not negotiable.
The Public-Domain Part Matters More Than It Sounds
Most well-validated personality inventories are commercial products, licensed per administration, which is why a facet-level profile has historically cost money and arrived through a certified practitioner. The International Personality Item Pool was built to change that: the items are in the public domain and the project states plainly that anyone may use any IPIP scale for any purpose, commercial use included, without permission.
The practical consequence is that a thirty-facet profile can be offered free, which would otherwise be impossible. The less obvious consequence is scientific: because the items are open, the instrument has been administered and analysed by many independent groups rather than by a single vendor with an interest in the outcome. Open items make a scale checkable.
It also means the IPIP-NEO-120 is not proprietary to any one site. What differs between providers offering it is the scoring, the norms used to convert raw scores into percentiles, and the quality of the write-up — not the instrument.
Reading a Facet Profile Without Overreading It
Start by finding your widest within-domain spread. The domain where your six facets disagree most is the domain where the summary score was misleading you, and it is the one worth reading properly. A domain whose six facets are all clustered together is a domain where the short-test answer was fine.
Then resist the temptation to treat thirty numbers as thirty findings. Facet scores are less reliable than domain scores — that is the unavoidable cost of splitting a fixed item budget — so a single facet sitting slightly above or below the others is noise until it recurs. What is trustworthy is a large, coherent pattern: two or three facets in a domain clearly separated from the rest.
And keep the same caution that applies to every self-report instrument. This is a description of how you currently see and report yourself, which correlates with behaviour but is not a measurement of it, and it will move somewhat with mood, life stage and how honestly you were willing to answer. There is a fuller account of what self-report can and cannot do in how personality tests work.
If you want the deep profile, the IPIP-NEO-120 is 120 items and about 30 minutes and returns all five domains with all thirty facets underneath. If you want to see how the model compares to its main rival — which adds a sixth domain rather than subdividing the five — HEXACO versus the Big Five covers that argument.