Reliability, structure, norms and stated limitations for the JobCannon Big Five assessment, computed from 3,472 completed first-party results collected between 2026-03-10 and 2026-08-07.
Every figure below was computed from real answer vectors through the same item-to-scale key the live test scores with. Nothing is quoted from the source instrument’s published studies, and nothing is hand-entered. Where we have no evidence — test-retest stability, criterion-related validity, adverse impact — section 9 says so instead of leaving the reader to assume.
The JobCannon Big Five is the IPIP 50-item Big-Five Factor Markers (Goldberg, 1992) — public domain, administered verbatim. It is not an adaptation, a shortened form, or an item set written in-house against the five-factor model: it is the published marker set, which means the substantial item-level literature on the IPIP-50 applies to this test directly rather than by analogy.
| Items | 50 |
| Scales | 5 (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) |
| Items per scale | 10 |
| Response format | 5-point Likert, Very Inaccurate → Very Accurate |
| Reverse-keyed items | 18 of 50 |
| Licence | Public domain (International Personality Item Pool) |
Scroll the table sideways for the rest of the columns.
Reverse-keyed items are distributed across all five scales — Openness 3, Conscientiousness 4, Extraversion 5, Agreeableness 4, Neuroticism 2 — so no scale can be answered high or low by a respondent who never reads past the first few words. Section 8 reports how often that happens anyway.
Self-administered in a browser, unproctored, with no time limit and no adaptive branching. Every respondent sees all 50 items in a fixed order with a fixed set of response options. Items are never rotated and never sampled from a larger pool, which is the property that makes a stored answer interpretable years later — see the integrity note at the end of section 3.
Scoring is a straight sum. Each response carries a weight along the published response scale — 5-point Likert, Very Inaccurate → Very Accurate — reverse-keyed items are inverted before summing, and the 10 item weights on a scale are added to a raw total. No response is imputed, no score is corrected for social desirability, and no score is adjusted after the fact.
The score JobCannon returns for a scale is the raw total divided by the maximum obtainable on that scale, expressed as a percentage. A score of 70 on Agreeableness means the respondent selected responses worth 70% of the ceiling — it does not mean they scored higher than 70% of people.
This distinction matters because the two metrics diverge sharply on some scales. In this sample the median Openness score is 72.5 and the median Agreeableness score is 70 — so a score of 70 on either is an ordinary result, not a high one. Section 6 publishes the observed cutoffs that convert a score into a percentile. Anyone building a threshold, a band or a cut score on top of these numbers should read them there first.
The analysis set is every completed Big Five result held on 2026-08-07: 3,697 rows scanned, 3,472 analysed. Two screens ran before anything was computed — 220 rows were dropped for holding a number of answers other than 50, and 5 for holding a response index outside the valid range. Nothing else was excluded: no outlier trimming, no removal of fast responders, no dropping of results that spoiled a coefficient.
It is a self-selected sample of people who chose to take a personality test on a consumer website, not a stratified population panel. That is the single most important caveat on the norms in section 6, and it is why the percentile table is labelled as this population rather than as a national norm. 955 of the results are attached to a registered account; the remainder were taken without signing in.
A stored answer is a position, not a text. If a test rotates its option order or serves questions from a pool, then a position recorded under one arrangement is meaningless under another, and a coefficient computed across both is computed against questions nobody was asked. Every figure in this manual therefore depends on the Big Five having exactly one fixed arrangement — so the generator asserts it rather than assuming it, and refuses to emit anything if the assertion fails.
On this run: 0 rows carried an option-order stamp, 0 carried a question-pool index, and all 3,697 rows were recorded under a single test version. The alphas below were also reproduced independently by /research/reliability, which is a different generator on its own monthly schedule, with its own sample: 3,372 results pulled 2026-08-06, against the 3,472 pulled 2026-08-07 here. Two scripts written apart, over two samples pulled a day apart — and the largest disagreement on any scale is .001. Where the two pages differ, this one names its own pull date beside every figure.
Internal consistency answers one question: do the ten items on a scale agree with each other enough to be averaged into a single score? Cronbach’s alpha summarises that agreement. Confidence intervals are bootstrap percentile intervals over 400 resamples of the respondents, computed with a fixed seed so the interval reproduces exactly on re-run. We use resampling rather than the standard F-based interval because the latter assumes essentially tau-equivalent items and multivariate normality, and a five-point Likert item satisfies neither.
| Scale | Alpha | 95% CI | SEM | Mean r | Band |
|---|---|---|---|---|---|
| Openness | .816 | .806–.825 | 7.3 | .308 | Good |
| Conscientiousness | .777 | .764–.791 | 8.2 | .257 | Acceptable |
| Extraversion | .862 | .853–.869 | 7.9 | .384 | Good |
| Agreeableness | .812 | .803–.822 | 7.8 | .309 | Good |
| Neuroticism | .858 | .849–.866 | 7.8 | .374 | Good |
| Mean across scales | .825 | — | — | — | Good |
Scroll the table sideways for the rest of the columns.
Reading the columns. By the conventional bands .70 is acceptable, .80 good and .90 excellent for a research or self-discovery instrument. SEM is the standard error of measurement expressed on the same 0–100 metric the product reports, so it can be used directly: a reported Extraversion score of 60 carries a 68% interval of roughly 52–68 and a 95% interval of roughly 44–76. Any banding or thresholding built on these scores should be at least that wide. Mean r is the average inter-item correlation, which unlike alpha does not rise simply because a scale is long — the .15–.50 range is the usual target, and all five scales sit inside it.
The two weakest scales are worth naming rather than averaging away. Conscientiousness (.777) and Agreeableness (.812) sit lowest, and the item table below shows where that comes from.
Each item is correlated with the sum of the other nine on its scale. An item below .30 is contributing little; across all 50 items, 2 fall below that line. They are named here rather than dropped, because the instrument is administered as published and quietly deleting an item would make this a different test from the one people took.
| Item | Stem | Key | r |
|---|---|---|---|
| Q45 | Spend time reflecting on things | Direct | .367 |
| Q20 | Am not interested in abstract ideas | Reverse | .445 |
| Q35 | Am quick to understand things | Direct | .445 |
| Q40 | Use difficult words | Direct | .456 |
| Q5 | Have a rich vocabulary | Direct | .466 |
| Q10 | Have difficulty understanding abstract ideas | Reverse | .499 |
| Q15 | Have a vivid imagination | Direct | .522 |
| Q30 | Do not have a good imagination | Reverse | .529 |
| Q25 | Have excellent ideas | Direct | .589 |
| Q50 | Am full of ideas | Direct | .658 |
Scroll the table sideways for the rest of the columns.
| Item | Stem | Key | r |
|---|---|---|---|
| Q13 | Pay attention to details | Direct | .284 |
| Q33 | Like order | Direct | .398 |
| Q48 | Am exacting in my work | Direct | .399 |
| Q38 | Shirk my duties | Reverse | .416 |
| Q8 | Leave my belongings around | Reverse | .462 |
| Q43 | Follow a schedule | Direct | .477 |
| Q3 | Am always prepared | Direct | .486 |
| Q18 | Make a mess of things | Reverse | .490 |
| Q23 | Get chores done right away | Direct | .506 |
| Q28 | Often forget to put things back in their proper place | Reverse | .508 |
Scroll the table sideways for the rest of the columns.
| Item | Stem | Key | r |
|---|---|---|---|
| Q26 | Have little to say | Reverse | .484 |
| Q36 | Don't like to draw attention to myself | Reverse | .501 |
| Q41 | Don't mind being the center of attention | Direct | .531 |
| Q11 | Feel comfortable around people | Direct | .538 |
| Q16 | Keep in the background | Reverse | .570 |
| Q1 | Am the life of the party | Direct | .589 |
| Q46 | Am quiet around strangers | Reverse | .597 |
| Q6 | Don't talk a lot | Reverse | .620 |
| Q21 | Start conversations | Direct | .623 |
| Q31 | Talk to a lot of different people at parties | Direct | .662 |
Scroll the table sideways for the rest of the columns.
| Item | Stem | Key | r |
|---|---|---|---|
| Q12 | Insult people | Reverse | .265 |
| Q2 | Feel little concern for others | Reverse | .367 |
| Q37 | Take time out for others | Direct | .481 |
| Q47 | Make people feel at ease | Direct | .489 |
| Q27 | Have a soft heart | Direct | .503 |
| Q22 | Am not interested in other people's problems | Reverse | .530 |
| Q42 | Feel others' emotions | Direct | .535 |
| Q7 | Am interested in people | Direct | .574 |
| Q32 | Am not really interested in others | Reverse | .583 |
| Q17 | Sympathize with others' feelings | Direct | .647 |
Scroll the table sideways for the rest of the columns.
| Item | Stem | Key | r |
|---|---|---|---|
| Q19 | Seldom feel blue | Reverse | .305 |
| Q9 | Am relaxed most of the time | Reverse | .397 |
| Q14 | Worry about things | Direct | .528 |
| Q24 | Am easily disturbed | Direct | .563 |
| Q44 | Get irritated easily | Direct | .606 |
| Q49 | Often feel blue | Direct | .612 |
| Q34 | Change my mood a lot | Direct | .640 |
| Q4 | Get stressed out easily | Direct | .650 |
| Q29 | Get upset easily | Direct | .663 |
| Q39 | Have frequent mood swings | Direct | .682 |
Scroll the table sideways for the rest of the columns.
Reliability says the items on a scale agree. Structure asks the harder question: are there five things here, and are they the five the scoring key claims? Two checks follow — how far apart the five scales sit from each other, and where the 50 items land when the factors are recovered from the data instead of assumed.
| Scale | Open. | Consc. | Extra. | Agree. | Neuro. |
|---|---|---|---|---|---|
| Openness | — | .207 | .171 | .318 | −.086 |
| Conscientiousness | .207 | — | .031 | .143 | −.266 |
| Extraversion | .171 | .031 | — | .275 | −.173 |
| Agreeableness | .318 | .143 | .275 | — | .043 |
| Neuroticism | −.086 | −.266 | −.173 | .043 | — |
Scroll the table sideways for the rest of the columns.
The five scales are close to independent. The largest correlation in the matrix is .318 (Agreeableness and Openness), which shares about 10% of its variance — enough to be a real relationship, nowhere near enough for the two scales to be measuring the same thing. That is what discriminant validity looks like on a five-factor instrument: five scores that carry five different pieces of information rather than one score dressed up five ways.
All 50 items were entered into a principal-components analysis of their correlation matrix, and 5 components were rotated to simple structure by varimax with Kaiser normalisation. The rotated components were then matched to the five keyed dimensions by their dominant loadings, and each item assigned to whichever component it loaded on most strongly.
| Components rotated | 5 |
| Distinct keyed dimensions recovered | 5 of 5 |
| Items loading on their keyed factor | 48 of 50 (96%) |
| Variance explained by the 5 components | 43.9% |
| First eight eigenvalues | 6.98, 5.38, 4.16, 2.84, 2.59, 1.94, 1.30, 1.19 |
Scroll the table sideways for the rest of the columns.
48 of 50 items load most strongly on the factor their scoring key assigns them to, and the five rotated components map onto five distinct dimensions rather than two of them collapsing into one. That is the result the instrument claims, recovered from the data without being told the answer.
What the eigenvalues also say. Kaiser’s criterion — retain components with an eigenvalue above 1 — would retain 8 here, not five, and it is more honest to print that than to report only the five that suit us. The reason five is still the right number to rotate: the first five eigenvalues (6.98, 5.38, 4.16, 2.84, 2.59) each stand clearly above the sixth (1.94), the scree bends there, and the five components that come out are interpretable as the five keyed dimensions. Components six and beyond are the minor item clusters any 50-item inventory produces — Kaiser’s rule is well known to over-retain them.
The items that did not fit. 2 of 50 loaded most strongly somewhere other than their keyed factor:
| Item | Stem | Keyed to | Loaded on | Loading |
|---|---|---|---|---|
| Q13 | Pay attention to details | Conscientiousness | Openness | 0.41 |
| Q18 | Make a mess of things | Conscientiousness | Neuroticism | −0.46 |
Scroll the table sideways for the rest of the columns.
Both are Conscientiousness items, and both drift in ways the literature would predict rather than at random — attention to detail sits close to the intellect facet of Openness, and making a mess of things loads on Neuroticism through the distress that comes with disorganisation. They are reported, kept in their keyed scale, and left in the alpha you see above. Section 4 shows both of them in the Conscientiousness item table, which is where its lower coefficient comes from.
This is the table that turns a score into a statement about a person. Each row is one scale; each column is the score at or below which that percentage of this sample falls. All values are on the 0–100 percent-of-maximum metric described in section 2.
| Scale | p5 | p10 | p25 | p50 | p75 | p90 | p95 | Mean | SD |
|---|---|---|---|---|---|---|---|---|---|
| Openness | 40 | 47.5 | 57.5 | 72.5 | 82.5 | 92.5 | 95 | 70.2 | 17.0 |
| Conscientiousness | 32.5 | 37.5 | 50 | 60 | 72.5 | 82.5 | 90 | 60.6 | 17.4 |
| Extraversion | 12.5 | 17.5 | 30 | 47.5 | 62.5 | 75 | 80 | 46.4 | 21.3 |
| Agreeableness | 35 | 42.5 | 55 | 70 | 80 | 90 | 95 | 67.2 | 17.9 |
| Neuroticism | 17.5 | 25 | 40 | 55 | 70 | 82.5 | 87.5 | 54.4 | 20.8 |
Scroll the table sideways for the rest of the columns.
Two things this table is for. First, converting a reported score into a percentile — the step that has to happen before any threshold is set, because a raw 70 on Agreeableness is the median of this population and not a high score. Second, showing how differently the five scales are distributed: Extraversion is near-symmetric and wide (mean 46.4, SD 21.3), while Openness is high and compressed (mean 70.2, SD 17.0), so the same five-point difference means considerably more on one than the other.
| Scale | Skew | Kurtosis | Shape |
|---|---|---|---|
| Openness | −0.42 | −0.19 | Ceiling-leaning |
| Conscientiousness | −0.07 | −0.22 | Near-symmetric |
| Extraversion | 0.06 | −0.57 | Near-symmetric |
| Agreeableness | −0.49 | −0.05 | Ceiling-leaning |
| Neuroticism | −0.24 | −0.39 | Near-symmetric |
Scroll the table sideways for the rest of the columns.
No scale is badly non-normal — the largest absolute skew on any scale is 0.49 and the largest absolute kurtosis is 0.57, both well inside the range where means, standard deviations and the parametric statistics above behave. The mild negative skew on Agreeableness and Openness is the usual pattern for self-report on socially valued traits, and it is one more reason to read a score against the percentile column rather than against 50.
Norm caveat. This is a norm for this population — people who chose to take a personality test on a consumer website between 2026-03-10 and 2026-08-07. It is not a census-weighted national norm, and it is not an applicant norm for any specific job. An employer comparing candidates should expect its own applicant pool to sit differently, and should build its bands against that pool rather than against this table.
The same 50 items, the same scoring key, answered in different languages. Because the instrument and the pipeline are identical across locales, the coefficients are directly comparable — which is unusual, and is the check that catches a translation that has quietly changed what an item asks. Cells below n = 100 are not published.
| Language | n | Open. | Consc. | Extra. | Agree. | Neuro. | Mean alpha |
|---|---|---|---|---|---|---|---|
| English | 2,133 | .821 | .797 | .867 | .800 | .873 | .832 |
| Not recorded | 734 | .803 | .707 | .848 | .819 | .810 | .797 |
| Japanese | 126 | .841 | .809 | .874 | .835 | .915 | .855 |
| Korean | 123 | .864 | .857 | .913 | .776 | .880 | .858 |
Scroll the table sideways for the rest of the columns.
Mean alpha runs from .797 to .858 across the 4 cells that clear the floor. The scales hold together about as well in Japanese and Korean as in English, which is evidence that the translations preserved what the items measure. “Not recorded” is results saved before the locale column was written, not a language.
What this is not. Comparable alphas are evidence of comparable internal structure, not proof of measurement invariance. Formal invariance testing — a multi-group confirmatory model constraining loadings and intercepts across languages — has not been run, so cross-language comparison of scores (as opposed to reliability) is not something this manual supports.
An unproctored self-report instrument has one structural weakness: it cannot make anyone answer carefully. The published estimate is that roughly 10–12% of respondents in unsupervised settings answer at least part of a long questionnaire without engaging with it (Meade & Craig, 2012). Rather than assume it away, the platform measures it.
| Longest run of an identical response — median | 4 items |
| Longest run — 95th percentile | 7 items |
| Results with a run of 10 or more | 69 (1.99%) |
| Completion time — median | 3 min 46 s |
| Completion time — 10th to 90th percentile | 2 min 16 s to 7 min 49 s |
| Results faster than one second per item | 1.61% |
Scroll the table sideways for the rest of the columns.
The two standard careless-responding indices both land low. The median respondent’s longest run of an identical answer is 4 items — far below the usual LongString threshold of half a scale length — and 1.99% of results contain a run of ten or more. 1.61% were completed faster than one second per item, the conventional lower bound for having read the stem.
These indices are computed on every result, and where a hiring flow surfaces them they are surfaced as flags for a human to read, never as an automatic adjustment to the score. That is deliberate: correcting a score for a suspected response style substitutes one unverified assumption for another, and the recommendation in the literature is to flag and let a person decide. The counts above are reported here because a manual that reports only its coefficients is telling you half of what it knows.
The sections above are the evidence we have. This section is the evidence we do not, and it is here rather than in a footnote because a technical manual that only lists its strengths is a brochure.
Stability over time cannot be reported. Across the whole set, 1 registered user has sat the assessment more than once, and 0 pairs of sittings are at least a day apart — so there is nothing to correlate. A test-retest coefficient is the right index for whether a score would be the same next month, and we do not have one. Internal consistency (section 4) is a different property and is not a substitute for it.
This manual contains no local study relating scores to job performance, tenure, training outcomes or any other criterion, because we hold no criterion data. The general literature on Big Five predictors of work outcomes applies to the IPIP-50 as an instrument, and that literature is where a claim about prediction has to come from — it is not a claim this dataset can make. A criterion study is job-specific and sample-specific by nature; an employer that wants one for its own role has to run it on its own applicants and outcomes.
No selection-rate analysis by protected class can be produced from this dataset, because it contains no demographic data. A voluntary self-disclosure schema exists and is built to the shape the Uniform Guidelines require — every field skippable, stored in a table entirely separate from scores, never readable by the org admin making the decision, exposed only as a k-anonymised aggregate — but it is not yet wired into any assessment flow and currently holds zero records. Until it does, adverse-impact analysis for a US selection procedure is the employer’s, computed on its own applicant flow.
Section 7 shows comparable reliability across four language cells. It does not show that the items function identically across them, which requires a multi-group confirmatory model that has not been run. Comparing a score obtained in Japanese against a score obtained in English is therefore outside what this manual supports.
The IPIP-50 carries no lie scale, no infrequency scale and no social-desirability correction, and none has been added. Section 8 measures careless responding, which is a different failure mode from deliberate impression management. In a high-stakes setting a respondent who wants to present a particular profile can, and no self-report personality instrument prevents that.
Neuroticism as measured here is a normal-range personality trait, not a mental health screen. No scale on this instrument diagnoses anything, and none should be read as though it does.
What the evidence here supports: using the five scores as a structured, reasonably precise description of normal-range personality — for self-understanding, for development conversations, for team composition discussions, and as one input among several in a selection process where a human makes the decision.
What it does not support: a score used alone as a pass/fail gate; a cut score set without reading section 6 and the standard error in section 4; comparison of scores obtained in different languages; or any inference about a clinical condition. Under the US Uniform Guidelines the employer using an assessment carries the burden of showing that what it measures is job-related for the position in question, and no vendor manual discharges that for a job it has never seen.
Every number on this page is emitted by one script, scripts/research/bigfive-tech-manual-stats.mjs, which reads raw answer vectors, parses the 50 items and their keys straight out of the live scoring configuration, and refuses to run if fewer than 50 parse. It writes a single JSON artifact; this page renders that artifact and computes nothing of its own. There is no path by which a figure here can drift from the data it claims to describe.
Methods, in brief: Cronbach’s alpha from the item covariance matrix; confidence intervals by bootstrap percentile over 400 resamples with a fixed seed; SEM as SD × √(1 − α); item-total correlations corrected by excluding the item from its own total; principal components by Jacobi eigendecomposition of the correlation matrix, rotated by varimax with Kaiser normalisation; percentiles by linear interpolation on the sorted population.
Published under CC BY 4.0. Quote the edition date with any figure — this is recomputed from live data, so a coefficient without a date is a coefficient nobody can check.
JobCannon (2026). Big Five Personality (IPIP-50) technical manual: internal consistency, structure and norms from 3,472 first-party results [Technical report]. JobCannon Research. Retrieved 2026-08-07, from https://jobcannon.io/research/big-five-technical-manual
Both, in part, and it says which part is missing. It reports internal-consistency reliability (Cronbach's alpha per scale, with bootstrap confidence intervals and standard errors of measurement) and structural validity (inter-scale correlations and a factor analysis of all 50 items). It does not report test-retest reliability or a criterion-related validity study against job performance, because we do not have the data for either — section 9 states that plainly rather than implying otherwise.
The IPIP 50-item Big-Five Factor Markers (Goldberg, 1992) — public domain. It is administered verbatim — 50 items, 10 per factor, a 5-point Likert, Very Inaccurate → Very Accurate response scale, 18 of them reverse-keyed. Because the item set is the published marker set rather than an adaptation, the item-level literature on IPIP-50 applies to this test directly.
From 3,472 completed result rows collected between 2026-03-10 and 2026-08-07, scored through the same item-to-scale key the live test uses. Nothing is quoted from the IPIP literature and nothing is hand-entered. One script reads the raw answer vectors and emits every figure on this page; re-running it can move a number down as well as up.
It means the person selected responses worth 70% of the maximum obtainable on that ten-item scale. It is not a percentile against a norm group. Section 6 publishes the observed percentile cutoffs for this population, which is what converts one to the other — on Openness, for example, the median score is 72.5, so 70 is around the middle of this sample rather than high.
Not on its own, and not without the employer doing work this manual cannot do for them. Under the US Uniform Guidelines the employer must be able to show that what it measures is job-related for the position in question, and must be able to compute selection rates by protected class on its own applicant flow. This manual establishes that the scales measure something coherent and stable in structure. It does not establish that a given score predicts performance in a given job, and no reputable Big Five manual can establish that for a job it has never seen.
On demand from live data, not on a fixed schedule — this edition was generated 2026-08-07. Quote the pull date alongside any figure you cite. Reliability across the whole JobCannon catalogue is recomputed monthly and published separately at /research/reliability.