Abstract Reasoning
n = 203 · 25 items · 4 scales
Built on Cognitive aptitude
- Analogies
- 0.85
- Pattern Series
- 0.79
- Rule Transformations
- 0.76
- Classification
- 0.75
- Overall
- 0.94
Reliability is the first number an I/O psychologist checks. This is Cronbach’s alpha — internal consistency — computed from 24,528 completed JobCannon results, not a textbook figure or a marketing claim.
Strongest so far:Abstract Reasoningα = 0.94over 203 takers
Most test sites quote the reliability of the instrument they adapted — a figure from the original paper, measured on the original sample, sometimes decades ago and in one language. That number describes a study, not the questionnaire you just filled in. Every alpha below was computed from 24,528 real answer vectors collected by this site, through the same scoring code that produced the results people were shown, across 508 items and 60 scales.
That cuts both ways, and it is the reason to publish it. A recomputed figure can fall as well as rise, and a scale that looks weaker here than in its source paper is telling you something real about how it behaves in the wild. 20 of 23 tests clear the conventional 0.70 threshold and 19 clear 0.80.
n = 203 · 25 items · 4 scales
Built on Cognitive aptitude
n = 209 · 25 items · 4 scales
Built on Cognitive aptitude
n = 117 · 25 items · 4 scales
Built on Cognitive aptitude
n = 149 · 12 items · 1 scale
Built on Grit Scale (Duckworth)
n = 152 · 12 items · 1 scale
Built on Clance IP Scale
n = 108 · 8 items · 1 scale
n = 130 · 24 items · 1 scale
Built on Attention and self-regulation patterns — Built on Hartmann (1997) and Hallowell & Ratey (2011). Original 24 items, not the ASRS v1.1 — we do not administer the WHO screener.
n = 112 · 25 items · 4 scales
Built on Cognitive aptitude
n = 178 · 13 items · 2 scales
Built on Copenhagen Burnout Inventory (Kristensen)
n = 262 · 10 items · 2 scales
Built on Empathy Quotient (Baron-Cohen)
n = 438 · 120 items · 5 scales
Built on IPIP-NEO-120
n = 252 · 20 items · 1 scale
Built on Autism-Spectrum Quotient constructs (Baron-Cohen et al. 2001) — Original 20 items, not the AQ-50 — we do not administer the clinical screener.
n = 291 · 32 items · 4 scales
Built on Cognitive aptitude
n = 434 · 10 items · 2 scales
Built on Rosenberg Self-Esteem Scale
n = 733 · 10 items · 2 scales
n = 15,354 · 20 items · 4 scales
Built on General cognitive aptitude
n = 3,372 · 50 items · 5 scales
Built on IPIP-50 (Goldberg)
Technical manual — confidence intervals, standard error of measurement, factor structure and percentile norms
n = 323 · 25 items · 4 scales
Built on Cognitive aptitude
n = 173 · 8 items · 1 scale
n = 1,101 · 12 items · 1 scale
Built on Type A/B personality
n = 141 · 8 items · 1 scale
n = 150 · 4 items · 1 scale
Built on Perceived Stress Scale (Cohen)
n = 146 · 10 items · 5 scales
Built on TIPI (Gosling 2003)
2 items per scale. Alpha rises with scale length, so a scale this short sits low by construction rather than by defect — the bands are the wrong ruler here. Why we publish it anyway
Cronbach’s alpha measures how consistently the items on a scale move together — whether the questions are really measuring one thing. By the standard bands: α ≥ 0.70 is acceptable, ≥ 0.80 good, and ≥ 0.90 excellent for a research or self-discovery instrument.
Every figure here is computed straight from real answer vectors using the exact same item-to-scale scoring as the live test — no hand-entered numbers. A test only appears once it has at least 100 completed results, so more tests join this page as volume grows. We never publish reliability for for-fun quizzes, even when an alpha could be computed — their raw responses are open anyway, as the entertainment quiz dataset, with the same warning attached there that this sentence makes here.
One test here serves rotating question sets — IQ Test (20) — so that two people taking it do not see the same questions. Item 3 of one set is a different question from item 3 of another, which means pooling every taker into one column would correlate answers to questions that were never asked together. The alpha above is computed inside each set and then averaged, weighted by how many people took it. Results recorded before we started storing which set a taker saw are excluded rather than guessed at.
Alpha is partly a function of how many questions a scale has. Lengthen a scale with equally good items and the coefficient rises on its own, which is why a two-question scale cannot reach the figure a fifty-question scale reaches, however well the two questions are written. Below about 3 items per scale the standard bands stop describing the instrument and start describing its length. One test here is built that way — Big Five Mini (TIPI-10), 10 items across 5 scales — and that is the design, not a defect: the whole point of an instrument that short is to cover a full model in about a minute.
Gosling, Rentfrow and Swann — the TIPI’s authors — say it themselves: “alphas are misleading when calculated on scales with small numbers of items” — and they recommend test-retest reliability as the appropriate index for a scale that short. We publish the figure as computed rather than adjusting it upward or quietly dropping the card, because a number you cannot see is worth less than a number you can put in context. Read the band label beside it as what alpha says, not as what the test is worth.
Ability tests (like the cognitive test) are split into their subscales so you can see each domain honestly rather than behind one flattering composite. Recomputed monthly. Last pulled 2026-08-06.
An alpha of 0.90 says the items on a scale agree with each other. It does not say the scale measures what its name claims. Ask ten near-identical questions about whether someone enjoys parties and you will get a superb coefficient and learn nothing about how they will do in a job — the questions agree because they are the same question. Consistency is the floor a scale has to clear before any other claim about it is worth discussing, not the claim itself.
So read this page as one specific check that passed, and read the Built on line under each test as the other half: it names the published instrument the item set is derived from, and that literature is where the evidence about what the scale predicts actually lives. Tests marked as framework-based rest on a model that is widely used but not validated the way a research inventory is, and the badge says so on every card rather than in a footnote.
None of these are clinical instruments, none diagnose anything, and none should decide a hire on their own. The distribution of results across the whole population is on the distributions page — a coherent scale and a plausible spread are two different things to check, and a test can pass one while failing the other.
These figures are published under CC BY 4.0. Reuse them with attribution, and quote the pull date alongside the number — they are recomputed monthly, so a coefficient without a date is a coefficient that cannot be checked.
JobCannon (2026). Internal-consistency reliability for 23 JobCannon assessments (n = 24,528) [Data set]. JobCannon Research. Retrieved 2026-08-06, from https://jobcannon.io/research/reliability
It asks whether the questions on a scale agree with each other. If a scale is really measuring one thing, someone who answers high on one of its items should tend to answer high on the rest, and alpha summarises how strongly that holds across every item and every respondent. It runs from 0 to 1. A low alpha means the items are pulling in different directions and the scale score is an average of things that do not belong together.
By the conventional bands, 0.70 and above is acceptable, 0.80 good, 0.90 excellent for a research or self-discovery instrument. Of the 23 tests published here, 20 clear 0.70 and 19 clear 0.80. The strongest is Abstract Reasoning at 0.94 over 203 takers. Very high alpha is not automatically better: past roughly 0.95 it usually means the same question was asked several ways rather than that the scale is unusually sound.
Because alpha climbs with scale length, a short scale is marked down by arithmetic before quality comes into it. Big Five Mini (TIPI-10) is scored from 10 items across 5 scales — 2 questions per scale. That is the design rather than a fault: the point of an instrument that short is to cover a whole model in about a minute, so it is judged by how closely it tracks the long form it compresses rather than by its own internal consistency. Gosling, Rentfrow and Swann — the TIPI’s authors — state that “alphas are misleading when calculated on scales with small numbers of items”, and recommend test-retest reliability as the appropriate index instead. We publish the computed coefficient anyway rather than adjusting it upward or dropping the card, and the validity badge reflects the published evidence about what the instrument predicts — a separate question from internal consistency.
No, and this is the distinction that matters most. Reliability is consistency; validity is whether the thing being measured consistently is the thing named on the label. A scale of ten near-identical questions about liking parties will have a superb alpha and still tell you nothing about job performance. Alpha is the floor, not the ceiling — a test that fails it cannot be measuring anything stable, but passing it proves only internal coherence.
From real answer vectors, using the same item-to-scale scoring the live test uses — nothing is hand-entered or copied from the source instrument's published figures. The set spans 24,528 completed results across 508 items and 60 scales, recomputed monthly and last pulled 2026-08-06. Because it is recomputed rather than quoted, a figure here can go down as well as up.
Two filters. A test needs at least 100 completed results before any alpha is published, so newer tests join as volume grows. And a test has to be built on a real psychometric model: for-fun quizzes are excluded even when an alpha can be computed for them, because publishing a reliability coefficient next to an entertainment quiz dresses it up as something it is not.
Yes, under CC BY 4.0 with attribution. Quote the pull date with the number, since these are recomputed monthly rather than fixed. Suggested form: JobCannon (2026). Internal-consistency reliability for 23 JobCannon assessments (n = 24,528) [Data set]. JobCannon Research. Retrieved 2026-08-06, from https://jobcannon.io/research/reliability