Before There Was a Test
The idea that reasoning could be separated from knowledge and measured on its own is barely a century old. Nineteenth-century mental testing measured reaction times, sensory discrimination and head size, and predicted almost nothing.
What changed the field was not a better apparatus. It was a statistical argument about what the numbers already collected had in common.
Spearman and the g Factor
In 1904 Charles Spearman noticed something in schoolchildren's marks that should not have been there: performance across unrelated subjects was positively correlated. Children good at classics tended to be good at mathematics and at pitch discrimination.
He proposed that a single general factor — g — contributed to performance on every cognitive task, alongside factors specific to each.
Why this mattered for testing
If a general factor exists, then some tasks tap it more purely than others. The measurement problem becomes a search: find the item type that loads most heavily on g and carries least specific content.
Abstract reasoning tests are, in a direct sense, the answer to the question Spearman posed. Everything that follows is the search he started.
Binet, and the Detour Through Content
Alfred Binet and Théodore Simon's 1905 scale, commissioned to identify French schoolchildren needing additional support, was the first practically useful intelligence test.
But it was heavily verbal and knowledge-laden. It worked well for the population it was built on and travelled badly to anyone outside it — a limitation that became visible the moment it was exported.
The Army tests and the first non-verbal push
The United States Army screened roughly 1.75 million recruits during the First World War. Army Alpha was written and required literacy; Army Beta was built for men who could not read English, using mazes, picture completion and figural tasks.
Army Beta was a landmark in scale and in method. It was also misused: results were cited in 1920s immigration debates in ways the instrument could not support, a caution that belongs in any honest history of the field.
Spearman's Students and the Analogy
Through the 1920s and 1930s, work in Spearman's circle converged on a specific item form. If g is about perceiving relationships, then the purest measurement is a task that presents relationships with no content attached.
The figural analogy — A is to B as C is to what — and its grid generalisation, the matrix, emerged from exactly this reasoning. The instrument that resulted was designed on theory before it was validated in practice, which was unusual then and remains so.
Raven's Progressive Matrices, 1938
John C. Raven, working with Spearman's ideas, published the Progressive Matrices in 1938. Each item is a grid of abstract figures with one cell missing; the candidate selects the piece that completes the pattern.
The design is austere and deliberate. No words, no numbers, no recognisable objects, instructions demonstrable by gesture, and items ordered so that each teaches something about the next.
Why it became the reference instrument
- It is the most g-loaded single test in wide use. No other instrument of comparable brevity correlates as strongly with full-scale scores.
- It crosses languages. One form, usable almost anywhere, which nothing verbal manages.
- It is administratively cheap. Group-administrable, objectively scored, requiring no trained interviewer.
- It has enormous accumulated data. Eighty-plus years of norms across dozens of populations.
Cattell and the Fluid–Crystallised Split
Raymond Cattell, another of Spearman's students, proposed in 1943 that general ability divides into two: fluid intelligence, reasoning applied to novel problems, and crystallised intelligence, accumulated knowledge and learned procedure.
The distinction explained a puzzle in the data — the two follow different trajectories across the lifespan, fluid ability peaking early and declining, crystallised ability holding or rising well into later life.
Cattell also built his own Culture Fair Intelligence Test in the 1940s, in the same figural tradition and with the same ambition of stripping content out.
Into Employment Selection
Abstract reasoning entered commercial hiring properly in the post-war decades, as test publishers built batteries around the same item forms.
The trajectory ran through occupational psychology in Britain and Europe particularly, where figural reasoning subtests became standard in graduate selection — a position they still hold, now delivered online and frequently adaptive.
What the meta-analytic evidence added
Frank Schmidt and John Hunter's 1998 synthesis of eighty-five years of selection research reported general mental ability as among the strongest single predictors of job performance and of training success, across job types.
That paper is the reason abstract reasoning survived the shift to competency-based hiring. Whatever its limitations, the predictive validity was documented, and the alternatives frequently were not.
Are These Tests Culture-Fair?
The phrase attached itself early and has been retreating ever since.
The strongest counter-evidence is the Flynn effect — the sustained rise in scores across the twentieth century, documented by James Flynn from the 1980s. Gains were largest on exactly the abstract, supposedly culture-free instruments, including Raven's.
A genuinely culture-independent measure of a stable trait should not drift upward by that much within a few generations. Something environmental is being measured alongside reasoning — schooling, exposure to abstract symbols, and familiarity with test formats are the usual candidates.
The defensible claim
Culture-reduced, not culture-fair. Removing language and content eliminates several sources of bias, and that is a real achievement. It does not eliminate the effects of an education that teaches people to think in diagrams.
Are They Still Trusted?
Yes, with the qualifications the century has attached to them.
Raven's remains in active clinical and research use. Abstract reasoning subtests remain standard in graduate selection. The predictive evidence has held up better than most findings of comparable age.
What has changed is the framing. These are measures of one important ability, interpreted against norms, in combination with other evidence — not the single-number verdicts on human worth that the field's least careful decades treated them as.