Aristotle and the First Formal System
Logic begins as a written discipline with Aristotle's Prior Analytics, composed in the fourth century BC. Its achievement was to treat the form of an argument as separable from its content.
The syllogism made that separable form visible. "All A are B; all B are C; therefore all A are C" is valid whatever A, B and C stand for — and once that is written down, arguments can be assessed mechanically.
The Stoics, principally Chrysippus, added the other half a century later: propositional logic, the reasoning of if–then rather than of categories. Conditional items on a modern test descend from the Stoic tradition, not the Aristotelian one.
Two Thousand Years of Almost Nothing
Aristotelian logic was preserved, taught, and elaborated through the medieval curriculum, where it formed part of the trivium and was examined by disputation — the oral defence of a thesis against objections.
That is an examination, but not a test. It had no standardised items, no comparison group, and no evidence that performance predicted anything outside the examination hall.
The nineteenth-century break
George Boole's Laws of Thought in 1854 recast logical relations as algebra. Gottlob Frege's Begriffsschrift in 1879 introduced quantification and built the predicate logic that the twentieth century ran on.
Neither man was interested in measuring people. But formalising inference into rules that could be applied mechanically is what later made it possible to write logical items with unambiguous answers.
Binet and the Move to Measurement
The 1905 scale of Alfred Binet and Théodore Simon, built to identify Parisian schoolchildren needing support, contained the first items recognisable as logical reasoning tasks: absurdity detection, analogies, and questions requiring inference from a described situation.
Binet's insight was procedural rather than logical. He established that reasoning could be sampled with short standardised tasks and scored against what children of a given age typically managed — the norm group, which is what turns an examination into a test.
1917: Reasoning at Industrial Scale
The US Army's group testing programme under Robert Yerkes produced Army Alpha, whose eight subtests included analogies, disarranged sentences, and number series — all inference tasks delivered to a hall of recruits against a clock.
The nonverbal Army Beta carried pattern and series items for recruits who were illiterate or did not read English. The pairing established the format that dominates today: logical structure delivered without language, so that reading ability is not silently part of the score.
The programme's interpretive conclusions were badly wrong and were used to support restrictive immigration policy in the 1920s. The format survived the conclusions, which is why claims about group differences from ability testing still warrant scrutiny about exposure rather than capacity.
Spearman, Raven, and the Matrix
Charles Spearman's 1904 paper proposed that performance across mental tasks shares a general factor, g. The search for the purest possible measure of it drove the next thirty years of test design.
John C. Raven's Progressive Matrices, published in 1938, was the result. A matrix with one cell missing, filled by extracting the rule governing rows and columns — no words, no arithmetic, no stored knowledge of any kind.
Why the format persisted
Matrix items turned out to be the most heavily g-loaded single format available, and they are administrable across languages with minimal translation. That combination is why nearly every modern battery contains a matrix section in some form.
The "culture-fair" label attached to them has aged less well. Raven's shows some of the largest Flynn-effect gains of any instrument — scores rising substantially across generations — which is difficult to reconcile with a measure supposedly independent of environment.
Wechsler and the Clinical Route
David Wechsler's adult scale of 1939 approached logical reasoning through several routes at once — Similarities for verbal abstraction, Picture Arrangement for sequential inference, and later Matrix Reasoning.
The fourth edition in 2008 restructured the scale into four indexes, placing Matrix Reasoning inside Perceptual Reasoning alongside Block Design and Visual Puzzles. That grouping is an explicit statement that non-verbal logical inference belongs with spatial reasoning rather than with verbal ability.
Wason and the Discovery That Reasoning Is Content-Dependent
Peter Wason's selection task, introduced in 1966, is the most consequential single experiment in the psychology of reasoning.
Participants were shown four cards and asked which to turn over to test a conditional rule. The overwhelming majority chose incorrectly, and the specific error — checking the case that could only confirm the rule rather than the one that could break it — proved remarkably robust.
The twist that mattered
The same logical problem, dressed in a familiar social rule about drinking age and identification, was solved easily by most people. Identical structure, radically different performance.
This is the finding that complicates every logical test ever built. Logical competence is not a single dial; it is heavily dependent on how the problem is framed — and the abstract framing that tests favour is the framing people handle worst.
The LSAT and the Employer Era
The Law School Admission Test, first administered in 1948, became the most influential logical reasoning instrument in the world, because its stakes made its item design worth perfecting.
Its Logical Reasoning sections — short arguments followed by questions about assumptions, flaws, and inference — are the template most employer critical-reasoning tests imitate. The Analytical Reasoning section, universally known as the logic games, ran constraint puzzles of unusual difficulty.
Those games were retired from August 2024, following a settlement over accessibility for blind and low-vision candidates, and replaced with an additional Logical Reasoning section. The instrument moved decisively toward argument analysis and away from formal puzzles.
Occupational testing
Saville and Holdsworth, founded in the UK in 1977, built the graduate-screening formats still dominant today, pairing abstract matrix items with prose critical-reasoning passages. Unproctored online delivery in the 2000s added randomised item banking and supervised verification retests to manage the obvious integrity gap.
Are These Tests Still Trusted?
They remain in wide use, and the evidence base has been revised rather than overturned.
Schmidt and Hunter's 1998 meta-analysis put general mental ability at the top of the predictor hierarchy, a claim repeated in textbooks for two decades. Sackett, Zhang, Berry and Lievens' 2022 re-examination argued the corrections had inflated it and produced markedly lower validities — lower, but still competitive.
The defensible position sits between the 1998 enthusiasm and dismissal: logical tests measure something real, they work better as one input among several than as a gate, and Wason's finding that framing dominates performance has never stopped applying.