Before Anyone Measured Space
Spatial skill was recognised long before it was tested. Apprenticeship in stonemasonry, carpentry, navigation and cartography selected for it ruthlessly β but by watching people work, not by scoring them.
The shift to measurement required two things that did not exist before the twentieth century: standardised items, and a norm group to read a score against.
Binet and the First Non-Verbal Items
The 1905 scale of Alfred Binet and ThΓ©odore Simon, built to identify Parisian schoolchildren needing extra support, included tasks that did not depend on reading β copying a figure, comparing forms, and reproducing a design from memory.
These were not conceived as spatial measures. They were included because Binet needed items that worked for children whose verbal skills were exactly what was in question, and non-verbal material was the practical solution.
That accident set a pattern the field never abandoned: spatial items became the standard way of measuring reasoning without language getting in the way.
1917: The Army Beta
The US Army's wartime testing programme under Robert Yerkes needed to assess a very large number of recruits, many of whom could not read English.
Army Beta, the non-verbal counterpart to Army Alpha, carried maze tracing, picture completion, and figure-based items delivered by demonstration rather than written instruction.
The programme demonstrated that visual reasoning tasks could be administered at scale, to groups, against a clock. That is the format employer testing still uses.
The caution the programme also demonstrated
The interpretive conclusions drawn from these scores were badly wrong, and were used in the 1920s to support restrictive immigration policy. Scores from candidates unfamiliar with the medium of testing were read as evidence about capacity.
The format survived; the conclusions did not. It is the reason claims about group differences on non-verbal tests still warrant scrutiny about exposure and familiarity before anything else.
Kohs and the Block Design Task
In 1920 Samuel Kohs published the Block Design Test: coloured cubes to be arranged to match a printed pattern.
The design was elegant because it made the reasoning observable. An examiner could watch the strategy unfold β whether the candidate worked from the corners, decomposed the pattern into cells, or drifted without a plan.
David Wechsler adapted the task into his adult scale in 1939, and Block Design has remained on the Wechsler scales through every revision since. Few instruments in psychology have been that durable.
Thurstone and Space as a Primary Ability
Louis Leon Thurstone's 1938 factor analysis was a direct challenge to the idea that a single general factor explained mental performance.
His seven primary mental abilities included Space as an independent factor, distinct from Number, Verbal Comprehension, Word Fluency, Memory, Perceptual Speed and Reasoning.
The claim had a practical consequence: if spatial ability was separable, it could be strong where others were weak, and a selection system measuring only verbal and numerical skill would miss people.
Raven and the Matrix
John C. Raven published the Progressive Matrices in 1938 β a matrix of figures with one cell missing, completed by extracting the rule governing rows and columns.
Raven's is a reasoning test delivered in visual form rather than a spatial test in the strict sense; it asks for rule extraction, not rotation. But it fixed the visual multiple-choice item as the default vehicle for measuring reasoning without language.
The culture-fair claim and its limits
Matrices were promoted as culture-fair, and the portability across languages is genuine. The label nonetheless overstates the case.
Raven's shows some of the largest Flynn-effect gains of any instrument, with average scores rising substantially across generations within the same populations. Something environmental is clearly being captured β most plausibly familiarity with abstract diagrams and testing conventions.
The Selection Era: Form Boards and Engineering
Between the 1930s and 1950s spatial testing moved decisively into occupational selection, and the instruments built then are still recognisable.
- The Minnesota Paper Form Board Test. Assembling flat pieces mentally into a whole shape; long used in technical and engineering selection.
- The Differential Aptitude Tests. First published in 1947, with a Space Relations subtest built around folding flat nets into solids.
- The Bennett Mechanical Comprehension Test. Published in 1940, pairing spatial visualisation with physical reasoning about mechanisms.
These were built for a specific claim: that spatial visualisation predicted success in technical training in a way that general ability alone did not. That claim has held up better than most from the period.
Shepard, Metzler, and Rotation as a Measurable Process
The most important single experiment in the field was published by Roger Shepard and Jacqueline Metzler in 1971.
Participants judged whether two three-dimensional figures were the same object or mirror images. Response time rose in almost perfect proportion to the angular difference between them.
That linear relationship was the finding. It indicated that people were performing something like an actual rotation, at a roughly constant rate, rather than comparing abstract descriptions β turning a private mental act into something with a measurable speed.
What followed from it
The Vandenberg and Kuse Mental Rotations Test of 1978 turned the paradigm into a standardised paper instrument, and it became the workhorse of the research literature β including the studies on sex differences and on training effects.
The Purdue Spatial Visualization Test, developed in the 1970s for engineering education, did the same for rotation and view-taking in a technical training context.
Screens, Surgery and the Modern Era
Two developments revived interest after a quiet period.
The first was minimally invasive surgery. Operating from a two-dimensional screen while manipulating instruments in three dimensions is a spatial task of an unusually demanding kind, and it generated an active literature on whether spatial testing predicts who acquires the skill quickly.
The second was Wai, Lubinski and Benbow's 2009 longitudinal analysis, which followed a very large American cohort across decades and showed spatial ability predicting STEM entry and achievement beyond mathematical and verbal scores. It made the case that decades of selection had been quietly discarding relevant information.
Are These Tests Culture-Fair?
Fairer than verbal tests, and not fair in any absolute sense.
What they remove is real: no language to translate, no culturally specific vocabulary, no reliance on schooling in a particular subject. That is why they are used in cross-national research.
What they do not remove is exposure. Familiarity with diagrams, technical drawing conventions, construction toys, and multiple-choice testing itself all plausibly contribute, and none of these are evenly distributed.
The defensible term is culture-reduced. That is a meaningful improvement over a vocabulary test, and it is not the same as measuring something untouched by environment.