Before There Were Tests
Formal verbal ability testing is barely more than a century old, but the idea behind it is older. Imperial China's civil service examinations, running in various forms for well over a thousand years, were essentially verbal: candidates were assessed on composition, classical exegesis, and argument.
What those examinations lacked was the machinery that makes a modern test a test — standardised administration, a norm group, and any way of knowing whether the result predicted anything.
Francis Galton's anthropometric laboratory in 1880s London attempted measurement of a sort, but measured the wrong things: reaction time, grip strength, sensory discrimination. James McKeen Cattell's "mental tests" of the 1890s followed the same sensory logic. The results correlated with almost nothing anyone cared about, and by the turn of the century the sensory programme had run out of road.
Binet: The Practical Break
The break came in Paris. In 1905, Alfred Binet and Théodore Simon published a scale commissioned to solve an administrative problem: identifying schoolchildren who needed additional support, without relying on teachers' impressions.
Their solution abandoned sensory measures entirely in favour of tasks that looked like the work of thinking — following instructions, defining words, spotting absurdities in sentences, repeating digit strings, and reasoning about everyday situations.
A large share of those tasks were verbal, and that was no accident. Binet's items sampled judgement as it is ordinarily exercised, and ordinary judgement is conducted in language. The scale worked in the narrow sense that mattered: it ranked children in a way that matched their observed classroom performance.
It also carried an early caution that later practice discarded — Binet regarded his scale as a diagnostic instrument for identifying children needing help, not as a measure of fixed capacity.
Lewis Terman's Stanford revision of 1916 brought the scale to the United States, extended it upward through adolescence, and attached the ratio-based intelligence quotient that gave the field both its shorthand and a century of trouble.
1917: Testing at Industrial Scale
The First World War turned a clinical instrument into a mass one. The US Army needed to sort a very large intake quickly, and a committee under Robert Yerkes produced two group-administered tests: Army Alpha for literate English-reading recruits, and Army Beta, deliberately nonverbal, for those who were illiterate or did not read English.
Alpha's structure is recognisably the ancestor of every graduate verbal test in use today: synonym-antonym pairs, analogies, sentence rearrangement, and information items, all delivered on paper to a room full of candidates working against a clock.
The existence of Beta as a separate instrument is itself the important historical fact — the designers understood that a verbal test measures verbal ability plus language exposure, and that the second term becomes an artefact when the candidate's language is not the test's.
The wartime programme's interpretive conclusions were badly wrong and were used to support restrictive immigration arguments in the 1920s. The format survived; the conclusions did not, and the episode remains the standing reason that verbal instruments require scrutiny for group differences arising from exposure rather than ability.
Thurstone and the Case for Separate Abilities
Through the 1920s and 1930s the field argued about structure. Spearman's 1904 general factor implied that a verbal score was mostly a window onto g. Louis Thurstone's 1938 Primary Mental Abilities disagreed, isolating a set of distinguishable factors — among them Verbal Comprehension, Word Fluency, Number, Space, Memory, Perceptual Speed, and Reasoning.
The practical consequence was the profile. If verbal ability is a distinguishable factor rather than a slice of one general capacity, then reporting a single overall score discards information, and a candidate strong in verbal comprehension but weak in spatial reasoning is telling you something an aggregate would hide. Every multi-scale battery on the market descends from that argument.
Wechsler: Verbal Ability Gets Its Own Index
David Wechsler's adult scale, first published in 1939 and reissued as the WAIS in 1955, made the verbal-versus-performance split structural: two halves, separately scored. That architecture held for decades and eventually proved too coarse, because the verbal half was mixing crystallised knowledge with working memory.
The modern resolution came with the fourth edition in 2008, which replaced the old halves with four index scores. Verbal Comprehension — built from Similarities, Vocabulary, and Information — became the clean measure of crystallised verbal ability, while Arithmetic and Digit Span moved into a separate Working Memory Index where they belonged.
The Verbal Comprehension Index is the reference standard against which shorter verbal instruments are still implicitly judged.
Cattell, Horn, and Why Verbal Scores Age Well
Raymond Cattell's distinction between fluid and crystallised intelligence, developed from the 1940s and elaborated with John Horn in the 1960s, explained something the verbal data had been showing for years: verbal scores hold up across adulthood in a way that novel problem-solving scores do not.
Vocabulary and verbal comprehension continue to accrue with exposure long after the abilities that depend on processing speed and working memory have begun their decline.
James Flynn's work from the 1980s onward added a second historical wrinkle. The generational score gains that carry his name have been consistently larger on abstract and fluid measures than on vocabulary and general-information items — a pattern that fits the crystallised-versus-fluid distinction and complicates any naive comparison of scores across decades.
SHL and the Employer Era
Occupational testing became an industry in its own right in the 1970s. Saville and Holdsworth — SHL — was founded in the United Kingdom in 1977 and built the graduate-recruitment format that dominates today: short, sharply timed verbal batteries in which a business-style passage is followed by statements to be judged true, false, or "cannot say."
That third option is the design's defining feature and the source of most candidate error. It converts the task from reading comprehension into evidential discipline: the question is not whether a statement is true, but whether this passage establishes it. Competing publishers built variants, the graduate milk-round standardised around the format, and it has since been ported wholesale to online delivery.
Online Delivery and the Cheating Problem
Moving tests to unproctored online administration in the 2000s solved logistics and created an obvious integrity gap.
The standard industry response was randomised item banking — drawing each candidate's paper from a large pool so that no two sittings are identical — usually paired with a shorter verification test taken under supervision at a later stage, whose score is checked for consistency against the unsupervised one.
Item-response-theory scoring, which estimates ability from the difficulty of the items a candidate got right rather than from a raw total, made that comparison statistically workable across non-identical papers. Most large-volume graduate verbal testing today runs on some version of this two-stage architecture.
Are These Tests Still Trusted?
Cognitive measures remain among the more predictive tools in selection research, but the size of the effect has been actively revised. Schmidt and Hunter's 1998 meta-analysis placed general mental ability at the top of the predictor hierarchy, and that figure was repeated in textbooks for two decades.
Sackett, Zhang, Berry and Lievens' 2022 re-examination argued that the standard corrections had been applied in a way that inflated the estimates, and produced substantially lower corrected validities — lower, but not negligible, and still competitive with most alternatives.
The practical position among careful practitioners is now narrower than the 1998 enthusiasm and broader than the sceptics' dismissal: verbal reasoning tests measure something real and job-relevant, they are more defensible as one input among several than as a gate, and the historical lesson from 1917 — that a verbal test is always partly a language-exposure test — has never stopped being true.